Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Vision01:24

Vision

54.2K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
54.2K
Genetic Lingo01:11

Genetic Lingo

103.5K
Overview
103.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

G6PD inhibition prevents abdominal aortic aneurysm formation in mice induced by Ang II plus high salt.

Atherosclerosis·2026
Same author

Metal-driven nanoassembly of hexahistidine-tagged melittin enables superior phytopathogen biofilm degradation with attenuated toxicity.

Journal of nanobiotechnology·2026
Same author

Piperidine-functionalized formononetin derivatives: Design, synthesis and evaluation of antibacterial activity, modes and phenotypes.

Pest management science·2026
Same author

Single cell and single nucleus RNA sequencing in liver tissues: applications and prospects in model and non-model organisms.

Frontiers in genetics·2026
Same author

Human obese gut microbiota induces lipid metabolic disorder and hypophagia without weight gain in normocaloric mice.

Science China. Life sciences·2026
Same author

Design, Synthesis, and Biological Activity Studies of Flavonol Derivatives Containing Isopropanolamine and Piperazine.

Journal of agricultural and food chemistry·2026

Related Experiment Video

Updated: Aug 6, 2025

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

9.0K

LCM-Captioner: A lightweight text-based image captioning method with collaborative mechanism between vision and text.

Qi Wang1, Hongyu Deng1, Xue Wu2

  • 1State Key Laboratory of Public Big Data, College of Computer Science and Technology, Guizhou University, China.

Neural Networks : the Official Journal of the International Neural Network Society
|March 19, 2023
PubMed
Summary

This study introduces LCM-Captioner, a lightweight method for text-based image captioning (TextCap). It efficiently integrates visual and textual information for improved image understanding and reduced computational costs.

Keywords:
Collaborative attention mechanismFeature transformationLightweight networkMultimodal informationText-based image captioning

More Related Videos

A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
12:39

A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers

Published on: January 18, 2020

7.7K
SIVQ-LCM Protocol for the ArcturusXT Instrument
07:37

SIVQ-LCM Protocol for the ArcturusXT Instrument

Published on: July 23, 2014

8.7K

Related Experiment Videos

Last Updated: Aug 6, 2025

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

9.0K
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
12:39

A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers

Published on: January 18, 2020

7.7K
SIVQ-LCM Protocol for the ArcturusXT Instrument
07:37

SIVQ-LCM Protocol for the ArcturusXT Instrument

Published on: July 23, 2014

8.7K

Area of Science:

  • Computer Vision
  • Artificial Intelligence
  • Natural Language Processing

Background:

  • Existing image captioning models often overlook textual information within images.
  • Complex architectures in current methods lead to inefficiency and deployment challenges.
  • Text-based image captioning (TextCap) requires models to process both visual and textual content for deeper comprehension.

Purpose of the Study:

  • To develop a lightweight and efficient TextCap method balancing performance and computational cost.
  • To address the limitations of existing complex models in multimodal information integration.
  • To improve the semantic alignment between visual and textual data for enhanced image description.

Main Methods:

  • Proposed TextLighT: a feature-lightening transformation for learning low-dimensional multimodal representations and reducing memory usage.
  • Developed VTCAM (Visual-Text Collaborative Attention Module): a module for semantic alignment of visual and textual information.
  • Implemented a lightweight captioning method, LCM-Captioner, integrating TextLighT and VTCAM.

Main Results:

  • The proposed LCM-Captioner method demonstrates effectiveness on the TextCaps dataset.
  • TextLighT successfully maps features to lower dimensions, reducing memory costs.
  • VTCAM facilitates semantic alignment, uncovering key visual and textual content.

Conclusions:

  • LCM-Captioner offers a high-efficiency, high-performance solution for text-based image captioning.
  • The developed techniques effectively model the relationship between vision and text.
  • The method overcomes deployment challenges associated with complex architectures.