Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

15.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.4K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.8K
3.8K
Language Development01:22

Language Development

1.0K
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
1.0K
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

460
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
460

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

MSCANet: a cross-attention-based multi-scale convolutional fusion neural network for EEG motor imagery classification.

Cognitive neurodynamics·2026
Same author

Road pedestrian detection and tracking algorithm based on improved YOLOv5s and DeepSORT.

PloS one·2025
Same author

Surface defect detection on industrial drum rollers: Using enhanced YOLOv8n and structured light for accurate inspection.

PloS one·2025
Same author

Structurally Oriented Carbon Dots as ROS Nanomodulators for Dynamic Chronic Inflammation and Infection Elimination.

ACS nano·2024
Same author

Three-dimensional shape measurement method based on composite cyclic phase coding.

Applied optics·2023
Same author

Directional Migration and Distribution of Magnetic Microparticles in Polypropylene-Matrix Magnetic Composites Molded by an Injection Molding Assisted by External Magnetic Field.

Materials (Basel, Switzerland)·2022

Related Experiment Video

Updated: Mar 25, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K

Attention re-alignment in multimodal large language models via intermediate-layer guidance.

Yanming Chen1, Pandong Wang1, Guofeng Qin1

  • 1Tongji University, Shanghai, China.

Scientific Reports
|March 24, 2026
PubMed
Summary

Multimodal large language models (MLLMs) struggle with visual details due to language bias. The proposed Attention Re-alignment module (ARA) enhances visual grounding by re-weighting attention maps, improving performance on visual question answering tasks.

Keywords:
Attention alignmentHallucination mitigationMultimodal large language modelsVisual question answering

Related Experiment Videos

Last Updated: Mar 25, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K

Area of Science:

  • Computer Science
  • Artificial Intelligence
  • Machine Learning

Background:

  • Multimodal large language models (MLLMs) excel at visual question answering (VQA).
  • MLLMs often fail to focus on fine-grained visual details during image analysis.
  • This limitation stems from language priors diluting visual attention in deeper model layers.

Purpose of the Study:

  • To address the issue of MLLMs neglecting fine-grained visual details.
  • To enhance the visual grounding capabilities of existing MLLMs.
  • To improve the accuracy and sensitivity of MLLMs in VQA tasks.

Main Methods:

  • Proposing a plug-and-play Attention Re-alignment (ARA) module.
  • Conducting layer-wise analysis of attention distributions in image-centric heads.
  • Employing a confidence-aware layer selection based on attention peak and entropy.
  • Dynamically aggregating attention maps from informative layers to guide semantic mask generation.

Main Results:

  • The ARA module effectively enhances suppressed visual grounding.
  • Semantic masks generated using ARA emphasize salient visual regions and suppress noise.
  • Consistent performance improvements observed across multiple VQA benchmarks.
  • Demonstrated effectiveness in improving MLLMs' sensitivity to visual details.

Conclusions:

  • The ARA module is a viable solution for improving MLLMs' visual detail perception.
  • ARA can be seamlessly integrated into existing MLLM architectures.
  • Enhanced visual grounding via ARA leads to superior performance in VQA tasks.