Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Power transformer online monitoring and state assessment method based on 2DCNN-BiLSTM multi-source feature fusion.

PloS oneĀ·2026
Same author

Differentially Private Accelerated Distributed Algorithm for Aggregative Optimization.

IEEE transactions on neural networks and learning systemsĀ·2026
Same author

Corrigendum to "Polyploid giant cancer cells: A novel target in future cancer therapy" [Eur. J. Cell Biol. 105/1 (2026) 151529].

European journal of cell biologyĀ·2026
Same author

Amylose content of sorghum starches measured by four different methods in relation to molecular structures of sorghum amylose and amylopectin.

Carbohydrate polymersĀ·2025
Same author

Stand spatial structure promotes tree growth and sapling diversity in northern tropical karst seasonal rainforest.

Frontiers in plant scienceĀ·2025
Same author

Leaf functional trait variation and environmental filtering across hierarchical levels in complex karst peak-depression landscapes.

Tree physiologyĀ·2025

Related Experiment Video

Updated: Sep 6, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

677

An Improved Transformer-Based Neural Machine Translation Strategy: Interacting-Head Attention.

Dongxing Li1, Zuying Luo1

  • 1School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China.

Computational Intelligence and Neuroscience
|July 1, 2022
PubMed
Summary

This study introduces an interacting-head attention mechanism to improve neural machine translation (NMT) models by enabling deeper head interactions and avoiding low-rank bottlenecks, enhancing translation performance.

More Related Videos

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.9K
Construction of an Improved Multi-Tetrode Hyperdrive for Large-Scale Neural Recording in Behaving Rats
10:04

Construction of an Improved Multi-Tetrode Hyperdrive for Large-Scale Neural Recording in Behaving Rats

Published on: May 9, 2018

11.4K

Related Experiment Videos

Last Updated: Sep 6, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

677
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.9K
Construction of an Improved Multi-Tetrode Hyperdrive for Large-Scale Neural Recording in Behaving Rats
10:04

Construction of an Improved Multi-Tetrode Hyperdrive for Large-Scale Neural Recording in Behaving Rats

Published on: May 9, 2018

11.4K

Area of Science:

  • Natural Language Processing
  • Artificial Intelligence
  • Machine Learning

Background:

  • Transformer models significantly advance neural machine translation (NMT).
  • Multi-head attention is key, but faces challenges like same-subspace computation and low-rank bottlenecks.
  • Existing solutions for low-rank bottlenecks struggle with variable sequence lengths and parameter bloat.

Purpose of the Study:

  • To propose an interacting-head attention mechanism for NMT.
  • To address limitations of standard multi-head attention in transformers.
  • To enhance NMT performance by fostering deeper interactions across attention heads.

Main Methods:

  • Introduced an interacting-head attention mechanism.
  • Employed low-dimension computations in different token subspaces.
  • Selected an appropriate number of heads to prevent low-rank bottlenecks.

Main Results:

  • Achieved performance improvements across multiple machine translation tasks (IWSLT2016 DE-EN, WMT17 EN-DE, WMT17 EN-CS).
  • Demonstrated gains in BLEU, WER, METEOR, ROUGE_L, CIDEr, and YiSi scores on evaluation and test sets.
  • Outperformed original multi-head attention in all tested translation directions.

Conclusions:

  • The interacting-head attention mechanism effectively enhances NMT performance.
  • The proposed method overcomes limitations of standard multi-head attention.
  • This approach offers a more robust and efficient way to utilize attention heads in NMT.