Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

2.6K
2.6K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Fuzzy-Match Repair Guided by Quality Estimation.

IEEE transactions on pattern analysis and machine intelligence·2020
Same author

Kalman filters improve LSTM network performance in problems unsolvable by traditional recurrent nets.

Neural networks : the official journal of the International Neural Network Society·2003
See all related articles

Related Experiment Video

Updated: Jul 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

596

Non-Fluent Synthetic Target-Language Data Improve Neural Machine Translation.

Victor M Sanchez-Cartagena, Miquel Espla-Gomis, Juan Antonio Perez-Ortiz

    IEEE Transactions on Pattern Analysis and Machine Intelligence
    |November 17, 2023
    PubMed
    Summary

    Generating synthetic data for neural machine translation can be improved by using non-fluent target sentences within a multilingual framework. This approach enhances translation performance, robustness, and reduces hallucinations.

    More Related Videos

    Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia
    10:15

    Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia

    Published on: July 2, 2013

    17.9K
    Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
    09:09

    Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

    Published on: September 27, 2024

    464

    Related Experiment Videos

    Last Updated: Jul 11, 2025

    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
    03:14

    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

    Published on: December 6, 2024

    596
    Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia
    10:15

    Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia

    Published on: July 2, 2013

    17.9K
    Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
    09:09

    Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

    Published on: September 27, 2024

    464

    Area of Science:

    • Natural Language Processing
    • Machine Translation
    • Artificial Intelligence

    Background:

    • Neural machine translation (NMT) models require large parallel corpora for effective training.
    • Data augmentation techniques, like generating synthetic parallel sentences, are crucial when parallel data is scarce.
    • Current methods often assume synthetic data must mimic in-domain parallel corpora, potentially limiting performance.

    Purpose of the Study:

    • To investigate the impact of non-fluent synthetic target sentences on NMT performance.
    • To propose a novel approach using non-fluent synthetic data within a multilingual NMT framework.
    • To evaluate the effectiveness of this method across various resource scenarios.

    Main Methods:

    • Generating synthetic parallel sentences with non-fluent target sides.
    • Integrating these synthetic sentences into a multilingual NMT system, treating them as data from another language.
    • Conducting comparative experiments against state-of-the-art synthetic data generation methods.

    Main Results:

    • The proposed method consistently improved translation performance across ten low-resource and four high-resource tasks.
    • Performance gains were observed compared to existing state-of-the-art synthetic data generation techniques.
    • The improvements were independent of the original training corpus size.

    Conclusions:

    • Non-fluent synthetic training data can enhance NMT performance when utilized in a multilingual setting.
    • This approach leads to more robust NMT systems, less susceptible to domain shift and hallucination.
    • The findings challenge the assumption that synthetic data must strictly adhere to target-side fluency for optimal results.