Related Experiment Video
Updated: Jun 16, 2026

A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Critical analysis of datasets for sign language translation
Bittor Alkain1, Adrián Núñez-Marcos1, Carlos Escolano2
1HiTZ Center - Ixa, University of the Basque Country UPV/EHU, San Sebastián, Spain.
Introduction:
In recent years, significant progress has been made in Machine Translation (MT), including multilingual and low-resource settings. However, Sign Language Translation (SLT) remains underdeveloped, largely due to the scarcity of high-quality datasets and the overreliance on a few small, widely used benchmarks. This study aims to critically assess the datasets most commonly used in SLT research to determine whether their characteristics may lead to overfitting and misleading evaluation results.
Methods:
We then conduct a detailed empirical study comparing training and test set similarity for PHOENIX14T, CSL-Daily, and LSE-Health. Using both gloss-based (TwoStream-SLT) and gloss-free (GFSLT-VLP) models, we evaluate the extent to which models memorize training data and how this affects BLEU scores.
Results:
Our analysis reveals that PHOENIX14T exhibits substantial overlap between training and test sets, leading to inflated BLEU scores and can even mask signs of overfitting. CSL-Daily shows less overlap and more robust generalization. We also show that a small subset of "training-like" sentences disproportionately contributes to BLEU scores.
Discussion:
We recommend that future SLT research move away from overused benchmarks and adopt larger, more diverse datasets such as How2Sign, CSL-News, and FLEURS-ASL. We also advocate for a shift toward gloss-free approaches and more careful interpretation of evaluation metrics, especially in low-resource settings.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy