Related Concept Videos
Language and Cognition
Termination of Translation
You might also read
Related Articles
Articles linked to this work by shared authors, journal, and citation graph.
Time series for blind biosignal classification model.
Unsupervised quality estimation model for English to German translation and its application in extensive supervised evaluation.
Related Experiment Video
Updated: Apr 28, 2026

A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
iSentenizer-μ: multilingual sentence boundary detection model.
Derek F Wong1, Lidia S Chao1, Xiaodong Zeng1
1NLPCT Laboratory, Department of Computer and Information Science, University of Macau, Macau.
A new multilingual sentence boundary detection (SBD) system, iSentenizer-μ, accurately processes mixed genres and languages. It uses incremental learning to adapt without retraining, outperforming existing SBD models.
More Related Videos
08:32Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
Published on: September 5, 2019
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Area of Science:
- Natural Language Processing
- Computational Linguistics
Background:
- Sentence boundary detection (SBD) systems are typically sensitive to training data genres and languages.
- Retraining SBD models for new data requires discarding previous work and starting from scratch.
Purpose of the Study:
- To introduce iSentenizer-μ, a novel multilingual SBD system.
- To develop an adaptable SBD system capable of handling diverse text genres and languages.
Main Methods:
- Utilized an incremental tree learning architecture, specifically the i(+)Learning algorithm.
- Developed a system adaptable to various text topics and Roman-alphabet languages through incremental knowledge merging.
- Designed iSentenizer-μ to revise existing models rather than requiring complete retraining.
Main Results:
- iSentenizer-μ demonstrated high accuracy in detecting sentence boundaries across a mixture of text genres and languages.
- The system was extensively evaluated on Danish, German, English, Spanish, Dutch, French, Italian, Portuguese, Greek, Finnish, and Swedish.
- Outperformed two state-of-the-art SBD systems, Punkt and MaxEnt, on all tested datasets.
Conclusions:
- The proposed iSentenizer-μ system offers a robust and adaptable solution for multilingual SBD.
- Incremental learning enables efficient adaptation to new data, overcoming limitations of traditional retraining approaches.
- iSentenizer-μ provides superior performance compared to existing SBD systems in diverse linguistic and topical contexts.