Related Experiment Video
Updated: Jul 11, 2026

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved (Non-model) Organisms
Published on: May 9, 2017
A reappraisal of sentence and token splitting for life sciences documents
Katrin Tomanek1, Joachim Wermter, Udo Hahn
1Jena University Language and Information Engineering (JULIE) Lab, Friedrich-Schiller-Universität Jena, Germany.
Abstract:
Natural language processing of real-world documents requires several low-level tasks such as splitting a piece of text into its constituent sentences, and splitting each sentence into its constituent tokens to be performed by some preprocessor (prior to linguistic analysis). While this task is often considered as unsophisticated clerical work, in the life sciences domain it poses enormous problems due to complex naming conventions. In this paper, we first introduce an annotation framework for sentence and token splitting underlying a newly constructed sentence- and token-tagged biomedical text corpus. This corpus serves as a training environment and test bed for machine-learning based sentence and token splitters using Conditional Random Fields (CRFs). Our evaluation experiments reveal that CRFs with a rich feature set substantially increase sentence and token detection performance.
Related Concept Videos
Pre-mRNA Processing: RNA Splicing
RNA Splicing
Nonsense-mediated mRNA Decay
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Interpreting ¹H NMR Signal Splitting: The (n + 1) Rule
Improving Translational Accuracy
Improving Translational Accuracy
