Related Experiment Video
Updated: May 4, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Recognition of the script in Serbian documents using frequency occurrence and co-occurrence analysis.
Darko Brodić1, Zoran N Milivojević2, Cedomir A Maluckov1
1Technical Faculty in Bor, University of Belgrade, Vojske Jugoslavije 12, 19210 Bor, Serbia.
This study presents a novel method for distinguishing Serbian Latin and Cyrillic scripts using statistical analysis of letter occurrences and co-occurrences. The approach effectively identifies script types in various document formats.
Area of Science:
- Computational Linguistics
- Natural Language Processing
- Document Analysis
Background:
- Serbian language documents can utilize either Latin or Cyrillic scripts.
- While visually similar, these scripts exhibit distinct statistical properties.
- Accurate script identification is crucial for text processing and analysis.
Purpose of the Study:
- To develop and evaluate a method for automatically distinguishing between Serbian Latin and Cyrillic scripts.
- To leverage statistical measures of letter occurrence and co-occurrence for script recognition.
Main Methods:
- Modeling individual letters based on their position within the baseline area to determine script type.
- Performing frequency analysis on the occurrence of modeled script types.
- Computing a co-occurrence matrix for script elements.
- Utilizing the co-occurrence matrix as a criterion for script differentiation.
Main Results:
- Significant statistical dissimilarities were observed in the occurrence of modeled letters between Latin and Cyrillic scripts.
- The co-occurrence matrix analysis provided a strong basis for distinguishing between the two scripts.
- Experiments on a diverse database of printed and web documents yielded encouraging results.
Conclusions:
- The proposed method effectively differentiates Serbian Latin and Cyrillic scripts based on statistical script analysis.
- The approach demonstrates high accuracy and robustness across various document types.
- This technique offers a reliable solution for automatic script identification in Serbian texts.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
14:34A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Related Concept Videos
Relative Frequency Histogram
Relative Frequency Distribution
Social Scripts
Determination of Expected Frequency
Expected Frequencies in Goodness-of-Fit Tests
Frequency-dependent Selection