Related Experiment Video
Updated: Jun 13, 2025

Assessing the Coherence of Parents' Short Narratives Regarding their Child Using the Five-Minute Speech Sample Procedure
Published on: September 19, 2019
Balinese story texts dataset for narrative text analyses.
I Made Satria Bimantara1, Diana Purwitasari1, Ngurah Agus Sanjaya Er2
1Informatics Department, Faculty of Intelligent Electrical and Informatics Technology, Institut Teknologi Sepuluh Nopember, Surabaya 60111, Indonesia.
This study introduces the first annotated Balinese story dataset for computational linguistic tools, enabling character identification and classification in low-resource languages. The dataset aids in developing advanced machine learning models for narrative text analysis.
Area of Science:
- Computational Linguistics
- Natural Language Processing (NLP)
- Artificial Intelligence (AI)
Background:
- Computational linguistic tools for narrative text analysis are advancing, but are primarily English-focused due to data scarcity in other languages.
- Character identification is a critical first step for deeper narrative analysis, yet remains challenging in diverse linguistic contexts.
- Low-resource languages like Balinese lack sufficient annotated datasets for developing such analytical tools.
Purpose of the Study:
- To present the first annotated Balinese story texts dataset for narrative text analysis.
- To facilitate character identification, alias clustering (named entity linking), and character classification in Balinese narratives.
- To support the development of computational linguistic tools and machine learning models for low-resource languages.
Main Methods:
- Manual annotation of 120 Balinese stories by native speakers, including a sociolinguistics expert.
- Creation of four sub-datasets for character identification (word and sentence level), alias clustering, and protagonist/antagonist classification.
- Calculation of inter-annotator agreement using Cohen's Kappa, Jaccard Similarity, and Mean F1-score to ensure dataset reliability.
Main Results:
- A comprehensive dataset comprising 89,917 annotated words, 6,634 annotated sentences, and 930 character groups.
- Successful annotation for character identification, alias clustering, and classification of 848 character groups into protagonists (66.16%) and antagonists (33.84%).
- Demonstrated reliability and consistency of the dataset through high inter-annotator agreement scores.
Conclusions:
- The developed Balinese narrative dataset is a valuable resource for advancing computational linguistics and AI in low-resource languages.
- This dataset enables enhanced research in character identification, network development, and relationship extraction within narrative texts.
- It provides a foundation for building sophisticated machine learning and deep learning models for analyzing non-English narrative structures.
Related Concept Videos
Data Collection by Survey
Life Histories
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Data Collection I
Data Collection III
The principles to begin the physical assessment include conducting a comprehensive or problem-related history in a quiet, well-lit room, emphasizing privacy and comfort for the...
Longitudinal Research

