Related Experiment Videos
Dataset of authorship attribution from short-text multi topic representative in Bahasa Indonesia
Yohan Muliono1, Ford Lumban Gaol1, Andry Chowanda2
1Computer Science Department, BINUS Graduate Program, Doctor of Computer Science Program, Bina Nusantara University, Jakarta 11480, Indonesia.
None:
Authorship attribution plays an important role in various applications, including digital forensics, Yauthorship verification, and the analysis of online textual content. While extensive benchmark datasets are available for high-resource languages such as English, publicly accessible datasets for authorship attribution in the Indonesian language remain limited. This limitation restricts systematic experimentation and comparative evaluation across different authorship attribution approaches in bahasa Indonesia. To address this gap, this data article introduces an Indonesian authorship attribution dataset consisting of short-text documents collected from social media. Short-text data were obtained from Twitter using the official Academic API, The dataset covers three topical categories politic, financial, and mix and includes anonymized author labels to preserve authorship consistency without revealing personal identities. In addition to the original text, the dataset provides explicit stylometric attributes, including word count and character count for each document, which are commonly used features in authorship attribution research. All data are consolidated into a single machine-readable CSV file to facilitate reuse, reproducibility, and further development of authorship attribution methods for the Indonesian language.
Related Concept Videos
Attribution Theory
Fundamental Attribution Error
Attribution
Inductive Reasoning
Correspondence Bias
Theory of Attribution I: Correspondent Inference Theory