Related Experiment Videos
Cyberbullying victimization identification and large language model-assisted assessment: a study of cyberbullying
Xingyun Liu1,2, Yuehan Liao1,2, Fan Feng1,2
1Key Laboratory of Adolescent Cyberpsychology and Behavior(CCNU), Ministry of Education, Wuhan, China.
Introduction:
Cyberbullying poses a global mental health threat, yet its accurate identification remains challenging due to biases in self-reporting and help-seeking barriers.
Methods:
Based on large-scale social media data, the present study constructed a Chinese cyberbullying victimization lexicon from three dimensions-cyberbullying methods (the types of cyberbullying experienced by the individual), perceived harm (the harm perceived by the victim), and coping strategies (the behavioral responses adopted by the victim)-using Weibo texts, psychological lexicons, and cyberbullying questionnaires. This approach aims to improve the precision of victim identification and facilitate timely intervention. Lexicon validity was evaluated by examining correlations between word-frequency statistics derived from 500 Weibo posts and expert ratings (n = 3). In addition, based on 3,442 RedNote posts, we preliminarily explored the lexicon's cross-platform applicability. To assess whether DeepSeek-R1 and GPT-4o could function as research assistants in dictionary construction, the present study replicated the manual lexicon development process-including text classification, vocabulary screening, weight assignment, and victimization severity assessment-and compared model outputs with human evaluations under simple and complex prompts using Cohen's Kappa, intraclass correlation coefficients (ICC), recall, and precision.
Results:
(1) The lexicon comprised 442 words across three dimensions: cyberbullying methods, perceived harm, and coping strategies. The lexicon demonstrated strong validity in identifying cyberbullying victimization expressions in social media text across each of the three sub-dimensions and the overall dimension (cyberbullying methods: r=0.500, p < 0.001; perceived harm: r=0.408, p < 0.001; coping strategies: r=0.509, p < 0.001; overall cyberbullying victimization expression: r = 0.870, p < 0.001), and effectively assessed the degree of cyberbullying victimization (r = 0.533, p < 0.001); (2) While DeepSeek-R1 performed well on small-scale text classification (Kappa = 0.775-0.781), both models showed significant limitations in large-scale processing (12,600 entries). Vocabulary selection and weight assignment tasks revealed substantial discrepancies with human evaluation (Kappa as low as -0.874), though DeepSeek-R1 achieved moderate consistency in dimension partitioning (Kappa = 0.780).
Discussion:
As the first Chinese cyberbullying victimization dictionary, this research demonstrates that large language models offer preliminary utility in structured tasks but require human oversight for complex, large-scale research phases, supporting a human-machine collaborative approach for optimal outcomes.
Related Concept Videos
Bullying
Language and Cognition
Social Foundations of Self IV: Self in Digital Communication