Preprocessing Large-Scale Conversational Datasets: A Framework and Its Application to Behavioral Health Transcripts

Paz Mor Naim1, Shiri Sadeh-Sharvit2,3, Samuel Jefroykin2

  • 1Department of Cognitive and Brain Sciences, Hebrew University of Jerusalem, Mount Scopus, Jerusalem, 9190500, Israel, 972 025882888.

JMIR Formative Research
|October 24, 2025
PubMed
Summary

This study introduces a framework using large language models (LLMs) to filter noisy conversational transcripts, improving data quality for behavioral health research. The hybrid approach effectively distinguishes therapy sessions from non-sessions, enhancing data usability.