Related Experiment Video
Updated: Jan 14, 2026

08:00
Decoding Natural Behavior from Neuroethological Embedding
Published on: October 3, 2025
587
Preprocessing Large-Scale Conversational Datasets: A Framework and Its Application to Behavioral Health Transcripts
Paz Mor Naim1, Shiri Sadeh-Sharvit2,3, Samuel Jefroykin2
1Department of Cognitive and Brain Sciences, Hebrew University of Jerusalem, Mount Scopus, Jerusalem, 9190500, Israel, 972 025882888.
JMIR Formative Research
|October 24, 2025
Summary
This study introduces a framework using large language models (LLMs) to filter noisy conversational transcripts, improving data quality for behavioral health research. The hybrid approach effectively distinguishes therapy sessions from non-sessions, enhancing data usability.
Area of Science:
- Computational Linguistics
- Health Informatics
- Artificial Intelligence
Background:
- Automatic transcription of conversations generates noisy datasets with errors and unintended recordings.
- Preprocessing and filtering are crucial for the research utility of large conversational transcript datasets.
- Accurate conversation representation is vital for deriving insights in behavioral health contexts.
Purpose of the Study:
- To present a framework for preprocessing and filtering large conversational transcript datasets.
- To remove non-session transcripts unrelated to behavioral treatment sessions.
- To enhance the utility of behavioral health transcripts for research.
Main Methods:
- Integrated feature extraction, human annotation, and large language models (LLMs).
- Utilized LLM perplexity to measure transcript noise and zero-shot prompting for classification.
- Prioritized data security and anonymity throughout the process.
Main Results:
- Approximately one-third of transcripts contained errors, including incomprehensible segments and speaker diarization issues.
- LLM perplexity showed higher scores in non-sessions, but moderate classification performance alone.
- Zero-shot LLM prompting achieved high agreement with expert ratings (κ=0.71) in distinguishing sessions from non-sessions.
Conclusions:
- The hybrid approach effectively characterizes errors and distinguishes text types in conversational datasets.
- Provides a foundation for ensuring data quality and usability in mental health research.
- Emphasizes integrating clinical experts with AI tools while prioritizing data security.
Keywords:
artificial intelligencebehavioral healthclinical documentationclinical textsconversational transcriptsdata preprocessingdata quality assessmenthealth informaticshealth information systemslarge language modelsnatural language processingpsychotherapytext classificationMore Related Videos
Related Concept Videos
Automatic Processing and Automatic Social Behavior
213
Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
213
Behavioral Genetics and Its Designs
1.0K
Behavior genetics explores how genetic inheritance influences human behavior. It focuses on how genes, passed from parents to offspring, contribute to the development of behavioral traits and tendencies. This branch of genetics seeks to understand the complex interplay between inherited genetic factors and environmental influences in shaping our behaviors.
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
1.0K

