Related Experiment Video
Updated: Sep 20, 2025

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Talk to your data: Introducing text embedding similarity analysis (TESA) in psychological research
Juul Vossen1, Evy Kuijpers2, Joeri Hofmans1
1Department of Work and Organizational Psychology, Vrije Universiteit Brussel, Pleinlaan 2, 1050, Brussels, Belgium.
This study introduces Text Embedding Similarity Analysis (TESA), a novel approach using retrieval augmented generation (RAG) and large language models (LLM) for hypothesis testing in qualitative research.
Area of Science:
- Computational Linguistics
- Qualitative Data Analysis
- Natural Language Processing
Background:
- Qualitative research is valuable but struggles with hypothesis testing due to limitations in statistical modeling of text data.
- Traditional text quantification methods are labor-intensive, susceptible to bias, and overlook semantic nuances.
- Existing novel approaches often demand extensive data and are primarily inductive.
Purpose of the Study:
- To introduce a novel retrieval augmented generation (RAG)-based approach, Text Embedding Similarity Analysis (TESA), for hypothesis-driven analysis of text data.
- To enable researchers to formulate both hypothesis-based and open-ended questions for qualitative data.
- To overcome the limitations of traditional methods in hypothesis testing with text data.
Main Methods:
- TESA transforms hypotheses into population/sample and variable search terms.
- Pretrained large language models (LLM) extract semantic embeddings for search terms and text data.
- Cosine similarity is employed to match embeddings, enabling hypothesis testing via similarity score distribution analysis.
Main Results:
- The proposed TESA method facilitates hypothesis testing by analyzing the alignment of similarity scores.
- It allows for quantitative assessment of relationships between variables within specified populations in text data.
- Demonstrates a novel application of RAG and LLM for structured qualitative data analysis.
Conclusions:
- TESA offers a robust, scalable, and less biased method for hypothesis testing in qualitative research.
- This approach integrates semantic understanding with statistical rigor for text analysis.
- TESA empowers researchers to derive quantitative insights from qualitative text data efficiently.
More Related Videos
08:17A Semantic Priming Event-related Potential ERP Task to Study Lexico-semantic and Visuo-semantic Processing in Autism Spectrum Disorder
Published on: April 12, 2018
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Empathy
Causes of Similarity-Dissimilarity Effect
Stereotype Content Model
Factors Influencing Attraction III: Similarity
Stereotypes, Prejudice, and Discrimination
Automatic Processing and Automatic Social Behavior