Related Experiment Video
Updated: Jun 5, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Artificial Allies: Validation of Synthetic Text for Peer Support Tools through Data Augmentation in NLP Model
Josué Godeme1, Julia Hill2, Stephen P Gaughan3
1Research Computing and Data Services, Information, Technology & Consulting, Dartmouth College, Hanover, NH 03784, USA, josue.f.godeme.26@dartmouth.edu.
Synthetic text can augment training data for Natural Language Processing (NLP) models in peer support. While participants struggled to distinguish AI from human text, findings support synthetic data for enhancing NLP in peer support interventions.
Area of Science:
- Natural Language Processing (NLP)
- Human-Computer Interaction
- Artificial Intelligence (AI)
Background:
- Augmenting training data is crucial for developing robust Natural Language Processing (NLP) models.
- Peer support tools can benefit from advanced NLP for improved user interaction and support fidelity.
- Evaluating the detectability of synthetic text is key to its effective use in AI model training.
Purpose of the Study:
- To investigate the efficacy of synthetic text in augmenting training data for NLP models within peer support contexts.
- To assess the ability of human raters, including professional peer supporters and AI-proficient individuals, to differentiate between AI-generated and human-written sentences.
- To analyze classification accuracy and confidence levels using signal detection theory and confidence-based metrics.
Main Methods:
- Surveyed 22 participants (13 professional peer supporters, 9 AI-proficient individuals) on distinguishing AI-generated from human-written text.
- Employed signal detection theory and confidence-based metrics to evaluate rater performance.
- Analyzed rater agreement, classification accuracy, and biases using statistical methods.
Main Results:
- No significant difference in rater agreement was found between peer supporters and AI-proficient individuals (p = 0.116).
- Overall classification accuracy was below chance levels (mean accuracy = 43.10%, p < 0.001), indicating difficulty in distinguishing text types.
- Peer supporters showed a significant bias towards misclassifying low-fidelity sentences as AI-generated (p = 0.007).
- AI-proficient raters demonstrated a significant negative correlation between errors and confidence (r = -0.429, p < 0.001).
Conclusions:
- Synthetic text shows feasibility for mimicking human communication, supporting its use for augmenting NLP training data.
- The findings have significant implications for enhancing the fidelity of peer support interventions through improved NLP model development.
- Further research is warranted to refine synthetic text generation and its application in specialized domains like peer support.
Related Concept Videos
Improving Translational Accuracy
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Non-equilibrium in the Cell
Transfer RNA Synthesis
Each of these chemical modifications is carried by a specific enzyme, post-transcription. All of these enzymes have unique base and site-specificity. Methylation, the most common chemical modification, is carried by at least nine different enzymes, with...
Proofreading
Errors During Replication are Corrected by the DNA Polymerase...

