Related Experiment Video
Updated: Sep 12, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
682
Improving MedDRA/J Coding Accuracy with a Fine-Tuned Text Embedding Model.
Shoya Wada1,2, Masaharu Okamoto2, Kento Sugimoto2
1Department of Transformative System for Medical Information, Graduate School of Medicine, The University of Osaka.
Studies in Health Technology and Informatics
|August 8, 2025
Summary
Fine-tuning text embedding models on adverse event (AE) data improves MedDRA coding accuracy in Japan. This approach enhances search precision and recall for better regulatory compliance.
Area of Science:
- Pharmacovigilance
- Natural Language Processing
- Medical Coding
Background:
- MedDRA (Medical Dictionary for Regulatory Activities) coding in Japan presents challenges due to discrepancies between source expressions and dictionary terms.
- Accurate adverse event (AE) reporting is crucial for patient safety and regulatory compliance.
Purpose of the Study:
- To enhance the accuracy and efficiency of MedDRA coding in Japan.
- To develop and evaluate a fine-tuned text embedding model using in-house AE data.
Main Methods:
- A text embedding model was fine-tuned on a dataset of 50,000 in-house adverse event entries.
- The performance of the fine-tuned model was compared against baseline methods for MedDRA term ranking and recall.
Main Results:
- The fine-tuned model demonstrated significant improvements in MedDRA term ranking and recall.
- Achieved nDCG@20 of 76.2% and Recall@20 of 90.8% with 50K in-house entries.
- Outperformed baseline approaches in accurately mapping source expressions to MedDRA terms.
Conclusions:
- Fine-tuning text embedding models on real-world AE data is a viable strategy to improve MedDRA/J coding.
- This advanced approach can significantly benefit pharmacovigilance activities by increasing search accuracy and efficiency.
- The methodology offers a scalable solution for addressing coding discrepancies in regulatory settings.
Related Concept Videos
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Margin of Error
4.5K
The margin of error is also called the maximum error of an estimate. The margin of error is the maximum possible or expected difference between the observed sample parameter value and the actual population parameter value. For proportion, it is the maximum difference between the value of sample proportion obtained from the data and the true value of population proportion. As the true value of the population parameter is not known, the margin of error is calculated using the sample statistic.
4.5K
Extraction: Advanced Methods
533
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
533
RNA Editing
9.2K
RNA editing is a post-transcriptional modification where a precursor mRNA (pre-mRNA) nucleotide sequence is changed by base insertion, deletion, or modification. The extent of RNA editing varies from a few hundred bases, in mitochondrial DNA of trypanosomes, to a just single base, in nuclear genes of mammals. Even a single base change in the pre-mRNA can convert a codon for one amino acid into the codon for another amino acid or a stop codon. This type of re-coding can significantly affect the...
9.2K
Mismatch Repair
40.6K
Overview
40.6K
Upsampling
311
Managing signal sampling rates is essential in digital signal processing to maintain signal integrity. A decimated signal, characterized by a reduced frequency range due to its lower sampling rate, can be upsampled by inserting zeros between each sample. This upsampling process expands the original spectrum and introduces repeated spectral replicas at intervals dictated by the new Nyquist frequency. To refine this zero-inserted sequence, it is passed through a lowpass filter with a cutoff...
311

