Related Experiment Video
Updated: Jan 21, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Accelerating dataset generation for machine learning using large language models: a pharmaceutical additive
Paola Carou-Senra1, Lucía Rodríguez-Pombo1, Carmen Alvarez-Lorenzo1
1Departamento de Farmacología, Farmacia y Tecnología Farmacéutica, I+D Farma (GI-1645), Facultad de Farmacia, Instituto de Materiales (iMATUS) and Health Research Institute of Santiago de Compostela (IDIS), Universidade de Santiago de Compostela 15782 Santiago de Compostela, Spain.
None:
Creating high-quality datasets for training machine learning models in specialized domains like pharmaceutical research is often constrained by the manual effort required to extract and compute critical parameters from heterogeneous literature. A novel deep prompt-engineering framework was developed to transform GPT-4 into a robust tool for automated and accelerated generation of structured datasets. Using a multi-set prompt strategy, GPT-4 analysed 70 full-text articles from literature on pharmaceutical inkjet printing to extract and compute 23 domain-relevant variables. These variables were organized into three main parameter groups: (i) printing parameters, (ii) rheological properties, and (iii) drug dose parameters, which were analysed using dedicated prompts. The outputs were benchmarked against a human-curated dataset compiled over four months by four domain experts, previously used to train machine learning models for predicting inkjet printability. Iterative prompt engineering yielded an overall accuracy of 0.942 across 4,217 individual variable-level data points, with computed variables reaching 0.983 accuracy despite multi-step calculations and unit conversions. Inter-day reproducibility was stable, and sensitivity, specificity, and predictive values all exceeded 0.900. Remarkably, the workflow reduced processing time from hours of human effort to <3.5 min per article. This novel prompt-engineering approach enabled GPT-4 to generate reliable, high-quality, literature-derived datasets, dramatically reducing manual effort while maintaining expert-level accuracy. Ultimately, this novel strategy facilitates the scalability of machine learning in pharmaceutical, and other data-intensive domains.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
04:04Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Simplified Synchronous Machine Model
In this model, each generator is connected to a...
Wind Turbine Machine Models
Induction machines interact through the rotating magnetic field generated by the stator and the rotor. The key parameter is slip, which is the difference between synchronous speed and rotor speed relative to synchronous speed. Slip is...
Conjugate Addition (1,4-Addition) vs Direct Addition (1,2-Addition)
Conjugate addition results in a thermodynamically stable product. The reaction retains the stronger C=O bond at the expense of the weaker C=C π bond. The process is slow as the β carbon is less electrophilic than the carbonyl carbon.
Direct addition products are...
Pharmaceutical Equivalents
Components of Language