Related Experiment Video
Updated: Sep 12, 2025

Asymmetric Thermoelectrochemical Cell for Harvesting Low-grade Heat under Isothermal Operation
Published on: February 5, 2020
Autogenerating a Domain-Specific Question-Answering Data Set from a Thermoelectric Materials Database to Enable
Odysseas Sierepeklis1, Jacqueline M Cole1,2
1Cavendish Laboratory, University of Cambridge, J. J. Thomson Avenue, Cambridge CB3 0HE, U.K.
We developed a method to automatically create a large question-answering (QA) dataset for thermoelectric materials. Fine-tuning a BERT model on this domain-specific data significantly improves its performance in materials science applications.
Area of Science:
- Materials Science
- Computational Linguistics
- Artificial Intelligence
Background:
- Domain-specific datasets are crucial for training high-performing language models.
- Existing generic QA datasets may not capture the nuances of specialized scientific fields like thermoelectric materials.
- Small language models (SLMs) offer computational efficiency but often require tailored training data.
Purpose of the Study:
- To present a method for autogenerating a large, domain-specific question-answering (QA) dataset for thermoelectric materials.
- To evaluate the performance of a fine-tuned BERT model on this dataset compared to generic datasets.
- To investigate the impact of mixing domain-specific and generic QA data on model performance.
Main Methods:
- Autogeneration of a 99,757 QA pair dataset from a thermoelectric materials database.
- Fine-tuning a BERT language model on the autogenerated dataset, the generic SQuAD-v2 dataset, and a mixed dataset.
- Evaluation of model performance using exact match and F1 scores on a dedicated test set.
Main Results:
- The BERT model fine-tuned on the autogenerated domain-specific dataset outperformed the model trained on SQuAD-v2.
- Mixing the domain-specific and generic datasets resulted in the best performance.
- The best model achieved an exact match score of 67.93% and an F1 score of 72.29% on the test data.
Conclusions:
- Autogenerated, domain-specific QA datasets can significantly enhance the performance of small language models in specialized fields.
- Combining domain-specific data with generic datasets offers synergistic benefits for language model training.
- This method enables the development of high-performing SLMs with modest computational resources for materials science applications.
Related Concept Videos
Thermodynamic Potentials
Thermodynamic Systems
Consider an example of tea boiling in a kettle. The...
First Law Of Thermodynamics: Problem-Solving
The following strategies can be used to solve any problem involving the first law of thermodynamics.
Heat Capacity: Problem-Solving
Determine the type of gas: The heat capacity of a gas depends on its molecular structure and the degree of freedom of its molecules. Different types of...
Maxwell-Boltzmann Distribution: Problem Solving
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
Thermosensation

