Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Band Gap Prediction of Two-Dimensional Materials Using a Gradient-Boosted Feature Selection Approach.

Journal of chemical information and modeling·2026
Same author

Machine-Learning Predictions of Photoluminescence in Molecules Exhibiting Thermally Activated Delayed Fluorescence with Implicit Experimental Validation.

Journal of chemical information and modeling·2026
Same author

Automatic Generation of a Mechanical Properties Question-Answering Data Set for Language Model Benchmarking: A Comparative Study of BERT, XLNet, and LLaMA Models.

Journal of chemical information and modeling·2026
Same author

A dataset of Curie and Néel temperatures auto-generated with ChemDataExtractor and the Snowball algorithm.

Scientific data·2025
Same author

Automated Determination of the Molecular Substructure from Nuclear Magnetic Resonance Spectra Using Neural Networks.

Journal of chemical information and modeling·2025
Same author

Autogenerating a Domain-Specific Question-Answering Data Set from a Thermoelectric Materials Database to Enable High-Performing BERT Models.

Journal of chemical information and modeling·2025

Related Experiment Video

Updated: May 28, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

478

Cost-Efficient Domain-Adaptive Pretraining of Language Models for Optoelectronics Applications.

Dingyun Huang1, Jacqueline M Cole1,2

  • 1Cavendish Laboratory, Department of Physics, University of Cambridge, J. J. Thomson Avenue, Cambridge CB3 0HE, U.K.

Journal of Chemical Information and Modeling
|February 11, 2025
PubMed
Summary

We developed three optoelectronics-specific BERT language models that outperform general models on optoelectronics NLP tasks. A cost-effective pretraining method significantly reduces resource needs.

More Related Videos

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.4K
Lensless Fluorescent Microscopy on a Chip
11:23

Lensless Fluorescent Microscopy on a Chip

Published on: August 17, 2011

17.6K

Related Experiment Videos

Last Updated: May 28, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

478
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.4K
Lensless Fluorescent Microscopy on a Chip
11:23

Lensless Fluorescent Microscopy on a Chip

Published on: August 17, 2011

17.6K

Area of Science:

  • Optoelectronics
  • Natural Language Processing (NLP)
  • Artificial Intelligence (AI)

Background:

  • Pretrained language models (PLMs) show versatility in NLP and applications like data mining in optoelectronics.
  • Bidirectional Encoder Representations from Transformers (BERT) is a widely adopted architecture in scientific domains.
  • General NLP models may not fully capture the nuances of optoelectronics research literature.

Purpose of the Study:

  • To introduce novel optoelectronics-aware BERT models (OE-BERT, OE-ALBERT, OE-RoBERTa).
  • To evaluate their performance against general English models and larger counterparts on optoelectronics-related NLP tasks.
  • To demonstrate an efficient domain-adaptive pretraining (DAPT) method for scientific NLP.

Main Methods:

  • Development of three specialized BERT architectures: OE-BERT, OE-ALBERT, and OE-RoBERTa.
  • Domain-adaptive pretraining (DAPT) applied to RoBERTa for optoelectronics text.
  • Comparative evaluation of model performance on various optoelectronics NLP tasks.

Main Results:

  • OE-BERT, OE-ALBERT, and OE-RoBERTa models surpassed general English BERT models and larger models in optoelectronics NLP tasks.
  • The DAPT method for RoBERTa achieved significant computational savings (>80%) in pretraining.
  • Performance was maintained or enhanced using the cost-effective DAPT approach.

Conclusions:

  • Optoelectronics-specific BERT models offer superior performance for NLP tasks in this domain.
  • Domain-adaptive pretraining is an effective and resource-efficient strategy for developing specialized scientific language models.
  • The developed models and datasets are publicly available to the optoelectronics research community.