Related Experiment Video
Updated: Jun 12, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning
Jessica N Lalonde1, Defne Circi2, Babetta L Marrone1
1Bioscience Division, Los Alamos National Laboratory, P.O. Box 1663, Los Alamos, New Mexico 87545, United States.
Abstract:
Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.
Related Concept Videos
Classification and Mechanical Properties of Synthetic Polymers
Bioplastics
Microbial Bioremediation of Plastics
Polymers: Molecular Weight Distribution