Related Experiment Video
Updated: May 24, 2025

Systematic Approach to Identify Novel Antimicrobial and Antibiofilm Molecules from Plants' Extracts and Fractions to Prevent Dental Caries
Published on: March 31, 2021
A machine learning driven automated system to extract multiple information fields from safety data sheet documents
Misbah Khan1, Julia Penfield1, Aatish Suman1
1VelocityEHS, Chicago, USA.
Abstract:
Safety Data Sheets (SDS) provide essential safety and health information for various substances and products. They are widely used in industries that require cataloguing information on chemicals such as green chemistry, industrial hygiene, and regulatory compliance, among others within the Environment, Health, and Safety (EHS) and the Environment, Social, and Governance (ESG) sectors. Over time, chemical data management has evolved from storing physical copies of SDS on work sites, to tabulating key fields from the SDS into a database, which is essential to successful inventory and chemical risk management. This process of extracting and structuring of essential information from SDS is known as SDS "indexing". SDS indexing is a critical task for inventory management and regulatory compliance, and is commonly done manually. Manual SDS indexing can be resource-intensive, as it requires personnel to process each document individually, often resulting in significant costs and extended processing times. Not all the fields from an SDS are always required to be indexed and the need is variant across companies and applications. However, there are 5 fields that are commonly needed across various applications and the process to extract them is referred to as "standard indexing". In this paper, we propose an automated system for standard indexing of SDS documents using a multi-step method with a combination of machine learning models and expert systems being executed sequentially. The system specifically extracts the fields Product Name, Product Code, Manufacturer Name, Supplier Name, and Revision Date. Our design achieves a precision range of 0.96-0.99 across the five fields, when evaluated on 150,000 SDS documents annotated for this purpose.
More Related Videos
09:04Identifying Per- and Polyfluorinated Chemical Species with a Combined Targeted and Non-Targeted-Screening High-Resolution Mass Spectrometry Workflow
Published on: April 18, 2019
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023