Related Experiment Video
Updated: Jan 21, 2026

Wet Chemistry and Peptide Immobilization on Polytetrafluoroethylene for Improved Cell-adhesion
Published on: August 15, 2016
Linking in silico MS/MS spectra with chemistry data to improve identification of unknowns
Andrew D McEachran1,2, Ilya Balabin3, Tommy Cathey4
1Oak Ridge Institute for Science and Education (ORISE) Research Participation Program, United States Environmental Protection Agency, 109 T.W. Alexander Dr., Research Triangle Park, Durham, NC, 27711, USA. admceachran@gmail.com.
Abstract:
Confident identification of unknown chemicals in high resolution mass spectrometry (HRMS) screening studies requires cohesive workflows and complementary data, tools, and software. Chemistry databases, screening libraries, and chemical metadata have become fixtures in identification workflows. To increase confidence in compound identifications, the use of structural fragmentation data collected via tandem mass spectrometry (MS/MS or MS2) is vital. However, the availability of empirically collected MS/MS data for identification of unknowns is limited. Researchers have therefore turned to in silico generation of MS/MS data for use in HRMS-based screening studies. This paper describes the generation en masse of predicted MS/MS spectra for the entirety of the US EPA's DSSTox database using competitive fragmentation modelling and a freely available open source tool, CFM-ID. The generated dataset comprises predicted MS/MS spectra for ~700,000 structures, and mappings between predicted spectra, structures, associated substances, and chemical metadata. Together, these resources facilitate improved compound identifications in HRMS screening studies. These data are accessible via an SQL database, a comma-separated export file (.csv), and EPA's CompTox Chemicals Dashboard.
Related Concept Videos
Emission Spectra
Chemistry of Carbohydrates
Chemistry of Carbohydrates
Covalently Linked Protein Regulators
These groups modify specific amino acids in a protein....
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used;...
Estimating Population Mean with Unknown Standard Deviation
William S. Gosset (1876–1937) of the...

