Related Experiment Video
Updated: Jan 24, 2026

Supervised Machine Learning for Semi-Quantification of Extracellular DNA in Glomerulonephritis
Published on: June 18, 2020
Robust identification of molecular phenotypes using semi-supervised learning
Heinrich Roder1, Carlos Oliveira1, Lelia Net1
1Biodesix Inc, 2970 Wilderness Pl, Ste100, Boulder, CO, 80301, USA.
Background:
Modern molecular profiling techniques are yielding vast amounts of data from patient samples that could be utilized with machine learning methods to provide important biological insights and improvements in patient outcomes. Unsupervised methods have been successfully used to identify molecularly-defined disease subtypes. However, these approaches do not take advantage of potential additional clinical outcome information. Supervised methods can be implemented when training classes are apparent (e.g., responders or non-responders to treatment). However, training classes can be difficult to define when assessing relative benefit of one therapy over another using gold standard clinical endpoints, since it is often not clear how much benefit each individual patient receives.
Results:
We introduce an iterative approach to binary classification tasks based on the simultaneous refinement of training class labels and classifiers towards self-consistency. As training labels are refined during the process, the method is well suited to cases where training class definitions are not obvious or noisy. Clinical data, including time-to-event endpoints, can be incorporated into the approach to enable the iterative refinement to identify molecular phenotypes associated with a particular clinical variable. Using synthetic data, we show how this approach can be used to increase the accuracy of identification of outcome-related phenotypes and their associated molecular attributes. Further, we demonstrate that the advantages of the method persist in real world genomic datasets, allowing the reliable identification of molecular phenotypes and estimation of their association with outcome that generalizes to validation datasets. We show that at convergence of the iterative refinement, there is a consistent incorporation of the molecular data into the classifier yielding the molecular phenotype and that this allows a robust identification of associated attributes and the underlying biological processes.
Conclusions:
The consistent incorporation of the structure of the molecular data into the classifier helps to minimize overfitting and facilitates not only good generalization of classification and molecular phenotypes, but also reliable identification of biologically relevant features and elucidation of underlying biological processes.
Related Concept Videos
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Molecular Models
Molecular and Ionic Solids
Molecular Solids
Molecular crystalline solids, such as ice, sucrose (table sugar), and iodine, are solids that are composed of neutral molecules as their constituent units. These molecules are held together by weak intermolecular forces such as London dispersion forces, dipole-dipole interactions, or hydrogen bonds, which...
Molecular Orbital Theory II
Molecular Orbital Theory I
Predicting Molecular Geometry

