Related Experiment Video
Updated: Oct 19, 2025

Analyzing and Building Nucleic Acid Structures with 3DNA
Published on: April 26, 2013
Accurate prediction of B-form/A-form DNA conformation propensity from primary sequence: A machine learning and free
Abhijit Gupta1, Mandar Kulkarni2, Arnab Mukherjee1
1Department of Chemistry, Indian Institute of Science Education and Research, Pune, Maharashtra 411008, India.
Abstract:
DNA carries the genetic code of life, with different conformations associated with different biological functions. Predicting the conformation of DNA from its primary sequence, although desirable, is a challenging problem owing to the polymorphic nature of DNA. We have deployed a host of machine learning algorithms, including the popular state-of-the-art LightGBM (a gradient boosting model), for building prediction models. We used the nested cross-validation strategy to address the issues of "overfitting" and selection bias. This simultaneously provides an unbiased estimate of the generalization performance of a machine learning algorithm and allows us to tune the hyperparameters optimally. Furthermore, we built a secondary model based on SHAP (SHapley Additive exPlanations) that offers crucial insight into model interpretability. Our detailed model-building strategy and robust statistical validation protocols tackle the formidable challenge of working on small datasets, which is often the case in biological and medical data.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Predicting Molecular Geometry
DNA as a Genetic Template
Protein Organization
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Protein Folding

