Related Experiment Video
Updated: Jun 24, 2025

DNA Virus Detection System Based on RPA-CRISPR/Cas12a-SPM and Deep Learning
Published on: May 10, 2024
Identification of Family-Specific Features in Cas9 and Cas12 Proteins: A Machine Learning Approach Using Complete
Sita Sirisha Madugula1, Pranav Pujar2, Bharani Nammi2
1Department of Pharmaceutical Sciences, University of North Texas System College of Pharmacy, University of North Texas Health Science Center, 3500 Camp Bowie Blvd, Fort Worth, Texas 76107, United States.
Abstract:
The recent development of CRISPR-Cas technology holds promise to correct gene-level defects for genetic diseases. The key element of the CRISPR-Cas system is the Cas protein, a nuclease that can edit the gene of interest assisted by guide RNA. However, these Cas proteins suffer from inherent limitations such as large size, low cleavage efficiency, and off-target effects, hindering their widespread application as a gene editing tool. Therefore, there is a need to identify novel Cas proteins with improved editing properties, for which it is necessary to understand the underlying features governing the Cas families. In this study, we aim to elucidate the unique protein features associated with Cas9 and Cas12 families and identify the features distinguishing each family from non-Cas proteins. Here, we built Random Forest (RF) binary classifiers to distinguish Cas12 and Cas9 proteins from non-Cas proteins, respectively, using the complete protein feature spectrum (13,494 features) encoding various physiochemical, topological, constitutional, and coevolutionary information on Cas proteins. Furthermore, we built multiclass RF classifiers differentiating Cas9, Cas12, and non-Cas proteins. All the models were evaluated rigorously on the test and independent data sets. The Cas12 and Cas9 binary models achieved a high overall accuracy of 92% and 95% on their respective independent data sets, while the multiclass classifier achieved an F1 score of close to 0.98. We observed that Quasi-Sequence-Order (QSO) descriptors like Schneider.lag and Composition descriptors like charge, volume, and polarizability are predominant in the Cas12 family. Conversely Amino Acid Composition descriptors, especially Tripeptide Composition (TPC), predominate the Cas9 family. Four of the top 10 descriptors identified in Cas9 classification are tripeptides PWN, PYY, HHA, and DHI, which are seen to be conserved across all Cas9 proteins and located within different catalytically important domains of the Streptococcus pyogenes Cas9 (SpCas9) structure. Among these, DHI and HHA are well-known to be involved in the DNA cleavage activity of the SpCas9 protein. Mutation studies have highlighted the significance of the PWN tripeptide in PAM recognition and DNA cleavage activity of SpCas9, while Y450 from the PYY tripeptide plays a crucial role in reducing off-target effects and improving the specificity in SpCas9. Leveraging our machine learning (ML) pipeline, we identified numerous Cas9 and Cas12 family-specific features. These features offer valuable insights for future experimental and computational studies aiming at designing Cas systems with enhanced gene-editing properties. These features suggest plausible structural modifications that can effectively guide the development of Cas proteins with improved editing capabilities.
Insights
This study identifies unique protein features distinguishing Cas9 and Cas12 gene editing proteins from non-Cas proteins using machine learning. These identified features can guide the development of improved CRISPR-Cas systems for genetic diseases.
Area of Science:
- Molecular Biology
- Bioinformatics
- Genetics
Background:
- CRISPR-Cas technology offers gene editing for genetic diseases, but Cas proteins have limitations like size and off-target effects.
- Understanding the features of Cas protein families is crucial for developing improved gene editing tools.
Purpose of the Study:
- To elucidate unique protein features of Cas9 and Cas12 families.
- To identify distinguishing features between Cas9, Cas12, and non-Cas proteins using machine learning.
Main Methods:
- Random Forest (RF) binary and multiclass classifiers were built using 13,494 protein features.
- Models were trained to distinguish Cas12 and Cas9 from non-Cas proteins, and to differentiate between Cas9, Cas12, and non-Cas proteins.
- Rigorous evaluation was performed on test and independent datasets.
Main Results:
- Binary models achieved high accuracy (92% for Cas12, 95% for Cas9) on independent datasets.
- The multiclass classifier achieved an F1 score of approximately 0.98.
- Key distinguishing features identified include Quasi-Sequence-Order descriptors for Cas12 and Amino Acid Composition/Tripeptide Composition for Cas9.
Conclusions:
- Specific tripeptides (PWN, PYY, HHA, DHI) in Cas9 are linked to DNA cleavage and specificity.
- Identified Cas9 and Cas12 family-specific features provide insights for designing enhanced gene-editing systems.
- These findings can guide structural modifications for improved Cas protein editing capabilities.
Related Concept Videos
CRISPR
Protein Families
Caspases
CRISPR and crRNAs
The CRISPR-Cas system stores a copy of foreign DNA in the host genome and uses it to identify the foreign DNA upon reinfection. CRISPR-Cas has three different...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...

