Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

SVM classification of human intergenic and gene sequences.

Y H Qiao1, J L Liu, C G Zhang

  • 1Biomechanics and Medical Information Institute, Beijing University of Technology, Beijing 100022, China.

Mathematical Biosciences
|May 17, 2005
PubMed
Summary

This study uses language analysis and Support Vector Machines (SVM) to improve gene prediction. The new method identifies keywords in DNA sequences, achieving 93% accuracy in predicting gene regions with fewer false positives.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Transient Large-Scale Anisotropy in TeV Cosmic Rays due to an Interplanetary Coronal Mass Ejection.

Physical review letters·2026
Same author

[Values of ATRX protein expression and ALT activation in differential diagnosis of uterine leiomyosarcoma].

Zhonghua bing li xue za zhi = Chinese journal of pathology·2026
Same author

First Detection of Ultrahigh Energy Emission from Gamma-Ray Binary LS I +61° 303.

Physical review letters·2026
Same author

Evidence of Cosmic-Ray Acceleration up to Sub-PeV Energies in the Supernova Remnant IC 443.

Physical review letters·2026
Same author

[Clinical characteristics of pediatric <i>Mycoplasma pneumoniae</i> necrotizing pneumonia with co-infections].

Zhonghua er ke za zhi = Chinese journal of pediatrics·2026
Same author

Precise Measurement of the Cosmic Ray Helium Spectrum above 0.1 PeV.

Physical review letters·2026

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Genomics

Background:

  • Gene-finding programs struggle with accurate automatic gene discovery.
  • Existing methods require significant manual refinement.
  • Limitations in current prediction accuracy hinder genomic research.

Purpose of the Study:

  • To develop a novel approach for accurate gene prediction using language analysis.
  • To enhance the identification of gene and intergenic sequences.
  • To improve the efficiency and correctness of automatic gene discovery.

Main Methods:

  • Gene and intergenic sequences analyzed as distinct linguistic subjects using a four-letter alphabet (A,C,G,T).
  • High-frequency simple sequences identified as keywords using a relative repeat ratio measurement (alpha(l(tau))).

Related Experiment Videos

  • DNA sequences mapped to a 178-dimensional Euclidean space; Support Vector Machines (SVM) employed for gene region prediction.
  • Main Results:

    • 178 short sequences selected as keywords after noise elimination.
    • Cross-validation demonstrated 93% prediction accuracy for gene sequences with 7% false positives.
    • Tested on a long genomic sequence, the method improved nucleotide-level specificity by 21% and correctly identified over 60% of predicted genes.

    Conclusions:

    • The developed gene-finding program significantly improves prediction accuracy and reduces false positives.
    • This language analysis-based approach offers a more robust method for automatic gene discovery.
    • The findings have implications for advancing genomic research and understanding gene organization.