Perspectives on Knowledge Discovery Algorithms Recently Introduced in Chemoinformatics: Rough Set Theory, Association

Eleanor J Gardiner1, Valerie J Gillet1

  • 1Information School, University of Sheffield , Regent Court, 211 Portobello, Sheffield S1 4DP, United Kingdom.

Summary

This review explores four data mining techniques: Rough Set Theory, Association Rule Mining, Emerging Pattern Mining, and Formal Concept Analysis, highlighting their chemoinformatics applications. These methods offer descriptive power for knowledge discovery in large datasets.

Related Concept Videos

Drug Discovery: Overview01:26

Drug Discovery: Overview

Drug discovery is a multifaceted process involving extensive screening, testing, and optimization of lead compounds to identify potential new drugs for therapeutic use. It combines several approaches, including screening large numbers of natural products, chemical modification of known active molecules, identification of new drug targets, and rational design based on biological mechanisms and drug-receptor structure. These approaches are carried out in both academic research laboratories and...
13.3K
Inductive Effects on Chemical Shift: Overview01:27

Inductive Effects on Chemical Shift: Overview

The protons in unsubstituted alkanes are strongly shielded with chemical shifts below 1.8 ppm. Methine, methylene, and methyl protons appear at approximately 1.7, 1.2 and 0.7 ppm, while the proton signal from methane appears at 0.23 ppm. An electronegative substituent, such as chlorine, withdraws the electron density from the protons, increasing their chemical shift. Progressive substitution of the hydrogens in methane by chlorine shifts the proton signals increasingly downfield, to 3.05 ppm in...
2.6K
Ligand Binding Sites02:40

Ligand Binding Sites

Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
15.8K
Protein-protein Interfaces02:04

Protein-protein Interfaces

Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
15.0K
Chemical Shift: Internal References and Solvent Effects01:17

Chemical Shift: Internal References and Solvent Effects

In an NMR sample, precise measurement of the absolute absorption frequencies of nuclei is difficult. A standard internal reference compound is added, and the frequency difference between the reference signal and sample signals is measured.
The internal reference compound generally used in NMR spectroscopy is tetramethylsilane (TMS). TMS is preferred because it is chemically inert, soluble in NMR solvents, and easily removable. Also, the highly shielded methyl protons in TMS yield an intense...
1.6K
Mass Spectrometry of Amines01:15

Mass Spectrometry of Amines

In mass spectroscopy, amines undergo fragmentation to give parent ions with odd molecule weights. This observed mass spectrum follows the nitrogen rule; a molecule with an odd number of nitrogen atoms produces a molecular ion with an odd molecular weight. Amines undergo fragmentation through α cleavage, producing nitrogen-containing cations—iminium ions—and alkyl radicals. Mass spectra of aromatic and cyclic aliphatic amines exhibit strong molecular ion peaks, but acyclic...
5.6K