Related Experiment Video
Updated: Jul 13, 2026

CorrelationCalculator and Filigree: Tools for Data-Driven Network Analysis of Metabolomics Data
Published on: November 10, 2023
Non-linear correlation of content and metadata information extracted from biomedical article datasets
Theodosios Theodosiou1, Lefteris Angelis, Athena Vakali
1Department of Informatics, School of Natural Sciences, Aristotle University of Thessaloniki, 54124 Thessaloniki, Greece. theodos@csd.auth.gr
Non-Linear Canonical Correlation Analysis (NLCCA) effectively combines information from diverse biomedical document sources into a single dataset. This machine learning approach enhances knowledge extraction and organization from large scientific literature databases.
Area of Science:
- Biomedical Informatics
- Machine Learning
- Data Science
Background:
- Biomedical literature databases are crucial for accessing current scientific knowledge.
- Organizing these databases and extracting novel insights necessitates efficient machine learning methods.
- Many machine learning methods require documents to be represented as large multivariate datasets.
Purpose of the Study:
- To investigate Non-Linear Canonical Correlation Analysis (NLCCA) for integrating information from different document representations.
- To address the challenge of combining voluminous information from various sources into a concise, representative dataset.
- To describe documents using a single, consolidated dataset.
Main Methods:
- Utilized Non-Linear Canonical Correlation Analysis (NLCCA), a multivariate statistical approach.
- Exploited correlations among variables from different document representations (text words, MeSH, GO terms).
- Experimented with document datasets represented by text words, Medical Subject Headings (MeSH), and Gene Ontology (GO) terms.
Main Results:
- Demonstrated the effectiveness of NLCCA in combining information from diverse biomedical data sources.
- Showcased NLCCA's ability to create a concise dataset that efficiently represents a corpus of documents.
- Validated the approach using document datasets represented by text words, MeSH, and GO terms.
Conclusions:
- NLCCA is an effective method for integrating heterogeneous biomedical data sources.
- The approach facilitates the creation of concise, informative datasets for improved knowledge discovery.
- NLCCA shows promise for enhancing the organization and extraction of knowledge from biomedical literature.
More Related Videos
Related Concept Videos
Correlation
Two variables, for example, a and b, are said to be positively correlated if both variables move in the same direction. In other words, a positive correlation exists between two variables, a and b, if:
Correlations
Correlation and Regression
Calculating and Interpreting the Linear Correlation Coefficient
Drug Concentration Versus Time Correlation
Two pivotal parameters are the minimum effective concentration (MEC) and the minimum toxic concentration (MTC). The MEC is the lowest drug...
MALDI-TOF Mass Spectrometry

