Related Experiment Video
Updated: Jun 11, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Semantic Space models for classification of consumer webpages on metadata attributes
Guocai Chen1, Jim Warren, Patricia Riddle
1Department of Computer Science, The University of Auckland, New Zealand. joe.g.chen@hotmail.com
This study introduces a method using language patterns to automatically classify online health information, improving quality and relevance for users. Machine learning accurately categorizes webpages by medical perspective, disease stage, and author credentials.
Area of Science:
- Computational linguistics
- Health informatics
- Machine learning
Background:
- Online health information presents challenges in quantity and quality.
- Web portals can curate reliable resources for specific health topics or user communities.
Purpose of the Study:
- To develop and evaluate automated methods for classifying online health webpages.
- To assess the accuracy of machine learning models in categorizing content based on language patterns.
Main Methods:
- Utilized Hyperspace Analogue to Language (HAL) to create Semantic Space models of webpage language.
- Applied machine learning algorithms: Support Vector Machine (SVM), Decision Forest, and Summed Similarity Measure (SSM).
- Classified webpages based on metadata attributes within the Breast Cancer Knowledge Online portal.
Main Results:
- Achieved over 93% accuracy in classifying 'medical' vs. 'supportive' perspectives.
- Reached over 92% accuracy for 'early' vs. 'advanced' disease stages.
- Attained over 90% accuracy in distinguishing 'lay' vs. 'clinician' author credentials.
Conclusions:
- Language use patterns in webpages can be effectively modeled using Semantic Spaces.
- Automated classification of online health content is feasible with high accuracy.
- This approach can enhance the organization and accessibility of reliable online health information.
Related Concept Videos
Stereotype Content Model
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Classification of Systems-II
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Self-Schemas
Schemas