SMASH: A Data-driven Informatics Method to Assist Experts in Characterizing Semantic Heterogeneity among Data

William Brown1, Chunhua Weng2, David K Vawdrey3

  • 1Department of Biomedical Informatics, Columbia University, New York, NY; HIV Center for Clinical and Behavioral Studies, NY State Psychiatric Institute & Columbia University, New York, NY.

Summary

Semantic heterogeneity (SH) in healthcare data hinders interoperability. A new method, SMASH, combined with expert review, effectively identified SH in HIV data, revealing most issues stem from differing terms for the same concept.

Related Concept Videos

Variability: Analysis01:11

Variability: Analysis

Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
596
Factors Affecting Dissolution: Polymorphism, Amorphism and Pseudopolymorphism01:21

Factors Affecting Dissolution: Polymorphism, Amorphism and Pseudopolymorphism

Polymorphism refers to the existence of a drug substance in multiple crystalline forms, known as polymorphs. Recently, this term has been expanded to include solvates (forms containing a solvent), amorphous forms (non-crystalline forms), and desolvated solvates (forms from which the solvent has been removed).
Some polymorphic crystals possess lower aqueous solubility than their amorphous counterparts, leading to incomplete absorption. For instance, the oral suspension of Chloramphenicol, which...
803
Test for Homogeneity01:23

Test for Homogeneity

The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.5K
Data: Types and Distribution01:19

Data: Types and Distribution

In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
2.1K
How Data are Classified: Categorical Data01:11

How Data are Classified: Categorical Data

A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
47.0K
Law of Independent Assortment02:03

Law of Independent Assortment

While Mendel’s Law of Segregation states that the two alleles for one gene are separated into different gametes, a different question of how different genes are inherited remains. For example, is the gene for tall plants inherited with the gene for green peas? Mendel asked this question by experimenting with a dihybrid cross; a cross in which both parents are homozygous for two distinct traits resulting in an F1 generation that are heterozygous for both traits.
63.8K