Related Experiment Video
Updated: Feb 27, 2026

05:12
ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
12.0K
SOTXTSTREAM: Density-based self-organizing clustering of text streams
Avory C Bryant1,2, Krzysztof J Cios1,3
1Department of Computer Science, Virginia Commonwealth University, Richmond, VA, United States of America.
Plos One
|July 8, 2017
Summary
A novel algorithm, SOTXTSTREAM, enhances density-based stream clustering for text data. It overcomes limitations of prior methods, improving clustering performance on real-world datasets.
Area of Science:
- Data Mining
- Machine Learning
- Natural Language Processing
Background:
- Density-based clustering algorithms often struggle with heterogeneous densities.
- Existing stream clustering methods frequently employ a two-phase approach (online micro-clustering, offline macro-clustering).
- The SOSTREAM algorithm introduced self-organization in the online phase to achieve single-phase macro-clustering.
Purpose of the Study:
- To present a new density-based self-organizing algorithm for text stream clustering, SOTXTSTREAM.
- To address limitations of the SOSTREAM algorithm in text data clustering.
- To demonstrate improved clustering performance using SOTXTSTREAM on real-world text streams.
Main Methods:
- Building upon the SOSTREAM algorithm's density-based, self-organizing approach.
- Implementing local density determinations to handle varying cluster densities.
- Developing SOTXTSTREAM for single-phase macro-clustering of text streams.
Main Results:
- SOTXTSTREAM effectively handles heterogeneous densities in text data streams.
- The algorithm achieves improved clustering performance compared to SOSTREAM.
- Demonstrated success on multiple real-world text stream datasets.
Conclusions:
- SOTXTSTREAM offers an effective single-phase solution for density-based text stream clustering.
- The algorithm overcomes key limitations of previous stream clustering techniques.
- SOTXTSTREAM shows significant promise for analyzing dynamic text data.
Related Concept Videos
Cluster Sampling Method
15.1K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.1K
Sampling Plans
1.1K
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
1.1K
Vesicular Tubular Clusters
3.3K
After budding out from the ER membrane, some COPII vesicles lose their coat and fuse with one another to form larger vesicles and interconnected tubules called vesicular tubular clusters or VTCs. These clusters constitute a compartment at the ER-Golgi interface known as ERGIC (Endoplasmic Reticulum Golgi Intermediate Compartment). The ERGIC is a mobile membrane-bound cargo transport system that sorts proteins secreted from ER and delivers them to the Golgi.
With the help of motor proteins such...
With the help of motor proteins such...
3.3K
Classification of Signals
1.5K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.5K
Relative Frequency Histogram
6.6K
The relative frequency depicts the proportion of data points that have each value. The frequency tells the number of data points that have each value. Like the histogram, a relative frequency histogram also has the same shape with a horizontal scale (the x-axis), but the vertical scale (the y-axis) is marked with relative frequencies (percentages of the whole) instead of actual frequencies. A relative frequency histogram is a graphical representation of a frequency distribution where the...
6.6K
RNA-seq
12.2K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.2K

