Related Experiment Videos
Cross-entropy embedding of high-dimensional data using the neural gas model
Pablo A Estévez1, Cristián J Figueroa, Kazumi Saito
1Department of Electrical Engineering, University of Chile, Casilla 412-3, Santiago, Chile. pestevez@ing.uchile.cl
Summary
This study introduces a novel cross-entropy method for mapping high-dimensional data to low-dimensional embeddings. The approach enhances visualization and topology preservation compared to existing techniques like Neural Gas and Sammon
Area of Science:
- Machine Learning
- Data Visualization
- Dimensionality Reduction
Background:
- High-dimensional data presents challenges in visualization and analysis.
- Existing dimensionality reduction techniques like Sammon's nonlinear mapping (NLM) and vector quantization (e.g., Neural Gas, Self-Organizing Feature Map) have limitations in preserving topological relationships.
Purpose of the Study:
- To develop a novel cross-entropy approach for mapping high-dimensional data into a low-dimensional space.
- To simultaneously project input data and codebook vectors (from Neural Gas) into a low-dimensional output space.
- To preserve the neighborhood relationships defined by the Neural Gas algorithm during the mapping process.
Main Methods:
- A cross-entropy cost function is defined between input and output probabilities.
- The cost function is minimized using the Newton-Raphson method.
- The proposed method projects both input data and Neural Gas codebook vectors into the low-dimensional space.
Main Results:
- The new approach enables clear visualization of both data points and codebooks.
- It achieves superior mapping quality, specifically demonstrating improved topology preservation quantified by the q(m) measure.
- Performance is compared against Sammon's NLM and hierarchical approaches combining vector quantization with NLM.
Conclusions:
- The proposed cross-entropy method offers an effective strategy for dimensionality reduction.
- It provides enhanced visualization and superior topological accuracy compared to established methods.
- This technique is valuable for analyzing and understanding complex high-dimensional datasets.