Related Experiment Video
Updated: Aug 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Instability of LLM text embeddings for unsupervised dimension reduction of tabular data
Jun Li1, Yixuan Gou2, Shawn Su3
1Department of Applied and Computational Mathematics and Statistics, University of Notre Dame, Notre Dame, IN, United States.
Large language model (LLM) text embeddings are not reliable for unsupervised dimension reduction on tabular data. LLM-based methods show less stability compared to direct tabular approaches when dealing with biological and clinical datasets.
Area of Science:
- Bioinformatics
- Machine Learning
- Data Science
Background:
- Large language models (LLMs) offer a method for supervised learning on tabular data by converting observations into text embeddings.
- This approach is attractive for biomedical data due to its ability to handle mixed-type and missing values, creating complete numeric representations.
- However, the effectiveness of LLM embeddings for unsupervised tasks, such as dimension reduction, is not well understood.
Purpose of the Study:
- To evaluate the utility of LLM-derived text embeddings for dimension reduction of tabular data.
- Focus on biological and clinical datasets to assess performance in relevant domains.
- Compare the stability of LLM embedding-based methods against direct tabular approaches.
Main Methods:
- LLM text embeddings were generated by serializing tabular data observations into text.
- A direct tabular approach was used for comparison, calculating dissimilarities directly from original variables.
- Performance was assessed using stability under perturbation, as unsupervised dimension reduction lacks a ground truth.
Main Results:
- The LLM embedding-based approach demonstrated consistently lower stability than the direct tabular approach across various datasets and settings.
- Minor increases in missing data and random feature permutation significantly impacted the low-dimensional representations derived from LLM embeddings.
- These findings highlight the sensitivity of LLM embedding methods to data perturbations.
Conclusions:
- The direct application of LLM text embeddings is not a reliable strategy for unsupervised dimension reduction of tabular data.
- LLM-based methods are less robust to data variations compared to traditional direct tabular techniques.
- Further research may be needed to improve the stability and reliability of LLM embeddings for unsupervised tabular data analysis.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...
Dimensional Analysis
Dimensional analysis allows us to analyze and compare physical quantities on a...
Dimensional Analysis
Dimensional Analysis
In fluid mechanics, dimensional...
Dimensional Analysis
Conversion Factors and Dimensional Analysis
The unit...
Problem Solving: Dimensional Analysis