Video Experimental Relacionado
Updated: Feb 24, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.2K
Fenotipado Contextual de Cohorte de Sepsis Pediátrica Utilizando Modelos de Lenguaje Grandes
Aditya Nagori1,2, Ayush Gautam1,3, Matthew O Wiens4,5,6
1Department of Surgery, Duke university school of medicine, Durham NC, USA.
AMIA ... Annual Symposium proceedings. AMIA Symposium
|February 23, 2026
Resumen
Los modelos de lenguaje grandes (LLM) ofrecen agrupamiento avanzado de subgrupos de pacientes para atención personalizada, superando los métodos tradicionales en datos complejos de sepsis pediátrica de países de bajos ingresos. El agrupamiento de LLM revela perfiles de pacientes distintos para una mejor toma de decisiones.
Área de la Ciencia:
- Computational biology
- Medical informatics
- Artificial intelligence in healthcare
Sus antecedentes:
- Personalized care and resource optimization require effective patient subgroup clustering.
- Traditional methods face challenges with high-dimensional, heterogeneous healthcare data and lack contextual understanding.
- Pediatric sepsis datasets from low-income countries present unique challenges for analysis.
Objetivo del estudio:
- To evaluate Large Language Model (LLM)-based clustering against classical methods for pediatric sepsis patient subgroups.
- To assess the performance of different LLM embedding models and clustering objectives.
- To determine the potential of LLM clustering for contextual phenotyping in resource-limited settings.
Principales métodos:
- Patient records were serialized into text and clustered using LLM embeddings (LLAMA 3.1 8B, DeepSeek-R1-Distill-Llama-8B, Stella-En-400M-V5) with K-means.
- Classical methods (K-Medoids) were applied to dimensionality-reduced mixed data (UMAP, FAMD).
- Clustering quality was assessed using Silhouette scores and statistical tests on a pediatric sepsis dataset (2,686 records).
Principales resultados:
- Stella-En-400M-V5 achieved the highest Silhouette Score (0.86).
- LLAMA 3.1 8B with a clustering objective identified distinct patient subgroups based on nutritional, clinical, and socioeconomic profiles.
- LLM-based methods demonstrated superior performance over classical techniques by capturing richer context.
Conclusiones:
- LLM-based clustering effectively captures contextual information and outperforms classical methods for heterogeneous healthcare data.
- LLM clustering holds significant potential for contextual phenotyping and improved decision-making in resource-limited settings.
- This approach can enhance personalized care strategies by identifying nuanced patient subgroups.

