Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

2.5K
2.5K
Aggregates Classification01:29

Aggregates Classification

292
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
292
Stereotype Content Model02:16

Stereotype Content Model

13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Data Collection by Observations01:08

Data Collection by Observations

11.7K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
11.7K
Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

456
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
456
Classification of Systems-II01:31

Classification of Systems-II

129
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
129

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A scaling limit theorem for controlled branching processes with a size-divisible term.

Revista de la Real Academia de Ciencias Exactas, Fisicas y Naturales. Serie A, Matematicas·2026
Same author

Adoption and Practice Variability of Prostate Stereotactic Body Radiotherapy (SBRT) in Latin America: The Role of Continuous Education in Advancing Quality and Standardization.

Cureus·2026
Same author

Corrigendum to Protein tyrosine phosphatase nonreceptor type 2 controls colorectal cancer development.

The Journal of clinical investigation·2026
Same author

Association of Frailty and Perioperative Outcomes Across Surgical Specialties Between 2016 and 2018 Using the National Inpatient Sample Database: A Retrospective Cohort Study.

Health science reports·2026
Same author

Updating the German Psycholinguistic Word Toolbox with AI-Generated Estimates of Concreteness, Valence, Arousal, Age of Acquisition, and Familiarity.

Journal of cognition·2026
Same author

Association of Frailty and Postoperative Outcomes across Race and Ethnicity.

Journal of racial and ethnic health disparities·2025

Related Experiment Video

Updated: May 17, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

466

TELEIA: A Spanish language dataset for evaluating artificial intelligence models.

Marina Mayor-Rocher1, Nina Melero2,3, Elena Merino-Gómez4

  • 1Facultad de Filosofía y Letras, Universidad Autónoma de Madrid, c/ Francisco Tomás y Valiente 1, 28049 Madrid, Spain.

Data in Brief
|March 31, 2025
PubMed
Summary

This study introduces TELEIA, a new dataset for evaluating Spanish language knowledge in Large Language Models (LLMs). TELEIA offers Spanish-specific multiple-choice questions to improve LLM assessment beyond translated English tests.

Keywords:
EvaluationLarge language modelsMachine learningNatural language processingSpanish

More Related Videos

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

636
Eye-tracking Technology and Data-mining Techniques used for a Behavioral Analysis of Adults engaged in Learning Processes
10:43

Eye-tracking Technology and Data-mining Techniques used for a Behavioral Analysis of Adults engaged in Learning Processes

Published on: June 10, 2021

5.2K

Related Experiment Videos

Last Updated: May 17, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

466
Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

636
Eye-tracking Technology and Data-mining Techniques used for a Behavioral Analysis of Adults engaged in Learning Processes
10:43

Eye-tracking Technology and Data-mining Techniques used for a Behavioral Analysis of Adults engaged in Learning Processes

Published on: June 10, 2021

5.2K

Area of Science:

  • Natural Language Processing
  • Computational Linguistics
  • Artificial Intelligence

Background:

  • Existing Large Language Model (LLM) evaluations often rely on translated English tests.
  • These translations may not accurately assess Spanish language proficiency.
  • There is a need for dedicated Spanish language evaluation datasets for LLMs.

Purpose of the Study:

  • To introduce TELEIA, a novel dataset for evaluating Spanish language knowledge in Large Language Models (LLMs).
  • To provide a resource that complements existing English-centric LLM evaluation tools.
  • To facilitate more accurate and nuanced assessment of LLMs' Spanish language capabilities.

Main Methods:

  • Development of a multiple-choice question dataset (TELEIA) specifically designed for Spanish language evaluation.
  • Questions are structured to match the format and difficulty of human Spanish proficiency tests.
  • The dataset includes 100 questions prepared and validated by Spanish language experts.

Main Results:

  • TELEIA provides a specialized tool for assessing LLMs' understanding of Spanish.
  • The dataset's multiple-choice format allows for automated testing and integration into LLM Leaderboards.
  • TELEIA aims to enhance the evaluation of LLMs in Spanish language tasks.

Conclusions:

  • TELEIA represents a significant step towards more accurate evaluation of LLMs in Spanish.
  • The dataset will be incorporated into the first Leaderboard for Spanish LLMs, promoting standardized assessment.
  • This resource is crucial for advancing LLM development and application in the Spanish-speaking world.