Related Experiment Video
Updated: Jun 19, 2026

13:01
Industrialized, Artificial Intelligence-guided Laser Microdissection for Microscaled Proteomic Analysis of the Tumor Microenvironment
Published on: June 3, 2022
3.6K
Open-source LLMs for text annotation: a practical guide for model setting and fine-tuning
Meysam Alizadeh1, Maël Kubli1, Zeynab Samei2
1Department of Political Science, University of Zurich, 8050 Zurich, Switzerland.
Summary
Fine-tuning open-source Large Language Models (LLMs) significantly boosts performance on political science text classification tasks, outperforming zero-shot models and offering a practical alternative to few-shot training.
Area of Science:
- Political Science
- Computational Social Science
- Natural Language Processing
Background:
- Open-source Large Language Models (LLMs) are increasingly used for text analysis.
- Scholars require guidance on LLM performance for specific political science tasks.
- Establishing benchmarks for LLM effectiveness in social science research is crucial.
Purpose of the Study:
- To evaluate the performance of open-source LLMs in political science text classification.
- To compare zero-shot and fine-tuned LLM capabilities on tasks like stance, topic, and relevance.
- To provide a benchmark for LLM effectiveness and inform scholarly decision-making.
Main Methods:
- Assessed open-source LLMs using both zero-shot and fine-tuned approaches.
- Utilized news articles and tweets datasets for text annotation tasks.
- Compared fine-tuning against few-shot training with limited annotated data.
Main Results:
- Fine-tuning enhances open-source LLM performance, matching or exceeding zero-shot GPT-3.5 and GPT-4.
- Fine-tuned open-source LLMs still lag behind fine-tuned GPT-3.5.
- Fine-tuning is more effective than few-shot training with modest data.
Conclusions:
- Fine-tuned open-source LLMs are suitable for diverse text annotation applications in political science.
- The study provides a practical benchmark for LLM performance in social science research.
- A Python notebook is available to assist researchers in applying LLMs for text annotation.
Related Concept Videos
Light Acquisition
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...

