Related Experiment Video
Updated: Jul 25, 2025

Continuous Noninvasive Measuring of Crayfish Cardiac and Behavioral Activities
Published on: February 6, 2019
CsFEVER and CTKFacts: acquiring Czech data for fact verification.
Herbert Ullrich1, Jan Drchal1, Martin Rýpar1
1Artificial Intelligence Center, Faculty of Electrical Engineering, Czech Technical University in Prague, Charles Square 13, 120 00 Prague 2, Czech Republic.
This study introduces methods for acquiring Czech data for automated fact-checking, creating datasets like CsFEVER-NLI and CTKFactsNLI. These resources aid in classifying textual claim veracity and Natural Language Inference tasks.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Information Retrieval
Background:
- Automated fact-checking requires labeled datasets for training veracity classification models.
- Existing datasets are often language-specific, necessitating the creation of resources for underrepresented languages like Czech.
Purpose of the Study:
- To develop and evaluate methods for acquiring Czech data for automated fact-checking.
- To create novel Czech datasets for fact-checking and Natural Language Inference (NLI).
- To analyze datasets for potential biases and provide baseline models.
Main Methods:
- Hybrid approach of machine translation and document alignment to create a Czech version of the FEVER dataset (CsFEVER-NLI).
- Annotation of a new dataset (CTKFacts) using Czech News Agency articles, followed by cleaning and error analysis.
- Development of standalone NLI datasets (CsFEVER-NLI, CTKFactsNLI) from the acquired data.
Main Results:
- Publication of CsFEVER-NLI (127k translations) and CTKFactsNLI datasets.
- Analysis of datasets for spurious cues and annotator errors, with a typology of common mistakes.
- Provision of baseline models for fact-checking pipeline stages.
Conclusions:
- The developed methods and datasets facilitate automated fact-checking and NLI research in Czech.
- The study highlights the importance of dataset quality and bias analysis for robust fact-checking models.
- Published resources and tools aim to advance cross-lingual fact-checking capabilities.
Related Concept Videos
Data Validation
Key parameters for method validation include:
Censoring Survival Data
Data Reporting and Recording
Data Collection by Survey
Convenience Sampling Method
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...
Data Collection I

