Related Experiment Video
Updated: Dec 29, 2025

10:56
A User-friendly and Powerful R Analysis of Large-scale Datasets
Published on: November 4, 2025
231
Adverse Events in Twitter-Development of a Benchmark Reference Dataset: Results from IMI WEB-RADR
Juergen Dietrich1, Lucie M Gattepaille2, Britta Anne Grum3
1Pharmacovigilance, Bayer AG, Müllerstr. 170, 13353, Berlin, Germany. juergen.dietrich@bayer.com.
Drug Safety
|January 31, 2020
Summary
Researchers created a benchmark dataset from Twitter data to evaluate automated adverse event recognition systems. This dataset aids in analyzing social media for drug safety information and improving pharmacovigilance.
Area of Science:
- Pharmacovigilance and Drug Safety
- Computational Linguistics
- Public Health Informatics
Background:
- Social media platforms like Twitter are increasingly recognized as valuable sources for pharmacovigilance, supplementing traditional safety surveillance methods.
- Automated systems are being developed to detect adverse events from unstructured text data, but require robust datasets for performance evaluation.
Purpose of the Study:
- To create a benchmark reference dataset from Twitter data for evaluating automated adverse event recognition systems.
- To facilitate the assessment of methods for extracting drug safety information from social media.
Main Methods:
- A retrospective analysis of 5,645,336 English-language Tweets mentioning six specific medicinal products was conducted.
- Two independent teams of safety reviewers extracted and coded product-event and product-indication combinations from a sample of 57,473 Tweets.
- A dataset comprising 1056 positive controls (adverse event Tweets) and 56,417 negative controls (non-adverse event Tweets) was curated.
Main Results:
- The benchmark dataset contains 1056 "adverse event Tweets" and 56,417 "non-adverse event Tweets".
- These "adverse event Tweets" include 1396 product-event combinations, referencing 292 different MedDRA Preferred Terms, with 83.9% falling into four System Organ Classes.
- Indication information was present in 195 Tweets, covering 25 Preferred Terms.
Conclusions:
- A manually curated benchmark reference dataset derived from Twitter data has been successfully created.
- This dataset is now available to the research community for the validation of automated adverse event recognition systems.
- The dataset supports the advancement of using unstructured social media data for drug safety surveillance.

