Related Experiment Video
Updated: Sep 4, 2026

Analyzing Tumor Gene Expression Factors with the CorExplorer Web Portal
Published on: October 11, 2019
Investigating Online Discussions About Cancer Screening on Twitter (Subsequently Rebranded as X): Corpus Analysis
Martin-Pieter Jansen1,2, Hanneke Hendriks3, Suzan Verberne4
1Center for Language Studies, Radboud University Nijmegen, Nijmegen, Gelderland, The Netherlands.
Background:
While cancer screening is proven to be effective in the early detection of the disease and early detection enables better treatment options, screening uptake has been declining. Research shows that online health information helps people to make health-related decisions. However, not all online health information is credible, and misinformation might play a role in people's choice to take part in screening.
Objective:
This study aimed to analyze online discussions about cancer screening programs using corpus analysis. Specifically, we aimed to investigate the full dataset through corpus analysis and misinformation in a manually coded subset. This enabled us to study naturalistic discussions about cancer screening over time, what information people share, and how prevalent misinformation is in these discussions. We differentiated tweets on Twitter (subsequently rebranded as X) for cervical, breast, colorectal, and general screening.
Methods:
We extracted a corpus of 55,403 tweets from 2011 to 2023, tweeted by 22,493 users from a database containing over 5.9 billion tweets. We used specific search strings corresponding to different types of screening to gather our corpus. The corpus consisted of tweets, timestamps, hashtags, and shared URLs. We used a machine learning classifier trained on another dataset of tweets about cancer screening to automatically code whether a tweet fell within the scope of the study. We manually coded a randomly drawn stratified subset of 1200 tweets representative of the full corpus regarding year and screening program for the presence of misinformation.
Results:
Tweets were not uniformly distributed across different screening programs and over time (χ²36=4045.99, n=55,403; P<.001). Most tweets discussed population screening in general (n=35,199), and the volume of tweets increased around real-world events. Hashtags in the tweets predominantly focused on the screening programs that were discussed in those tweets. In our corpus, most shared URLs linked to other tweets (n=10,569) or news websites (n=2807). In our coded subset, information was shared in 679 tweets. Overall, 23 tweets contained misinformation. Topics in those tweets showed criticism toward the programs and policies, suspicions about conflicts of interest, and antivaccination beliefs regarding human papillomavirus (HPV). Most users used rhetorical questions, sarcasm, fearmongering, or expressed anger.
Conclusions:
Our findings reveal that cancer screening programs are actively debated across social media platforms. We observed that conversations tend to spike in response to real-world events, suggesting social media can serve as a valuable lens into public reactions to health policy changes. Link-sharing behavior was common, though we noted a tendency for sources to reference back to the same platform where discussions originated. Despite finding limited instances of misinformation, we caution that even modest amounts of inaccurate information may have meaningful consequences for public health messaging and screening uptake.
More Related Videos
08:53Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
Published on: May 31, 2019
07:41Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019