Related Experiment Video
Updated: Oct 7, 2026

An Affordable HIV-1 Drug Resistance Monitoring Method for Resource Limited Settings
Published on: March 30, 2014
Natural Language Processing Analysis of WhatsApp Messages to Identify Unrecognized HIV Prevention Needs Among Young
Jessica E Haberer1,2, Bernard Rono3, Rakesh Ravi4
1Center for Global Health, Massachusetts General Hospital, Boston, Massachusetts, USA.
Introduction:
Identifying HIV prevention needs is challenging. Common approaches (e.g. public health campaigns, HIV risk scores) have limited sensitivity. Natural language processing techniques, such as topic modelling and computational linguistics, can extract meaning and subjective information from text. This study used these techniques with WhatsApp messages to identify unrecognized HIV prevention needs among young women in Kenya.
Methods:
Prior work established key ethical protections (e.g. privacy, confidentiality). Between July and December 2024, women (18-24 years) who used WhatsApp were enrolled from four clinical sites in Kisumu, Kenya. After completing a socio-demographic questionnaire, participants securely transferred all WhatsApp messages from the prior 6 months and provided anonymized relationship types (e.g. friend, sexual partner) for their 20 most frequent contacts. Data were de-identified, and only outgoing messages were retained; all photos and incoming messages were deleted. WhatsApp data were pre-processed using a downloaded large language model to translate Swahili and Dholuo to English and account for slang linguistic patterns. We explored multiple approaches to topic modelling in the WhatsApp data using latent Dirichlet allocation. Non-negative matrix factorization and uniform manifold approximation and projection allowed for comparison of participant word topics with the VOICE HIV risk score (range 0-8, 8 indicating high risk).
Results:
Of 413 eligible young women, 13 (3%) declined to provide WhatsApp data, and 22 had insufficient data for analysis. Participants had a mean age of 22 years and a VOICE risk score of 5.7. The remaining 378 women provided a median of 14 chats per individual. Participants most commonly chatted with friends (62.4%), followed by sexual partners (18.1%). The median VOICE risk score was 7 for women whose most messaged relationship was a sexual partner versus 5 for those messaging a friend most often. Topic models had differential correlation with the VOICE risk score that was most strongly positive (0.12) when indicating romance/affection and most strongly negative (-0.09) when describing faith/transactions. Dimensional analysis of linguistic similarity identified discrete participant populations with distinct VOICE risk scores.
Conclusions:
This study provides proof-of-concept that natural language processing of routine WhatsApp messages can yield meaningful individual-level information about HIV prevention needs. Future work should compare such findings with HIV incidence to optimally determine its ability to identify unrecognized HIV prevention needs, thus potentially expanding HIV prevention services to individuals not currently engaged in care.

