Related Experiment Video
Updated: Feb 10, 2026

Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
Published on: December 9, 2022
[Diabetes prevalence estimated using a standard algorithm based on electronic health data in various areas of Italy]
Roberto Gnavi1, Ludmila Karaghiosoff, Daniela Balzi
1Servizio di epidemiologia ASL TO3, Regione Piemonte. roberto.gnavi@epi.piemonte.it
Aims:
the goal of this study was to estimate the prevalence of diabetes through record linkage of various data sources in four Italian areas.
Setting:
Aulss 12 Veneziana, Aulss 4 Alto vicentino, Torino, ASL10 of Firenze.
Participants:
all 2002 to 2004 residents in the four areas (n = 2,123,913 on 30th June 2003).
Main Outcome:
crude prevalence by age and gender and standardized prevalence by gender.
Methods:
we used three different data sources. The first was the set of files of all persons discharged from hospitals with a primary or secondary diagnosis of diabetes (ICD-9-CM code 250*) in the year of interest or in the four previous years. The second data source was the set of files of all prescriptions of antidiabetic drugs (ATC code: A10A* and A10B*) prescribed in the year of interest; we considered as persons with diabetes only those who had at least two prescriptions of antidiabetic drugs at two different times. The third source was the set of files of all subjects who obtained exemption from payment of drugs or laboratory testing due to a diagnosis of diabetes mellitus in the year of interest or in the 3 previous years. All data sources were matched by a deterministic linkage procedure. We defined as "prevalent case" those persons who were present in at least one of the three data sources. We compared the estimated prevalence in the four different areas.
Results:
in 2003, the prevalence of diabetes in the four areas ranged from 3.93% to 5.55% among men, and from 3.55% to 4.52% among women. After adjustment for age, differences among men were reduced and were no longer present among women. Prevalence is higher among the elderly and among men.
Conclusions:
using routinely collected data we were able to identify large cohorts of persons with known diabetes and to estimate the prevalence of the disease, which was shown to be highly homogeneous among participating centres, and similar to that reported in other studies conducted in Italy with more costly and time consuming methods.
More Related Videos
Related Concept Videos
Estimating Population Standard Deviation
Estimating Population Mean with Known Standard Deviation
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
Estimating Population Mean with Unknown Standard Deviation
William S. Gosset (1876–1937) of the...
Prevalence and Incidence
Prevalence indicates the proportion of individuals in a population who have a specific disease or health...
Data Reporting and Recording
Trial and Error and Algorithm

