Related Experiment Video
Updated: Oct 1, 2026

A User-friendly and Powerful R Analysis of Large-scale Datasets
Published on: November 4, 2025
[Influence of data structure on the selection of statistic analysis methods]
Isabel María Barroso Utra1, Mayilée Cañizares Pérez, Lydia Lera Marqués
1Instituto Nacional de Higiene, Epidemiología y Microbiología, Infanta 1158 entre Clavel y Llinas, La Habana, Cuba. ibarroso@inhem.sld.cu
Abstract:
In medical research, data is grouped either as per the design of the study or the selection of the sample. This structure must be taken into account in order to make correct estimates of the of the parameters and standard errors involved. This study is of the methodological type and is aimed at illustrating methods for estimating population-related methods and regression models with grouped data. For this purpose, nine variables from the First National Risk Factor and Preventive Measure Survey conducted in Cuba in 1995 are employed. The prevalence of high blood pressure is overestimated by 15% when the conventional estimators are used as compared with the weight-based and adjusted analysis. In the regression models for the body mass index based on the conventional procedures, sex, degree of schooling, degree of sedentariness, smoking habit, diastolic and systolic blood pressure were found to be significant. However, when the method taking into account the structure of conglomerates was employed, the degree of schooling and sedentariness ceased to be significant. When the random intercept model was adjusted, the 91.3% total variability was found to be explained by individual variables, the 8.7% variability being attributed to larger units. When estimating population-related parameters based on conglomerate-structure data involving inconsistent selection probabilities, the use of sample-related weights and analysis methods that take in the correlation among subjects (potential) for one same conglomerate. When adjusting regression models, it is not only important to efficiently estimate the coefficients, but rather the focus (aggregated or disaggregated) must be taken into account for modeling the problem under study.
More Related Videos
09:43Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
Published on: November 22, 2019
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Related Concept Videos
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
Statistical Methods for Analyzing Epidemiological Data
Introduction to Nonparametric Statistics
One of...
Data: Types and Distribution
Distributions in...
Statgraphics
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with data...