Related Experiment Video
Updated: Feb 13, 2026

The Other End of the Leash: An Experimental Test to Analyze How Owners Interact with Their Pet Dogs
Published on: October 13, 2017
Validity of a two-stage cluster sampling design to estimate the total number of owned dogs
Oswaldo Santos Baquero1, Marcos Amaku1, Ricardo Augusto Dias1
1Department of Preventive Veterinary Medicine and Animal Health, School of Veterinary Medicine and Animal Science, University of São Paulo, Av. Prof. Orlando Marques de Paiva, 87, Cidade Universitária, São Paulo, SP, CEP: 05508-270, Brazil.
Abstract:
Estimates of owned dog population size are necessary to calculate measures of disease frequency and to plan and evaluate population management programs. We calculated the error and bias of estimates of the total number of owned dogs using a two-stage cluster sampling design. The estimates were conditioned on sample composition as well as on size and heterogeneity of the spatial distribution of owned dog populations. For this, we simulated nine cities that differed systematically in size (number of census tracts) and heterogeneity (variance of the number of dogs per census tract). Then, we defined 16 scenarios to calculate the sample composition using an algorithm that incorporated data from a pilot sample, estimates of cost, and prior specifications of the expected error and confidence level. In three additional scenarios of predefined sample composition, the numbers of primary and secondary sampling units were: 30 × 30, 50 × 20 and 65 × 15. Finally, for each city and sample composition, we selected primary sampling units (census tracts) with probability proportional to its size and with replacement, and secondary sampling units (households) by simple random sampling. For each city and composition we selected 500 samples, totaling 85500 samples. The distribution of errors conditioned on the sample composition and city showed that estimates were accurate (average mean bias = 0.006%, maximum mean bias = 0.3%). All sample compositions resulted in errors between 4% and 7% in cities with low heterogeneity. In cities with high heterogeneity, the errors for the various compositions ranged as follows: 8-11% (calculated), 11-13% (65 × 15), 12-14% (50 × 20) and 15-17% (30 × 30). The sample size of predefined compositions was between 33% and 87% lower than the sample size of calculated compositions. Therefore, the predefined compositions have an operational advantage (reduced sampling effort) and simplify the sampling design (calculation of sample composition is not needed). Furthermore, the expected error of estimates under different scenarios is known for each predefined composition. In the absence of information about the heterogeneity of the cities, the 65 × 15 is the more conservative composition.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Reliability and Validity
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
Group Design
Vesicular Tubular Clusters
With the help of motor proteins such...
Data Validation
Key parameters for method validation include:

