Related Experiment Video
Updated: Feb 1, 2026

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
An empiric approach to identifying physician peer groups from claims data: An example from breast cancer care
Jeph Herrin1,2, Pamela R Soulos2,3, Xiao Xu2,4
1Section of Cardiovascular Medicine, Yale University School of Medicine, New Haven, Connecticut.
Objective:
To develop an empiric approach for evaluating the performance of physician peer groups based on patient-sharing in administrative claims data.
Data Sources:
Surveillance, Epidemiology and End Results-Medicare linked dataset.
Study Design:
Applying social network theory, we constructed physician peer groups for patients with breast cancer. Under different assumptions of key parameter values-minimum patient volume for physician inclusion and minimum number of patients shared between physicians for a connection-we compared agreement in group membership between split samples during 2004-2006 (T1) (reliability) and agreement in group membership between T1 and 2007-2009 (T2) (stability). We also compared the results with those derived from randomly generated groups and to hospital affiliation-based groups.
Principal Findings:
The sample included 142 098 patients treated by 43 174 physicians in T1 and 136 680 patients treated by 51 515 physicians in T2. We identified parameter values that resulted in a median peer group reliability of 85.2 percent (Interquartile range (IQR) [0 percent, 96.2 percent]) and median stability of 73.7 percent (IQR [0 percent, 91.0 percent]). In contrast, stability of randomly assigned peer groups was 6.2 percent (IQR [0 percent, 21.0 percent]). Median overlap of empirical groups with hospital groups was 32.2 percent (IQR [12.1 percent, 59.2 percent]).
Conclusions:
It is feasible to construct physician peer groups that are reliable, stable, and distinct from both randomly generated and hospital-based groups.
Related Concept Videos
Testing a Claim about Mean: Known Population SD
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...
Model Approaches for Pharmacokinetic Data: Compartment Models
Two primary types of compartment models are recognized: mammillary and catenary. The more...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Empirical Method to Interpret Standard Deviation
This rule is used widely in statistics to calculate the proportion of data values...
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used;...

