Probabilistic classifiers and automated cancer registration: an exploratory application
Sandro Tognazzo1, Bovo Emanuela, Fiore Anna Rita
1Venetian Tumour Registry, Registro Tumori del Veneto, Istituto Oncologico Veneto-IRCCS, 35128 Padua, Italy. sandro.tognazzo@ioveneto.it
Journal of Biomedical Informatics
|July 16, 2008
Summary
Random forests effectively classify cancer cases automatically, reducing manual checks. This machine learning approach shows promise for improving cancer registry efficiency with a low error rate.
Area of Science:
- Oncology
- Biostatistics
- Machine Learning
Background:
- Cancer registries rely on manual case definition, which is time-consuming.
- Accurate identification of primary and secondary cancer cases is crucial for epidemiological studies.
Purpose of the Study:
- To evaluate the performance of random forests and multinomial logit models for automated cancer case classification.
- To assess the potential reduction in manual review workload for cancer registries.
Main Methods:
- Utilized data from 5608 subjects registered by the Venetian Tumour Registry (1987-1996).
- Employed an eightfold cross-validation technique to estimate classification error.
- Included 63 predictive variables for model fitting using random forests and multinomial logit models.
Main Results:
- Random forests automatically classified 45% of subjects with <5% error.
- Multinomial logit models achieved a 31% classification error.
- Potential reduction of manually checked cases from 1750 to 960 per incidence year.
Conclusions:
- Random forests demonstrate a promising approach for automated cancer case definition.
- This method could significantly decrease the manual workload in cancer registries.
- Further refinement and application to other case categories are recommended.
