Related Experiment Video
Updated: Jun 21, 2026

Pore-scale Imaging and Characterization of Hydrocarbon Reservoir Rock Wettability at Subsurface Conditions Using X-ray Microtomography
Published on: October 21, 2018
Filling-well: An effective technique to handle incomplete well-log data for lithology classification using machine
Sherly Ardhya Garini1,2, Ary Mazharuddin Shiddiqi1, Widya Utama2
1Department of Informatics, Institut Teknologi Sepuluh Nopember, Indonesia.
Extreme gradient boosting (XGBoost) effectively handles missing well-log data for lithology classification, outperforming K-nearest neighbours (KNN) and artificial neural networks (ANN). XGBoost shows superior accuracy, especially with up to 30% missing data, crucial for oil and gas exploration.
Area of Science:
- Geoscience and Petroleum Engineering
- Machine Learning Applications in Earth Sciences
- Data Science for Resource Exploration
Background:
- Lithology classification is vital for sustainable oil and gas exploration.
- Missing values in well-log data (e.g., GR, NPHI, RHOB, RS, DTCO, DTSM, RD) significantly reduce machine learning classification accuracy.
- Addressing missing data is critical for reliable subsurface characterization.
Purpose of the Study:
- To evaluate machine learning algorithms for handling missing values in well-log datasets.
- To improve lithology classification accuracy in the presence of significant data gaps (up to 30%).
- To compare the performance of XGBoost, KNN, and ANN in imputing missing well-log data.
Main Methods:
- Application of Extreme Gradient Boosting (XGBoost) algorithm.
- Implementation of K-nearest neighbours (KNN) algorithm.
- Utilization of Artificial Neural Network (ANN) algorithm, including preprocessing with isolation forest and bias correction.
Main Results:
- XGBoost demonstrated the highest efficiency and accuracy, achieving the lowest Mean Absolute Percentage Error (MAPE) and Root Mean Square Error (RMSE) for RHOB, NPHI, DTCO, and DTSM.
- ANN performed well on GR, RS, and RD features post-preprocessing but showed potential for overfitting with smaller datasets.
- KNN exhibited limitations with missing-not-at-random (MNAR) data due to its reliance on distance metrics and the k parameter.
Conclusions:
- XGBoost is the most effective algorithm for handling extreme missing values in well-log data for lithology classification.
- ANN offers a viable alternative, particularly for specific log types, but requires careful tuning and sufficient data.
- KNN is less suitable for MNAR well-log data imputation compared to XGBoost and ANN.
More Related Videos
08:56Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
Published on: January 13, 2023
09:21Author Spotlight: Generating Neuronal Phenotypic Profiles - A Protocol to Culture and Image Human Midbrain Dopaminergic Neurons
Published on: July 7, 2023