Related Experiment Video
Updated: Sep 8, 2025

Use of Principal Components for Scaling Up Topographic Models to Map Soil Redistribution and Soil Organic Carbon
Published on: October 16, 2018
Pattern Recognition for Steam Flooding Field Applications Based on Hierarchical Clustering and Principal Component
Na Zhang1,2, Mingzhen Wei2, Baojun Bai2
1Shandong University of Science and Technology, Qingdao 266590, China.
This study uses machine learning to group steam flooding projects based on reservoir and fluid properties. By applying hierarchical clustering and principal component analysis, the researchers reduced eight parameters to two key components while retaining most of the data’s variation. The analysis identified five clusters of projects with similar conditions and production outcomes. These findings suggest that clustering can help engineers make better decisions when designing new steam flooding operations. The study shows that similar reservoirs tend to have similar production results, which can guide future operations.
Area of Science:
- Petroleum engineering data analysis
- Enhanced oil recovery techniques
- Machine learning in geoscience
Background:
Steam flooding is a widely used method to enhance oil recovery from reservoirs containing both heavy and light oils. While traditional data analysis has been applied to assess these projects, machine learning approaches remain underutilized. Prior research has focused on conventional statistical methods, leaving a gap in understanding how advanced clustering techniques might reveal hidden patterns. This gap motivated the use of hierarchical clustering and principal component analysis to explore steam flooding data. No prior work had resolved how these methods could group similar projects effectively. The need for better analog identification in reservoir engineering remains unmet. Existing studies lack a systematic way to extract design insights from global steam flooding data. This paper introduces a novel approach to address these limitations. The field requires more precise tools for operational decision-making based on historical performance.
Purpose Of The Study:
This study aims to develop a method for identifying patterns in global steam flooding projects using machine learning. The specific problem is the lack of systematic analog identification in reservoir engineering. By clustering similar projects, the study hopes to improve operational design and production performance. The motivation comes from the need to extract hidden patterns from large datasets. Traditional methods fail to capture complex relationships between reservoir and fluid parameters. This paper proposes a solution using hierarchical clustering and principal component analysis. The goal is to group projects with similar reservoir conditions and production outcomes. The study seeks to provide actionable insights for future steam flooding operations.
Main Methods:
The study applies hierarchical clustering analysis (HCA) along with principal component analysis (PCA) to global steam flooding data. PCA is used to reduce the dimensionality of eight reservoir and fluid parameters. This transformation results in two principal components that retain about 90% of the original variance. The reduced data is then clustered using HCA with Euclidean distance and Ward’s linkage method. The clustering process identifies five distinct groups of steam flooding projects. Each cluster is defined by unique ranges of reservoir and fluid properties. The method allows for the comparison of operational designs and production performance across clusters. The approach enables the extraction of insights from historical data for future applications.
Main Results:
The analysis identified five distinct clusters of steam flooding projects based on reservoir and fluid parameters. Each cluster represents a unique range of properties and production performance. Principal component analysis successfully reduced eight parameters to two components while retaining 90% of the variance. The clustering results show that similar reservoir conditions lead to similar operational designs and production outcomes. The first cluster includes projects with high pressure and low viscosity. The second cluster features low-pressure reservoirs with high oil saturation. The third cluster includes high-temperature fields with moderate viscosity. The fourth cluster is defined by high water cut and low production rates. The fifth cluster represents fields with high injection rates and low recovery efficiency. These findings suggest that clustering can guide operational decisions in steam flooding.
Conclusions:
The authors conclude that hierarchical clustering and principal component analysis can effectively group steam flooding projects based on reservoir and fluid parameters. The results suggest that similar reservoir conditions lead to similar operational designs and production outcomes. The study demonstrates that machine learning can extract hidden patterns from global steam flooding data. The findings indicate that clustering can guide operational decisions in steam flooding. The method allows for the comparison of production performance across clusters. The results support the use of data-driven approaches in reservoir engineering. The study proposes that clustering can improve the design of future steam flooding projects. The authors suggest that this approach can enhance decision-making in enhanced oil recovery.
Frequently Asked Questions
The main outcome is the identification of five distinct clusters of steam flooding projects based on reservoir and fluid parameters.
PCA reduces eight reservoir and fluid parameters to two principal components while retaining about 90% of the original variance.
Ward’s linkage minimizes the variance within each cluster, ensuring distinct groupings of similar steam flooding projects.
Each cluster represents a unique range of reservoir conditions and production performance, guiding operational design choices.
The two components capture 90% of the variance in eight reservoir and fluid parameters, simplifying data interpretation.
The authors propose that clustering can improve operational design and production performance by identifying similar reservoir conditions.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Design Example: Analyzing Capacity Contours for Flood Risk Assessment
Uniform Depth Channel Flow: Problem Solving
Turbulent Flow: Problem Solving
Temperature is a key factor in CO2 solubility. In this case, the CO2 gas and the liquid are cooled to 20°C. Lower temperatures...

