增强的机器学习组合方法用于估计储条件下的石油形成体积因子
Parsa Kharazi Esfahani1,2, Kiana Peiro Ahmady Langeroudy1,3, Mohammad Reza Khorsand Movaghar4
1Department of Petroleum Engineering, Amirkabir University of Technology (Tehran Polytechnic), 424 Hafez Avenue, Box 15875-4413, Tehran, 1591634311, Iran.
Scientific reports
|September 14, 2023
概括
本研究引入了先进的机器学习模型,包括XGBoost,以准确预测石油形成体积因子 (Bo),使用易于获得的井头数据. 新方法显著提高了预测准确度,与以前的方法相比,错误减少了四倍.
科学领域:
- 石油工程是石油工程中的一个.
- 储水库工程 储水库工程
- 机器学习应用 机器学习应用
背景情况:
- 石油形成体积因子 (Bo) 对于水库工程计算至关重要,包括估计现场原始石油和产量预测.
- 传统的Bo预测方法往往耗时或依赖于复杂的组成分析.
- 在石油中溶解的气体,压力,API重力和温度是影响石油体积变化的关键因素.
研究的目的:
- 开发和评估机器学习模型,准确预测石油形成体积因子 (Bo).
- 使用易于获取的井头数据作为输入,绕过需要昂贵和耗时的实验室组合分析.
- 将渐变增强决策树 (GBDT) 技术的性能与现有的相关性和其他机器学习方法进行比较.
主要方法:
- 使用GBDT技术开发机器学习模型,特别是极端梯度提升 (XGBoost),梯度提升和CatBoost.
- 开发的模型与传统的相关性和基于树的包装方法进行比较,例如额外树 (ETs),随机森林 (RF) 和决策树 (DTs).
- 使用统计和图形指标在各种水库压力条件 (气泡点上方和下方的压力) 中验证模型.
主要成果:
- 在所有测试的压力范围内,XGBoost模型在估计油形成体积因子 (Bo) 方面表现出卓越的性能.
- 拟议的XGBoost模型实现了平均绝对相对偏差仅为0.2598%,比以前的组合方法提高了四倍.
- 开发的模型只使用井头数据准确预测体积特性,从而消除了对常规实验室分析的需求.
结论:
- 机器学习模型,特别是XGBoost,为预测石油形成体积因子 (Bo) 提供了一个高度准确和高效的替代方案.
- 该研究成功地证明了使用简单的井头数据来准确估计水库流体属性的可行性.
- 这种方法大大降低了与石油工程应用的传统实验室分析相关的成本和时间.
相关概念视频
Estimation of the Physical Quantities
4.3K
On many occasions, physicists, other scientists, and engineers need to make estimates of a particular quantity. These are sometimes referred to as guesstimates, order-of-magnitude approximations, back-of-the-envelope calculations, or Fermi calculations. The physicist Enrico Fermi was famous for his ability to estimate various kinds of data with surprising precision. Estimating does not mean guessing a number or a formula at random. Instead, estimation means using prior experience and sound...
4.3K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
566
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
566
Residual Plots
4.6K
A residual plot is a statistical representation of data used to analyze correlation and regression results. It helps verify the requirements for drawing specific conclusions about correlation and regression. To obtain the residual plot, first, the residual for each data value is calculated, which is simply the vertical distance between the observed and the predicted value obtained from the regression equation.
When the residual values are plotted against the variable x, it is called a residual...
When the residual values are plotted against the variable x, it is called a residual...
4.6K
Extraction: Partition and Distribution Coefficients
2.5K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
2.5K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Mechanistic Models: Compartment Models in Individual and Population Analysis
64
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
64


