一个GAN-BO-XGBoost模型用于高质量的专利识别
Zengyuan Wu1, Jiali Zhao2, Ying Li2
1College of Economics and Management, China Jiliang University, No. 258, Xueyuan Street, Hangzhou, 310018, Zhejiang, People's Republic of China. wuzengyuan@cjlu.edu.cn.
Scientific reports
|April 26, 2024
概括
识别高质量的专利至关重要,但由于其普及率较低,因此具有挑战性. 我们开发了一个GAN-BO-XGBoost模型,增强专利分析,以便更好地做出研发决策.
科学领域:
- 知识产权管理知识产权管理
- 数据科学数据科学数据科学
- 机器学习 机器学习
背景情况:
- 专利申请的快速增长使得区分高质量的专利和低质量的专利成为一个挑战.
- 准确识别高质量的专利对于明智的研发决策和战略专利组合管理至关重要.
- 高品质专利的稀缺性使得它们的有效识别成为一个重大障碍.
研究的目的:
- 开发一个改进的模型,以有效和准确地识别高质量的专利.
- 通过结合新的功能来增强现有的专利质量评估系统.
- 解决专利质量分类中数据集不平衡的挑战.
主要方法:
- 通过增加与专利权人的技术实力相关的四个特征来重建专利质量指数系统.
- 整合重新采样技术,特别是生成对抗网络 (GAN),以扩大少数 (高质量专利) 样本.
- 应用一个集体学习算法,极端梯度提升 (XGBoost),优化与贝叶斯优化 (BO),形成GAN-BO-XGBoost模型.
主要成果:
- 拟议的GAN-BO-XGBoost模型在识别高质量的专利方面,与其他模型相比,表现出更高的性能和稳定性.
- 使用石版技术领域的专利数据进行的评估和十倍交叉验证证实了该模型的有效性.
- 整合GAN用于数据增强和BO-XGBoost用于分类,显著提高了识别准确性.
结论:
- GAN-BO-XGBoost模型为准确有效地识别高质量专利提供了强大而有效的解决方案.
- 这种方法提高了组织做出战略研发决策和优化专利布局的能力.
- 该研究强调了结合先进的机器学习技术来应对知识产权分析方面的挑战的潜力.
相关概念视频
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Pharmacokinetic Models: Comparison and Selection Criterion
69
Physiological and compartmental models are valuable tools used in studying biological systems. These models rely on differential equations to maintain mass balance within the system, ensuring an accurate representation of the dynamic processes at play.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
69
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
120
Drug disposition in the body is a complex process and can be studied using two major approaches: the model and the model-independent approaches.
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
120
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K


