全球创新指数面板数据的综合数据集 (2013-2022):用K-means进行集群和主要组件分析
Edilvando Eufrazio1,2, Helder Costa2
1Instituto Nacional de Tecnologia, Av. Venezuela, 82, Rio de Janeiro, 20081-312, RJ, Brazil.
Data in brief
|November 11, 2025
概括
本研究分析了使用全球创新指数 (GII) 2013-2022 年数据的国家创新格局. 它根据创新特征确定了五个不同的经济集群,帮助进行比较分析和战略决策.
科学领域:
- 经济学 经济学 经济学
- 创新研究 研究 创新研究
- 数据科学数据科学数据科学
背景情况:
- 创新是全球政策和研究重点.
- 全球创新指数 (GII) 提供国家创新数据.
- 了解国家创新格局对于经济发展至关重要.
研究的目的:
- 用GII数据分析国家创新格局.
- 通过集群识别具有类似创新特征的经济体.
- 为创新战略和政策提供数据驱动的框架.
主要方法:
- 利用了118个经济体 (2013-2022) 的全球创新指数 (GII) 数据.
- 应用K-means集群使用肘子方法来识别五个不同的经济集群.
- 采用主要组件分析 (PCA) 来增强集群差异化和减少维度.
主要成果:
- 根据创新支柱确定了五个不同的经济集群.
- 分析显示,当专注于输入,输出或组合支柱时,不同的创新概况.
- 通过PCA增强的集群提供了更清晰的划分和额外的分类标签.
结论:
- 该数据集提供了对国家创新的细分观点,超越了聚合得分.
- 这五个集群的框架有助于理解同行经济和长期创新趋势.
- 这些数据支持研究人员,政策制定者和商业领袖的基于证据的决策.
更多相关视频
07:50Global and Current Research Trends of Single-Cell Sequencing in Cancer: A Bibliometric and Visualization Study
Published on: April 18, 2025
839
09:01A Method for Investigating Age-related Differences in the Functional Connectivity of Cognitive Control Networks Associated with Dimensional Change Card Sort Performance
Published on: May 7, 2014
10.5K
相关概念视频
Cluster Sampling Method
13.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
13.9K
Outliers and Influential Points
5.9K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
5.9K
Variability: Analysis
430
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
430
Central Tendency: Analysis
459
Measures of central tendency are tools used in biostatistics to identify the average or center of a dataset. They offer a single representative value for understanding and summarizing data distribution.
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
459
Factorial Design
13.7K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.7K
Coefficient of Variation
8.1K
The coefficient of variation measures the dispersion of the data points or distribution around the mean. Using the coefficient of variation, we can compare two data series with drastically different means or different units of measurement. The coefficient of variation for a sample and a population is expressed as a percentage of the ratio of standard deviation to the mean.
The coefficient of variation is a practical statistical tool in finance. It allows investors to assess the volatility or...
The coefficient of variation is a practical statistical tool in finance. It allows investors to assess the volatility or...
8.1K
