在世界各地的社会科学和医学文章中使用数据
Brian Stacy1, Lucas Kitzmüller2, Xiaoyu Wang1
1World Bank, Development Data Group, 1818 H St NW, Washington, DC 20433, USA.
PNAS nexus
|June 27, 2025
概括
本研究介绍了一种方法来追踪学术数据的使用情况. 调查结果显示,高收入国家主导着数据驱动的研究,突出了基于证据的政策研究的全球差异.
科学领域:
- 社会科学 社会科学 社会科学
- 医学 医学 医学 医学 医学
- 公共政策 公共政策
背景情况:
- 基于证据的公共政策依赖于数据驱动的研究.
- 目前对研究中全球数据利用的理解是有限的.
- 识别研究缺口对于政策制定至关重要.
研究的目的:
- 开发一种方法来追踪学术数据的使用情况.
- 分析数据驱动研究的地理分布.
- 确定可以从增加数据生产或使用中受益的国家.
主要方法:
- 自然语言处理 (NLP) 应用于大量英语社会科学和医学文章.
- 开发一个模型来估计学术出版物中国家特定数据的使用情况.
- 对NLP模型与人类编码数据的验证 (相关性为0.99).
主要成果:
- 分析了超过14万篇学术文章.
- 高收入国家是大约50%使用数据的论文的主题.
- 在数据驱动研究中,高收入国家 (占全球人口的17%) 的比例不成比例.
结论:
- 在全球数据驱动研究中存在显著的不平衡,有利于高收入国家.
- 较贫穷的国家可能会从增加数据生产中受益,而较富裕的国家可以提高数据利用率.
- 开发的NLP方法提供了一个可扩展的方法来监测全球研究趋势并为政策提供信息.
相关概念视频
Data Collection by Experiments
Data collection is a systematic method of obtaining, observing, measuring, and analyzing accurate information. An experimental study is a standard method of data collection that involves the manipulation of the samples by applying some form of treatment prior to data collection. It refers to manipulating one variable to determine its changes on another variable. The sample subjected to treatment is known as “experimental units.”
An example of the experimental method is a public clinical trial...
An example of the experimental method is a public clinical trial...
Data Collection I
Data collection gathers information needed to make accurate judgments about a patient's present condition. During a health history interview, subjective data is collected from the patient, their caregivers, or family members, and objective data is collected through observations and physical assessment. Patients are the primary source of subjective data. Thus information gathered from patients through interviews, observations, and physical examination is primary data. Secondary sources of data...
Biostatistics: Overview
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
Data: Types and Distribution
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Distributions in...
Statistical Methods for Analyzing Epidemiological Data
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
Statistical Software for Data Analysis and Clinical Trials
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...


