模拟公众论:比较LLM和随机森林的分布和个人级预测
Fernando Miranda1, Pedro Paulo Balbi1,2
1Programa de Pós-Graduação em Engenharia Elétrica e Computação, Universidade Presbiteriana Mackenzie, São Paulo 01302-000, SP, Brazil.
Entropy (Basel, Switzerland)
|September 27, 2025
概括
大型语言模型 (LLM) 可以使用调查数据模拟个人意见. 与传统模型相比,LLM可以更好地预测综合意见分布,从而推进了计算社会科学.
科学领域:
- 计算社会科学 计算社会科学
- 人工智能的人工智能
- 政治科学 政治科学是指政治学.
背景情况:
- 传统的基于代理的模型使用简化的人类决策规则.
- 建模信息流是理解两极分化和意见动态的关键.
研究的目的:
- 探索大型语言模型 (LLM) 作为模拟意见的高保真代理.
- 评估LLM使用真实调查数据预测个体反应的能力.
- 将LLM模拟与传统模型进行比较.
主要方法:
- 根据2020年美国全国选举研究 (ANES) 调查数据,有条件的法学士.
- 用Jensen-Shannon距离和F1得分进行评估.
- 将LLM模拟与监督的随机森林模型进行比较.
主要成果:
- 在个人层面上,LLM的表现与Random Forest的表现相当.
- 总体而言,LLM系统始终以更接近实证数据的总体意见分布.
- 在零射击环境中,LLM表现出强大的预测准确性.
结论:
- 在模拟复杂的意见动态方面,LLM显得有前途.
- 在计算社会科学中,LLM提供了一种新的方法来建模信念系统.
- 在模拟中,LLM可以捕捉微妙的,情境敏感的人类决策.
更多相关视频
04:35Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
3.7K
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
490
相关概念视频
Distributions to Estimate Population Parameter
5.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
5.0K
Prediction Intervals
3.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.3K
Survival Tree
388
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
388
Choosing Between z and t Distribution
3.5K
The z and the Student t distribution estimate the population mean using the sample mean and standard deviation. However, to decide which distribution to use for a calculation, one needs to determine the sample size, the nature of the distribution, and whether the population standard deviation is known. If the population standard deviation is known and the population is normally distributed, or if the sample size is greater than 30, the z distribution is preferred. The Student t distribution is...
3.5K
Randomized Experiments
8.9K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
8.9K
Estimating Population Standard Deviation
3.3K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.3K
