在建议中确定大型语言模型的公平性
Wei Liu1, Baisong Liu2, Jiangcheng Qin1
1Faculty of Electrical Engineering and Computer Science, Ningbo University, Ningbo, 315211, China.
Scientific reports
|February 14, 2025
概括
大型语言模型 (LLM) 可以通过识别用户属性相关性来识别不公平的建议. 整合LLM可以显著提高推的公平性,以最小的效用损失,平衡股权和绩效.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 推系统是一个推系统.
背景情况:
- 推系统的公平性对于公平的用户待遇至关重要.
- 大型语言模型 (LLM) 表现出类似人类的行为,包括公平意识.
- 现有的推模型可能会无意中延续偏见.
研究的目的:
- 调查LLM作为推系统中的公平性识别器.
- 为了利用LLM的公平意识来构建公平的建议.
- 提出一种方法,将LLM整合到建议管道中.
主要方法:
- 使用了MovieLens和LastFM数据集进行评估.
- 具有公平策略和没有公平策略的可变自动编码器 (VAE) 的比较.
- 雇佣了ChatGLM3-6B和Llama2-13B来评估建议的公平性.
- 开发了一种混合方法,使用LLM来完善VAE建议.
主要成果:
- 通过将用户属性与结果相关联,LLM有效地识别不公平的建议.
- 拟议的方法显著提高了推的公平性.
- 在整合后观察到推实用性的最小损失.
- 公平与实用性比率有了显著的改善,例如,使用ChatGLM,从5-6到30-50左右.
结论:
- 在推系统中,LLM可以作为有效的公平检测器.
- 将LLM整合到推流程中,可以提供更好的公平性-实用性权衡.
- 这种方法对开发更公平的人工智能驱动系统充满希望.
相关概念视频
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
The Representativeness Heuristic
15.8K
The representative heuristic describes a biased way of thinking, in which you unintentionally stereotype someone or something. For example, you may assume that your professors spend their free time reading books and engaging in intellectual conversation, because the idea of them spending their time playing volleyball or visiting an amusement park does not fit in with your stereotypes of professors.
15.8K
Confirmation Biases
5.4K
The confirmation bias is the tendency to focus on information that confirms our existing beliefs and ignore information that is inconsistent with our expectations. For example, if you think that your professor is not very nice, you notice all of the instances of rude behavior exhibited by the professor while ignoring the countless pleasant interactions he is involved in on a daily basis. Have you ever fallen prey to the confirmation bias, either as the source or target of such bias?
5.4K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Social Proof
27.4K
Social proof is a form of persuasion based on comparison and conformity. People compare their behavior and actions to what others are doing and will change to conform to do what their peers do.
27.4K
Law of Independent Assortment
53.5K
While Mendel’s Law of Segregation states that the two alleles for one gene are separated into different gametes, a different question of how different genes are inherited remains. For example, is the gene for tall plants inherited with the gene for green peas? Mendel asked this question by experimenting with a dihybrid cross; a cross in which both parents are homozygous for two distinct traits resulting in an F1 generation that are heterozygous for both traits.
53.5K


