语言模型的上下文微调与分类器驱动的内容调节用于文本生成
Matan Punnaivanam1, Palani Velvizhy1
1Department of Computer Science and Engineering, College of Engineering Guindy, Anna University, Chennai 600025, India.
Entropy (Basel, Switzerland)
|January 8, 2025
概括
这项研究开发了一个框架,使用微调的大型语言模型 (LLM) 来生成和分类适合年龄的儿童故事. 微调模型显著提高了故事质量,BERT分类器准确地识别了不合适的内容.
科学领域:
- 人工智能的人工智能
- 自然语言处理自然语言处理.
- 儿童发展 儿童发展
背景情况:
- 确保儿童内容的适当性在数字时代至关重要.
- 自动生成文本 (例如,大型语言模型 (LLM)) 需要有效的内容过工具.
- 现有的方法与儿童文学的细微差别作斗争.
研究的目的:
- 开发一个强大的框架,根据适用性来生成和分类儿童故事.
- 为弥合儿童文学现有的内容调节工具的差距.
- 为了利用微调的LLM进行适合年龄的内容创建和分类.
主要方法:
- 微调的LLM (LLaMA,Mistral,Zephyr) 用于上下文的故事生成.
- 使用基于BERT的分类器进行内容合适性评估.
- 使用ROUGE,METEOR和BERT分数评估生成的故事.
主要成果:
- 微调的Mistral-7B和Zephyr-7B-Beta模型在故事生成质量方面比基础模型显著改善.
- 在识别不合适的内容时,BERT分类器实现了高精度 (0.95) 和回忆 (0.97).
- 微调的模型产生了更符合人类标准的内容.
结论:
- 先进的LLM提供了一个有希望的方法来生成适合年龄的儿童故事.
- 开发的框架加强了儿童安全数字环境的内容调节策略.
- 这项研究对教育技术,内容策划和家长控制系统有影响.
相关概念视频
Regulation of Expression at Multiple Steps
868
The gene expression in cells is regulated at different stages: (i) transcription, (ii) RNA processing, (iii) RNA localization, and (iv) translation. Transcriptional regulation is mediated by regulatory proteins such as transcription factors, activators, or repressors—these control gene expression by initiating or inhibiting the transcription of genes. Once a precursor or pre-mRNA is produced, it undergoes post-transcriptional modification, including 5' capping, splicing, and the...
868
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
Regulation of Expression Occurs at Multiple Steps
22.4K
Gene expression can be regulated at almost every step from gene to protein. Transcription is the step that is most commonly regulated. This involves the binding of proteins to short regulatory sequences on the DNA. This association can either promote or inhibit the transcription of a gene associated with the respective sequence.
Transcription results in the generation of precursor (pre-mRNA) that consists of both exons and introns, which needs further processing before being translated to a...
Transcription results in the generation of precursor (pre-mRNA) that consists of both exons and introns, which needs further processing before being translated to a...
22.4K
Genetic Lingo
100.8K
Overview
100.8K
Types of Errors: Detection and Minimization
1.4K
Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
1.4K
RNA Editing
8.9K
RNA editing is a post-transcriptional modification where a precursor mRNA (pre-mRNA) nucleotide sequence is changed by base insertion, deletion, or modification. The extent of RNA editing varies from a few hundred bases, in mitochondrial DNA of trypanosomes, to a just single base, in nuclear genes of mammals. Even a single base change in the pre-mRNA can convert a codon for one amino acid into the codon for another amino acid or a stop codon. This type of re-coding can significantly affect the...
8.9K


