Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Observational Learning01:12

Observational Learning

312
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
312
Reinforcement Schedules01:24

Reinforcement Schedules

242
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
242
Improving Translational Accuracy02:07

Improving Translational Accuracy

11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Reinforcement01:23

Reinforcement

341
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
341
Generalization, Discrimination, and Extinction01:24

Generalization, Discrimination, and Extinction

788
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
788
Language Development01:22

Language Development

449
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
449

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Overlayer-Engineered BiVO<sub>4</sub> Suppresses H<sub>2</sub>O<sub>2</sub> Decomposition to Enable Sustained Photocatalytic Production.

Angewandte Chemie (International ed. in English)·2026
Same author

Crystal Phase Engineering Accelerates Hydrogen Reverse Spillover for Efficient Alkaline Hydrogen Production.

Nano-micro letters·2026
Same author

Systemic LPS exposure suppresses pulsatile LH secretion and impairs hepatic function in goats, accompanied by immune activation and stress responses.

The Journal of reproduction and development·2026
Same author

Plasmonic Re-Excitation Enables Superoxide-Mediated Ethane Conversion to Acetic Acid under Visible Light.

Journal of the American Chemical Society·2026
Same author

Self-supplying hydrogen peroxide-driven gold nanoparticle aggregation via bio-orthogonal click reaction activates cGAS-STING-PERK pathway for multimodal tumor therapy.

Journal of nanobiotechnology·2026
Same author

Nonlinear hydrothermal associations between coupled landscape ecological risk and resilience in a major grain-producing region of China.

Journal of environmental management·2026

相关实验视频

Updated: Sep 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

681

用来自大型语言模型的背景知识提高强化学习的样本效率.

Fuxiang Zhang, Junyou Li, Yi-Chen Li

    IEEE transactions on neural networks and learning systems
    |August 14, 2025
    PubMed
    概括

    本研究引入了一种使用大型语言模型 (LLM) 来提取一般环境知识的新框架,显著提高了强化学习 (RL) 任务中的样本效率.

    科学领域:

    • 人工智能的人工智能
    • 机器学习 机器学习

    背景情况:

    • 强化学习 (RL) 面临的挑战是样本效率低.
    • 目前使用大型语言模型 (LLM) 为RL指导的方法缺乏跨任务的通用性.
    • 需要可重复使用的环境知识来加速RL政策学习.

    研究的目的:

    • 开发一个利用LLM来提取环境的可概括的背景知识的框架.
    • 将这些知识作为潜在的函数来表示,以便在RL中有效地塑造奖励.
    • 通过一次性知识表示,提高各种下游RL任务的样本效率.

    主要方法:

    • 用事先收集的环境经验为LLM奠定基础.
    • 促使LLM使用代码生成,偏好注释和目标赋值等方法来界定背景知识.
    • 将提取的知识表示为潜在函数,用于基于潜力的奖励塑造.

    主要成果:

    • 在多个RL任务中显示出样本效率的显著改善.
    • 验证了框架在Minigrid和Crafter领域的有效性.
    • 展示了为各种下游任务提取的环境知识的概括性.

    结论:

    更多相关视频

    Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
    05:47

    Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

    Published on: June 13, 2025

    578
    Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
    08:05

    Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques

    Published on: June 30, 2020

    7.7K

    相关实验视频

    Last Updated: Sep 11, 2025

    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
    03:14

    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

    Published on: December 6, 2024

    681
    Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
    05:47

    Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

    Published on: June 13, 2025

    578
    Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
    08:05

    Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques

    Published on: June 30, 2020

    7.7K
  • 拟议的框架有效地利用LLM来捕获和表示一般环境知识.
  • 这种方法显著提高了强化学习中的样本效率.
  • 该方法为创造更具适应性和效率的RL代理提供了一个有希望的方向.