Jove
Visualize
お問い合わせ
JoVE
x logofacebook logolinkedin logoyoutube logo
JoVEについて
概要リーダーシップブログJoVEヘルプセンター
著者向け
出版プロセス編集委員会範囲と方針査読よくある質問投稿
図書館員向け
推薦の声購読アクセスリソース図書館諮問委員会よくある質問
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experimentsアーカイブ
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教員リソースセンター教員サイト
利用規約
プライバシーポリシー
ポリシー

関連する概念動画

Improving Translational Accuracy02:07

Improving Translational Accuracy

15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.7K
3.7K

こちらも読む

関連記事

共著者、ジャーナル、引用グラフによってこの研究に関連する記事。

並び替え
Same author

SmartAlert - Implementing Machine Learning-Driven Clinical Decision Support for Inpatient Laboratory Utilization Reduction.

NEJM AI·2026
Same author

Improving Generalizability in Whole-Cell Antibiotic Discovery Through Active Learning.

bioRxiv : the preprint server for biology·2026
Same author

Why and How to Monitor Deployed AI Systems in Health Care.

NEJM catalyst innovations in care delivery·2026
Same author

BRIDGE: benchmarking large language models for understanding real-world clinical practice texts.

Nature biomedical engineering·2026
Same author

A Consensus Approach to the Incorporation of Total Neoadjuvant Therapy in a Treatment Algorithm for Stage I-III Resectable Rectal Cancer.

Current oncology (Toronto, Ont.)·2026
Same author

Micro-randomization trial design under operational constraints.

Contemporary clinical trials·2026

関連する実験動画

Updated: Feb 28, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K

MedAgentBench v2: 臨床EHRタスクにおけるLLMエージェント設計の改善

Eric Chen1, Sam Postelnik2, Kameron Black3

  • 1MIT, USA, eric25@alum.mit.edu.

Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
|February 27, 2026
PubMed
まとめ

本研究では、臨床電子カルテ(EHR)タスクにおける大規模言語モデル(LLM)エージェントの評価ベンチマークであるMedAgentBenchを紹介する。メモリを用いた改良により、成功率は98%に達し、AIの可能性を示す。

キーワード:
大規模言語モデルLLMエージェント臨床タスク電子カルテベンチマークAIヘルスケアインフォマティクス

さらに関連する動画

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.7K

関連する実験動画

Last Updated: Feb 28, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.7K

科学分野:

  • ヘルスケアにおける人工知能
  • 臨床インフォマティクス
  • 自然言語処理(NLP)

背景:

  • 臨床現場における大規模言語モデル(LLM)エージェントの評価は困難である。
  • 既存のベンチマークにはFHIR準拠の電子カルテ(EHR)統合が欠けている。
  • 以前のLLMエージェントは、複雑な臨床タスクやEHRインタラクションに苦労していた。

研究 の 目的:

  • FHIR準拠のEHR内での臨床タスクにおけるLLMエージェント初のベンチマークであるMedAgentBenchを導入する。
  • LLMエージェントのためのプロンプトエンジニアリング、ツール設計、およびメモリコンポーネントの改善を提示する。
  • エージェントの一般化と実世界での適用性を評価するための新しい臨床主導型タスクを開発する。

主な方法:

  • FHIR準拠のEHRを備えたMedAgentBenchを開発した。
  • 連鎖思考推論や少数例学習を含むプロンプトエンジニアリング技術を実装した。
  • EHRインタラクション、出力フォーマット、計算のための強化されたツール、およびメモリコンポーネントを導入した。

主要な成果:

  • GPT-4.1を使用し、メモリなしで91.0%、メモリありで98.0%の成功率を達成した。
  • メモリエントリーのないタスクでも性能が向上し、適応性を示唆した。
  • 医師と協力して、包括的な評価のために300の新しい多段階臨床タスクを開発した。

結論:

  • LLMエージェント設計における大幅な改善により、臨床EHRタスクのパフォーマンスが向上した。
  • メモリコンポーネントとプロンプトエンジニアリングは、高い成功率のために不可欠である。
  • 医療における責任あるAI展開のためには、EHRエージェントとベンチマークのさらなる開発が必要である。