Jove
Visualize
お問い合わせ

関連する概念動画

Comparing Experimental Results: Student's t-Test01:09

Comparing Experimental Results: Student's t-Test

2.2K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
2.2K
Binet's Contribution to Measures of Intelligence01:23

Binet's Contribution to Measures of Intelligence

1.4K
Alfred Binet, along with his student Théophile Simon, was tasked by the French Ministry of Education in 1904 to create a method for identifying students who struggled to learn through conventional classroom instruction. This initiative aimed to address overcrowding by placing such students in specialized schools. Binet and Simon developed an intelligence test comprising 30 tasks, ranging from simple commands, like touching one's nose or ear, to more complex tasks, such as drawing...
1.4K
Microsoft Excel: Student's t-Test01:25

Microsoft Excel: Student's t-Test

674
Student's t-test in Microsoft Excel is a statistical method used to compare the means of two groups to determine if they are significantly different from each other. It's commonly used to evaluate hypotheses, such as testing whether a treatment has an effect compared to a control group. Excel provides built-in functions to perform t-tests, making it accessible for users needing to conduct basic statistical analysis.
To conduct a t-test in Excel, use the T.TEST function or the "Data...
674
Introduction to Test of Independence01:21

Introduction to Test of Independence

2.5K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.5K
Reliability and Validity01:29

Reliability and Validity

13.2K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.2K
Test Cross01:39

Test Cross

42.4K
Alleles are different forms of the same gene. Humans and other diploid organisms inherit two alleles of every gene, one from each parent.
42.4K

こちらも読む

関連記事

共著者、ジャーナル、引用グラフによってこの研究に関連する記事。

並び替え
Same author

Words vs. worlds.

Science (New York, N.Y.)·2026
Same author

Quantum 'thermometer' takes temperatures inside living cancer cells.

Nature·2026
Same author

Chats with sycophantic AI make you less kind to others.

Nature·2026
Same author

Is the journal legitimate? Aletheia-Probe can help you decide.

Nature·2026
Same author

Eureka Cam: Movements reveal moment of mathematical discovery.

Scientific American·2025
Same author

Google AI aims to make best-in-class scientific software even better.

Nature·2025
JoVE
x logofacebook logolinkedin logoyoutube logo
JoVEについて
概要リーダーシップブログJoVEヘルプセンター
著者向け
出版プロセス編集委員会範囲と方針査読よくある質問投稿
図書館員向け
推薦の声購読アクセスリソース図書館諮問委員会よくある質問
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experimentsアーカイブ
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教員リソースセンター教員サイト
利用規約
プライバシーポリシー
ポリシー

関連する実験動画

Updated: Sep 24, 2025

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

924

試しに教えられた

Matthew Hutson1

  • 1Matthew Hutson is a journalist in New York City.

Science (New York, N.Y.)
|May 10, 2022
PubMed
まとめ

人工知能 (AI) のソフトウェアは 複雑なIQテストの質問に優れているが 単純な推論のタスクでは失敗する. AIのベンチマークの改善は 人工知能の進歩に不可欠です

科学分野:

  • 人工知能
  • 認知科学
  • サイコメトリック

背景:

  • 現在の人工知能システムは 特定の領域で高度な能力を発揮しており しばしば人間の能力を超えています
  • しかし,AIはしばしば脆さを示し,常識や一般的な推論を必要とするタスクでは予期せぬ失敗をします.
  • 標準化された知能テストは IQテストのように AIの進歩を比較するためにますます使用されています

研究 の 目的:

  • 総合的なサイコメトリック知能テストで高度なAIモデルのパフォーマンスを評価する.
  • 知覚的なタスクにおけるAIの特定の弱点と失敗モードを特定する.
  • 一般的な人工知能 (AGI) の評価のための強化されたベンチマークの有用性を探求する.

主な方法:

  • 大規模な言語モデル (LLM) を含む最先端のAIモデルを使用して,確立されたIQテストバッテリーの質問に答えました.
  • 言語的推論,抽象的思考,空間的視覚化といった 異なる認知領域における AI のパフォーマンスを分析した.
  • AIのエラーを分類して 失敗の性質を理解します

主要な成果:

  • AIモデルは多くのIQテストで高得点を達成し 洗練されたパターン認識と知識リコールを実証しました

さらに関連する動画

Mixed Reality for Education MRE Implementation and Results in Online Classes for Engineering
04:12

Mixed Reality for Education MRE Implementation and Results in Online Classes for Engineering

Published on: June 23, 2023

759
Making Sense of Listening: The IMAP Test Battery
11:25

Making Sense of Listening: The IMAP Test Battery

Published on: October 11, 2010

15.9K

関連する実験動画

Last Updated: Sep 24, 2025

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

924
Mixed Reality for Education MRE Implementation and Results in Online Classes for Engineering
04:12

Mixed Reality for Education MRE Implementation and Results in Online Classes for Engineering

Published on: June 23, 2023

759
Making Sense of Listening: The IMAP Test Battery
11:25

Making Sense of Listening: The IMAP Test Battery

Published on: October 11, 2010

15.9K
  • 基本的な常識,因果的な推論,暗黙の文脈の理解を必要とするタスクでは,重大な失敗が観察されました.
  • エラー分析は,AIが単純に思える問題で 論理的でない,または無意味な間違いをする傾向を示した.
  • 結論:

    • AIは特定の認知のタスクに 期待を示していますが 現在のモデルには 堅実な一般知能と常識的な推論が欠けています
    • 既存のベンチマークは,知識集約的なタスクに焦点を当てているため,AIの能力を過大評価する可能性があります.
    • より微妙で挑戦的なベンチマークの開発は,AGIに向けた将来のAI研究を導くために不可欠です.