エコネット++: 複数の言語でサッカー試合のオーディオコメントを集めたデータセット
Fahad Majeed1, Maria Nazir2, Marco Agus3
1College of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar. fama44316@hbku.edu.qa.
Scientific reports
|February 17, 2026
まとめ
この研究は,サッカー放送のためのオーディオ分析パイプラインを導入し,スピーチ認識と翻訳を改善します. このシステムは,プロのスポーツメディア向けに,スケーラブルで言語不可知性のオーディオ理解を提供します.
科学分野:
- * マルチモダル信号処理
- * コンピュータ言語学
- * スポーツアナリスト
背景:
- *プロのサッカー放送は,複数の重複するスピーチ・アンド・ノンスピーチ・シグナルを持つ複雑なオーディオ環境を提示します.
- * 既存の音声分析方法は,多言語のコンテンツと多様な音響条件で苦労することが多い.
- * 自動音声認識 (ASR) と翻訳は,コンテンツ分析に不可欠ですが,堅牢なパイプラインが必要です.
研究 の 目的:
- * プロのサッカー試合の放送のための包括的で再現可能なオーディオ分析パイプラインの開発と評価.
- * パイプラインの個々のコンポーネント (事前処理,セグメンテーション,トランスクリプション,翻訳) が全体的なパフォーマンスに与える影響を体系的に分析する.
- * 放送コンテンツのスケーラブルで,言語アグノスティックなオーディオ理解を可能にする.
主な方法:
- *オーディオ抽出,デノイシング (Demucs),スピーチセグメンテーション (Silero VAD),コメンテーター/視聴者ストリーム分類のための統一パイプライン.
- * 音声信号の隔離のための周波数域変換 (FFT) と帯域域フィルタリング (300 Hz-7 kHz).
- *様々な自動音声認識 (ASR) モデル (Whisperのバリエーション,Insane Fast Whisper) を使った多言語トランスクリプションと英語翻訳.
主要な成果:
- * 単語の誤差率が低く,完全なマッチのビデオで様々な音響条件で正確な言語検出を達成しました.
- *一貫したスピーチセグメンテーションと,解説者/観客のオーディオストリームの信頼性の高い分類が実証されています.
- * 複数の言語のオーディオを処理し,トランスクリプト,翻訳,メタデータを含む構造化されたJSON出力を生成しました.
結論:
- *開発されたパイプラインは,スポーツ放送における自動オーディオ処理のための堅牢な枠組みを提供します.
- * 体系的なコンポーネント分析は,現実世界のシナリオでASRと翻訳パフォーマンスを最適化するための洞察を提供します.
- *データセットとパイプラインの公開は,再現可能な研究と要約と分析のような下流アプリケーションを容易にする.
関連する概念動画
Air-entraining Agents
297
Air-entraining agents improve the durability and workability of concrete in climates with frequent freezing and thawing. These agents prevent cracks by introducing small air bubbles into the mix, creating spaces accommodating water expansion when temperatures drop. The air-entraining agents lower the surface tension of water, forming stable, small air bubbles. This method is more effective than having accidental large voids, as the intentional, smaller, and evenly distributed air voids improve...
297
Echo
1.0K
The human ear cannot distinguish between two sources of sound if they happen to reach within a specific time interval, typically 0.1 seconds apart. More than this, and they are perceived as separate sources.
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
1.0K
RACE - Rapid Amplification of cDNA Ends
7.3K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
7.3K


