arXiv論文メモ
新着一覧
cs.LG / q-bio.NC · 査読状況未確認

脳波基盤モデルの比較で入力処理の影響を点検する

Matched-Input Estimates Differ in Sign Across Architectures: Auditing EEG Foundation Models on Motor Imagery

Kevin Zhou, Sparsh Roy

この論文をやさしく読む

ひとことで言うと

脳波モデルの優劣が、事前学習そのものだけでなく、入力する周波数帯や比較モデルの選び方に左右されないかを調べます。

何に役立つ?

脳波の運動想起モデルを比較するときに、入力条件を揃え、検証データだけで設定を決める評価設計の参考になります。

この研究の面白いところ

入力を同じ広帯域にしても改善するモデルと悪化するモデルがありました。また、予測の確信度の較正を直しても、正解率の低さは別に残ります。

どこまで分かった?

入力を揃えた差はn=9での多重比較補正後に有意ではなく、交互作用を証明していません。2クラス課題で差を検出できなかったことも、同等性の証明ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

事前学習済みの脳波基盤モデルは、ブレイン・コンピュータ・インターフェースの汎用エンコーダとして提案されることが増えているが、最近のベンチマークでは、表現が下流課題へ移転する条件について結果が一致していない。本研究では運動想起を対象にLaBraMとCBraModを監査する。前処理、アーキテクチャ、最適化、層を固定する深さ、チェックポイント、温度、手法選択を、訓練セッションのデータだけで決める、検証によって固定した手順を用いる。4クラスのBCI Competition IV-2aでは、ここで評価したすべての教師あり比較手法が、検証で選択した微調整を含むすべての基盤モデル設定を上回る。 次に、基盤モデルと課題専用デコーダが通常異なる入力処理で評価されるという重要な交絡要因を調べる。三つの教師ありアーキテクチャを基盤モデルが使う広帯域の配列で再学習すると、入力を揃えたときの正解率の差の符号はアーキテクチャ間で異なった。広帯域入力はATCNetの正解率を0.078高める一方、EEG Conformerでは0.088低下させた。n=9で多重比較補正を行うと、入力を揃えた三つの個別項のいずれも有意ではない。そのため、この符号の違いは、アーキテクチャと入力処理の正式な交互作用としてではなく、記述的な結果として扱う。観測された符号の違いは、一つの比較手法だけでは、事前学習モデルと教師ありモデルの性能差を、アーキテクチャによらない形で分解できない可能性を示唆する。 また、4クラスでの劣位は運動想起データセット全体で一様には再現しない。2クラスのBNCI2014-004では、微調整したCBraModと教師あり比較手法の間に、同じ隔たりを検出できない。最後に、検証データで当てはめた温度スケーリングは、4クラスの正解率が大幅に低いにもかかわらず、基盤モデルの較正誤差を教師あり手法の範囲まで戻す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Pretrained EEG foundation models are increasingly proposed as general-purpose encoders for brain-computer interfaces, yet recent benchmarks disagree about when their representations transfer to downstream tasks. We audit LaBraM and CBraMod on motor imagery under a validation-locked protocol in which preprocessing, architecture, optimization, freeze depth, checkpoint, temperature, and method selection are determined using training-session data only. On four-class BCI Competition IV-2a, every supervised comparator evaluated here outperforms every foundation-model configuration, including validation-selected fine-tuning. We then examine a key confound: foundation models and task-specific decoders are normally evaluated using different input pipelines. Retraining three supervised architectures on the broadband arrays consumed by the foundation models produces matched-input accuracy differences of opposite sign across architectures: broadband input improves ATCNet by 0.078 accuracy while reducing EEG Conformer accuracy by 0.088. None of the three individual matched-input terms is significant after multiple-comparison correction at n = 9, so we treat the sign variation descriptively rather than as a formal architecture-by-pipeline interaction. These observed sign differences suggest that a single comparator may not provide an architecture-invariant decomposition of a pretrained-versus-supervised performance gap. The four-class deficit also does not reproduce uniformly across motor-imagery datasets: on two-class BNCI2014-004 we cannot detect the same separation between fine-tuned CBraMod and the supervised comparators. Finally, validation-fitted temperature scaling returns foundation-model calibration error to the supervised range despite substantially lower four-class accuracy.

arXiv ID: 2609.23924 / 要約の誤りについて