arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

時系列予測の正解率と説明の忠実さを分けて測る

TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction

Jie Gong, Maowei Jiang, Zhiwei Liu, Yankai Chen, Guojun Xiong, Xue Liu, Min Peng, Qianqian Xie, and Sophia Ananiadou

この論文をやさしく読む

ひとことで言うと

数値と文章を一緒に読んで予測するモデルが、本当に両方を使っているか、説明がその判断を正しく表しているかを確かめるテスト集です。

何に役立つ?

正解率が高くても説明をそのまま信用できない場面を見つけるのに役立ちます。金融と交通の入力を制御して、単なる正解と証拠に基づく判断を区別します。

この研究の面白いところ

説明で時間的要因に触れる割合が90%超でも、実際の予測行動による裏付けは22%未満という隔たりを測っています。言葉で関係を説明できることと予測に使えることを分離しています。

どこまで分かった?

評価対象は二分野の4,856件と10モデルです。HPCの数値はペア単位の正解率であり、通常の予測正解率ではありません。データなどは公開予定とされ、要旨だけでは公開済みとは確認できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)は、数値の時系列履歴と文章で記述した出来事から予測を行うために、ますます使われている。しかし正解率だけでは、正しい答えが二つの入力の有効な統合によるものなのか、それとも出来事の正負の傾向、単一モダリティの事前知識、表面的な手掛かりによるものなのかを見分けられない。同様に、もっともらしい説明も、モデルの振る舞いを駆動する証拠を忠実に反映せず、予測を後付けで正当化している可能性がある。 出来事を条件とする時系列予測における、モダリティをまたぐ理解と説明の忠実性を診断するベンチマーク、TimeLitmusを導入する。TimeLitmusは金融と交通の二分野にわたる4,856件の評価レコードを含み、自然な予測課題に、制御された反実仮想・対照介入、説明を対象とする忠実性テスト、系統的な近道利用の統制を組み合わせる。 代表的な10種類のLLMを評価すると、通常の予測正解率は、信頼できるモダリティ横断理解を大幅に過大評価していた。難しいペア対照課題(HPC)でペアの両方に正解する割合は、最高でも金融で19.2%、交通で11.7%にとどまり、10モデルすべてが金融の時系列側を操作する統制課題で期待を下回る一貫性を示した。モデルは場面間の関係を明示的には認識していても、個別に予測するときにはその関係を適用できないことが多い。 説明の忠実性にも同様の隔たりがある。交通では、多くのモデルが90%を超える事例で操作した時間的要因に言及する一方、振る舞いによる裏付けは22%未満にとどまる。人間の評価者は、条件をそろえた統制課題や難しいペア課題でLLMを上回り、これらの区別が入力から読み取れることを確認した。自然な課題だけを用いた追加適応は、証拠の選択や入力への感度を選択的に改善するが、統制課題や難しいペア課題で一貫した改善はもたらさない。ベンチマーク、評価スイート、教師あり適応用データを一般公開する予定である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cannot reveal whether correct answers reflect effective integration of the two inputs or instead arise from event polarity, unimodal priors, or superficial cues. Likewise, plausible explanations may rationalize predictions without faithfully reflecting the evidence that drives model behavior. We introduce TimeLitmus, a diagnostic benchmark for cross-modal understanding and explanation faithfulness in event-conditioned time-series prediction. TimeLitmus contains 4,856 evaluation records across Finance and Traffic, combining natural prediction with controlled counterfactual and contrastive interventions, explanation-targeted faithfulness tests, and systematic shortcut controls. Across ten representative LLMs, standard prediction accuracy substantially overstates reliable cross-modal understanding: Hard Paired Contrast (HPC) pair correctness peaks at only 19.2% in Finance and 11.7% in Traffic, and all ten models show lower-than-expected consistency on Finance series-side controls. Models often recognize scenario relations explicitly yet fail to apply them during independent prediction. Explanation faithfulness shows a similar gap: in Traffic, most models cite the manipulated temporal factor in over 90% of cases, while behavioral support remains below 22%. Human annotators outperform LLMs on matched controlled and hard-pair diagnostics, confirming that these distinctions are recoverable from the inputs. Natural-only adaptation yields selective gains in evidence selection and input sensitivity, but not consistent gains in controlled or hard-pair behavior. The benchmark, evaluation suite, and supervised adaptation data will be released publicly.

arXiv ID: 2609.24677 / 要約の誤りについて