人のAIへの信頼と根拠のずれを検証するCRiDiT
CRiDiT: Instantiating a run-time testbed for trust calibration in AI-infused systems
この論文をやさしく読む
ひとことで言うと
AIを人がどれくらい信じるかと、AIを信じる根拠がどれくらいあるかを比較し、ずれを検出して対処する仕組みを実装・検証しています。
何に役立つ?
信頼度を1つの点数にまとめて差を取る設計で、必要な情報が失われていないかを点検する参考になります。信頼のずれを見つける処理と、その後の対応を別々に検討できます。
この研究の面白いところ
設計どおりに機能したという報告だけでなく、リスク別のしきい値が期待した作動差を生まないなど、実装して初めて見えた問題を要件に戻しています。過信への説明と訂正の違いもログから検討しています。
どこまで分かった?
評価は3シナリオ、15セッション、144ステップで、訂正がずれを縮めた観察は6例です。実際の採用・金融・法律業務でリスクが減ったことを示す大規模な運用評価とは区別する必要があります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
AIが大規模な技術基盤に組み込まれるにつれ、人間の信頼をシステムの信頼に値する程度に合わせる「信頼較正」は、工学上の重要な課題となっている。信頼が過剰でも不足していても、運用や安全上のリスクにつながるためである。概念的な枠組みは信頼較正を理解する強い基盤を提供するが、動作するシステムへの実装は依然として難しい。人間の信頼の入力、機械の信頼性の根拠、ずれの検出と是正が、閉じたループの中で一体として動作するテスト環境が少ないためである。 本論文では、CRiDiT(Computational Risk-Sensitive biDirectional Trust)を実行時のテスト環境として実装する。機械側の信頼はDempster–Shafer理論とPCR5による再配分、人間側の信頼はSubjective Logicを使って扱い、較正にはしきい値に基づく信頼のずれを用いる。Design Science Researchの方法論に従い、採用、金融、法律という3つの重大な影響を持つシナリオで試作物を動かし、15セッションにわたる144の対話ステップを記録した。 解析の結果、この試作物は意図したとおり信頼較正の動きを捉える一方、実装された方策が設計要件から外れる点が3つ明らかになった。機械側の推定はタスクに関係する根拠ではなく全体的なベンチマークから始まること、リスクに応じたしきい値がリスクに応じた作動を生まないこと、そして較正方策が過信に対して説明を促すプロンプトを割り当てている一方、観察された6例すべてで訂正がずれを縮めたことである。最初の2点は、推定された2つのスカラーの差を較正基準にするという同じ設計判断に由来し、根拠から導出したスカラーではなく根拠そのものに基準を作用させるべきだという、共通の要件を示す。3点目は検出後の対応に関するもので、信頼の修復から引き継いだ行動の選択肢が、対話ログで有効と示された対応と一致しないことを示す。本研究は、試作物、その実行時の振る舞いの特徴付け、そしてこの特徴付けから導かれる要件を提供する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The integration of AI into larger technical infrastructures has made the alignment of human trust with system trustworthiness, known as trust calibration, a critical engineering concern, since misplaced trust in either direction leads to operational and safety risks. While conceptual frameworks provide a strong foundation for understanding trust calibration, their translation into running systems remains a challenge, because there are few testbeds in which human trust inputs, machine trustworthiness evidence, gap detection and remediation operate together within a closed loop. This paper instantiates CRiDiT (Computational Risk-Sensitive biDirectional Trust) as a run-time testbed, operationalising machine-side trust with Dempster-Shafer Theory and PCR5 redistribution, human-side trust with Subjective Logic, and calibration with a threshold-based trust gap. Following the Design Science Research methodology, we exercise the artifact across three high-stakes scenarios (hiring, financial, legal), producing 144 logged interaction steps across fifteen sessions. The analysis shows that the artifact captures trust calibration dynamics as intended, and reveals three points at which the instantiated policy departs from its design requirements: the machine-side estimate begins from a global benchmark rather than task-relevant evidence; risk-sensitive thresholds do not produce risk-sensitive triggering; and the calibration policy assigns explanatory prompts to over-trust, where corrections narrowed the gap in all 6 observed cases. Since the first two arise from the same design decision, to make the difference of two estimated scalars the calibration criterion, they point toward a common requirement: that the criterion should operate on the evidence rather than on scalars derived from it. The third concerns what follows detection, and shows that the action vocabulary inherited from trust repair does not align with what the interaction logs show to be effective. The work contributes the artifact, a characterisation of its run-time behaviour, and the requirements this characterisation elicits.
arXiv ID: 2609.24833 / 要約の誤りについて