同じ平均点でも契約文書の個別判断は変わる
Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding
この論文をやさしく読む
ひとことで言うと
契約文書AIの平均正解率だけでは、個別判断の安定した正しさが分からないと示した。
何に役立つ?
考えられる用途は、契約書を扱うモデルを費用、速度、個別判断の正しさで評価することである。
この研究の面白いところ
同じ契約と判断を保ち、問いの見せ方や出力順序だけを変えて比較する。
どこまで分かった?
ContractNLIでの比較であり、条件間の小差から一般的な安定性の優位は主張できないと明記する。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
契約書の推論では一つの文書について複数の判断が必要だが、平均正解率は個々の判断の変化を隠すことがある。繰り返し同じ答えを出すだけでも不十分で、同じ誤答を返し続ける場合がある。本論文はContractNLIでJevと九つの言語モデルを比べ、推論費用、応答時間、平均的な正しさ、依頼条件を変えて繰り返したときの正しさを評価する。制御した比較では契約書と対象の判断を固定したまま、仮説の見せ方、求める出力、出力順序を変える。Jevは評価した構成の中で費用と応答時間の中央値が最も低かったが、外部提供の言語モデルは基準条件での正解率が高かった。基準の正解率による順位は、すべての条件と反復で正解を保つ度合いによる順位とは異なった。ただし後者の小さな差から、一般的な安定性の優位は証明できない。開発段階の診断では、改善と悪化が打ち消し合う例や持続する誤りも見つかった。これらは、費用と応答時間に加えて、依頼の設定が変わっても個々の判断が正しいままかを評価する必要性を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating inference cost, response time, average correctness, and correctness across repeated request conditions. Controlled comparisons vary hypothesis visibility, requested outputs, and output order while keeping the contract and target judgment fixed. Jev has the lowest cost and median response time among the evaluated configurations, while hosted language models achieve higher baseline accuracy. Rankings by baseline accuracy differ from rankings by correctness across every condition and repeat, although small differences in the latter do not establish a general stability advantage. Development diagnostics further reveal compensating corrections and regressions, as well as persistent errors. These findings motivate evaluating cost and response time alongside whether individual judgments remain correct as the request configuration changes. Code: https://github.com/ZF-Utokyo/Jev-Benchmark
arXiv ID: 2609.27678 / 要約の誤りについて