指定された否定の意味を言語モデルは守れるか
Not What You Meant: Can LLMs Follow a Specified Negation Semantics?
この論文をやさしく読む
ひとことで言うと
同じ「否定」でも決めた解釈に従って答えが変わる問題で、言語モデルが指定された規則を守れるか調べた。
何に役立つ?
法令や医療など、否定の読み方を厳密に指定する必要がある推論システムの評価に役立つ。実務での安全性が確認されたという主張ではない。
この研究の面白いところ
正解を解法器で検証した上で、論理は同じまま文章の順序や言い方を変え、モデルが影響を受けるか測っている。
どこまで分かった?
報告された正答率はNAFBenchの評価設定に対するものである。要旨には、実際の法務・医療判断での評価結果はない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
否定の解釈は分野をまたいで一様ではない。法的、規制上、医療上の推論では、開世界か閉世界か、二値か三値か、あるいは許容的か懐疑的かという、採用する読み方によって意図する解釈が変わる。本研究は、大規模言語モデルが既定でどの否定解釈を選び、別の読み方を明示されたときにその傾向を切り替えられるかを調べる。そのために、SLDNF、整礎意味論、安定モデル意味論における許容的推論と懐疑的推論という四つの意味論的観点を含む、解法器で正解を確かめた問題を手続き的に生成するNAFBenchを導入する。生成器は、深さ、幅、循環構造を制御した基礎正規論理プログラムを出力する。各プログラムをSWI-Prolog、整礎意味論の解法器、clingoを使って四つの観点全てで解き、観点ごとに異なり得る最大四つのラベルを得る。その後、答えを変えない複数の表現や規則の順序でプログラムを自然言語に直す。 結果は一貫した隔たりを示した。オープンソースモデルでは、指定された否定の意味論に従う課題はまだ解決されておらず、最も強いモデルの正答率は四つの観点で59~74%、最も弱いモデルは31~67%だった。全モデルが、論理的には同じ規則の並べ替えの半数以上で順序の影響を受け、弱い方の二つのモデルは整礎意味論での「未定義」をしばしば過度に断定した。最先端モデルの二つは主となる固定複雑度の評価集合で100%に達し、三つ目のo4-miniもほぼ完全だったが、整礎意味論の「未定義」では81%に下がった。推論を解法器に委ねること、正解を検証した推論過程で追加学習すること、明示的な三値の判定を求めることは、それぞれ隔たりを部分的に縮めた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Negation does not carry a uniform interpretation across domains. In legal, regulatory, and medical reasoning, the intended interpretation depends on the reading in force -- open- versus closed-world, two- versus three-valued, and credulous versus skeptical. We study which reading of negation large language models adopt by default and whether they can override that preference when a different reading is explicitly specified. To this end, we introduce NAFBench, a procedural generator of solver-certified instances spanning four semantic viewpoints: SLDNF, well-founded semantics (WFS), and credulous and skeptical reasoning under stable-model semantics. The generator emits ground normal logic programs with controlled depth, width, and cycle structure. Each program is solved under all four viewpoints using SWI-Prolog, a well-founded semantics solver, and clingo, yielding up to four divergent labels. The programs are then verbalized into natural language under multiple framings and rule orderings that leave the answer invariant. The results expose a consistent gap. Across open-source models, following a specified negation semantics remains unsolved: the strongest models score 59--74% across the four semantic viewpoints, while the weakest score 31--67%. All models are order-sensitive on more than half of logically identical rule shufflings, while the two weaker models frequently overcommit on well-founded "undefined." Two frontier models reach 100% on the main fixed-complexity evaluation set, and a third, o4-mini, is near-perfect, falling only to 81% on well-founded "undefined." Delegating reasoning to a solver, fine-tuning on certified traces, or forcing an explicit three-valued verdict each partly closes the gap.
arXiv ID: 2609.27517 / 要約の誤りについて