arXiv論文メモ
新着一覧
cs.SE / cs.AI / cs.CL · 査読状況未確認

言語モデルの行動仕様にある矛盾を仕様文から検出する

Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning

Zichen Xie, Mrigank Pawagi, Lize Shao, Yang Hu, Wenxi Wang

この論文をやさしく読む

ひとことで言うと

AIへの行動指示を読み比べ、同じ状況で同時に守れないルールの組み合わせを探す方法です。

何に役立つ?

モデルの応答テストとは別に、仕様文の段階で矛盾候補を見つける用途に役立ちます。報告された評価では、候補を人が確認して5件の不整合を特定しています。

この研究の面白いところ

仕様を別の形式言語へ全面的に置き換えず、自然言語のままLLMに検証させます。権限レベルと話題を使って比較するルールを整理する点が特徴です。

どこまで分かった?

適合率38.5%という結果から、検出候補のすべてが確認済みの不整合というわけではありません。要旨が示す具体的適用先はOpenAI Model Specです。開発者による議論開始は、修正完了を意味しません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

モデル仕様は、大規模言語モデル(LLM)がどう振る舞うべきかを定め、アラインメント学習、推論時の振る舞い、評価の指針となる。しかし、仕様自体に欠陥が含まれることもある。個別には妥当な2つの原則が、同じ状況に適用されると両立しない振る舞いを要求し、両方を満たす応答が存在しなくなる場合がある。このような不整合の検出は難しい。自然言語の仕様を形式化すると微妙な区別が失われるおそれがあり、振る舞いに基づくテストでは、仕様の欠陥とモデルの振る舞いの違いを確実には区別できない。 本研究では、仕様文そのものを監査してモデル仕様の不整合を直接検出する初の手法、VeriSpecを提案する。中心となる着想は、仕様を自然言語のまま保持しつつ、LLMを検証器として使うことである。VeriSpecは文脈を考慮した構造化ルールを抽出し、同じ権限レベルにあって振る舞いに関係するルールを話題に基づくグラフでまとめ、検証器としてのLLMの推論によって不整合を検出する。OpenAI Model Specへの適用では405個のルールを抽出し、5件の不整合を手作業で確認した。すべて開発者へ報告しており、開発者は肯定的に応答し、内部での議論を開始した。 5つの比較手法に対し、VeriSpecは確認済み不整合を最も多く検出し、適合率は最高の38.5%、確認済み不整合1件当たりの費用は最少の11.12ドルだった。これらの結果は、モデルに影響を及ぼす前に欠陥を元から発見する直接的な仕様監査が、振る舞いに基づくアラインメント評価を補う実用的な手段となることを示している。コードは https://github.com/HIPREL-Group/VeriSpec で公開している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at https://github.com/HIPREL-Group/VeriSpec.

arXiv ID: 2610.01847 / 要約の誤りについて