arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

カナダ法の条文を根拠に議論するAIを評価するGRACE

GRACE: Grounded Adversarial Reasoning over Canadian Law

Jiakang Xu, Wantong Huo, Udom Silparcha and Jonathan H. Chan

この論文をやさしく読む

ひとことで言うと

カナダの法律の条文を読ませ、立場を擁護したり、不足する情報を扱ったり、複数の規定を組み合わせたりするAIを育てて評価するデータセットです。

何に役立つ?

条文を入力として使う小規模な法律向けモデルの研究に利用できます。実験では、条文を渡すことが根拠を伴う回答の改善に重要でした。

この研究の面白いところ

単純な法律知識の正誤だけでなく、対立する立場の議論や不確実な状況を評価対象にしています。条文を渡す場合と渡さない場合を比べ、資料への依存も可視化しています。

どこまで分かった?

概念実証の評価には教師出力との一致が含まれ、これは独立した専門家による法的正確性の確認とは同一ではありません。条文を与えないと根拠への結び付きが大きく低下し、実際の法律業務での有効性は要旨に示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルはさまざまな法律タスクで高い性能を示しているが、既存のベンチマークは、法的な立場を選んで擁護する能力、不完全な情報の下で推論する能力、複数の法規定を統合する能力をほとんど評価していない。法律分野の自然言語処理で十分に扱われてこなかったカナダ法では、この不足が特に顕著である。 本研究では、カナダの連邦法に基づく1,915件の質問・推論・回答からなるデータセットGRACE(Grounded Reasoning Adversarial Canadian LEgal examples)を導入する。GRACEは、対立する立場を擁護する議論、不確実性、適用に関する推論という3つの推論様式を扱う。生の条文テキストを分割し、事例に基づく質問と推論を生成し、モデルを用いない引用検証とLLMによる品質監査を通じて例を選別する処理系を開発した。 概念実証として、根拠に基づく法的推論のための軽量モデルCLeAR-4B(Canadian Legal Adversarial Reasoning)を微調整し、資料を参照できる設定とできない設定の両方で、変更を加えていない基盤モデルQwen3-4Bと比較評価する。関連する法律の条文を提供した場合、CLeAR-4Bは教師出力との一致と条文引用の振る舞いを大幅に改善する一方、条文を与えない場合には根拠への結び付きが急激に低下する。これらの結果は、GRACEが、提供された条文からより効果的に推論する軽量な法律モデルの開発を支援できることを示唆する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models have shown strong performance across a range of legal tasks, but existing benchmarks rarely evaluate the ability to take and defend a legal position, reason under incomplete information, or synthesize multiple statutory provisions. This gap is particularly pronounced for Canadian law, which remains underrepresented in legal NLP. We introduce GRACE (Grounded Reasoning Adversarial Canadian LEgal examples), a dataset of 1,915 question-reasoning-answer instances grounded in Canadian federal legislation. GRACE covers three reasoning modes: adversarial advocacy, uncertainty, and applied reasoning. We develop a pipeline that partitions raw statutory text, generates scenario-based questions and reasoning, and filters examples through model-free citation verification and LLM-based quality auditing. As a proof of concept, we fine-tune CLeAR-4B (Canadian Legal Adversarial Reasoning), a lightweight model for grounded legal reasoning, and evaluate it against the unmodified Qwen3-4B base model in open- and closed-book settings. CLeAR-4B substantially improves agreement with teacher outputs and statutory citation behavior when the relevant act text is provided, while its grounding degrades sharply when the statute is withheld. These results suggest that GRACE can support the development of lightweight legal models that reason more effectively from supplied statutory text.

arXiv ID: 2609.23726 / 要約の誤りについて