ドイツ憲法裁判所の判決文から解釈方法を分類
Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court
この論文をやさしく読む
ひとことで言うと
ドイツ連邦憲法裁判所の判決文を使い、LLMが文ごとの法解釈方法を分類できるか測る基準を作った。
何に役立つ?
司法文書の分析モデルやプロンプトを比較評価するために役立つ。
この研究の面白いところ
七つの分類課題で、専門家作成のプロンプトと自動最適化したプロンプトを比べている。
どこまで分かった?
結果は対象の判決、四つのモデル、試験したプロンプト構成での評価である。別の法体系への一般化は要旨で検証されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
司法の論証を大規模言語モデル(LLM)で分析することは、依然として難しい。本論文は、サヴィニーの伝統におけるラーレンツの解釈論に基づく解釈方法を、LLMが分類できるか評価するための文単位のベンチマークを提供する。貢献は三つある。第一に、この解釈の考え方を分類基準として操作可能な形にする。第二に、ドイツ連邦憲法裁判所の判決を文単位で注釈したデータセットを提供する。第三に、専門家が手書きしたプロンプトの下で、三つのモデル系列に属する四つのLLMの基準評価を報告し、Genetic-Pareto(GEPA)で最適化したプロンプトとの比較を行う。 七つの二値分類課題にわたる平均F1値は、モデル間で70.4~79.2に集中した。一般に、文法的解釈は最も見分けやすく、体系的解釈は最も難しかった。試験した構成では、GEPAで最適化したプロンプトは手書きのものを一貫して上回らず、専門家のプロンプトが有意義な基準であることを示唆する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny. Our contributions are threefold. First, we operationalize this conception of interpretation as classification criteria. Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level. Third, we report baseline evaluations of four LLMs from three model families under expert hand-written prompts, compared against prompts optimized with Genetic-Pareto (GEPA). Mean F1 over the seven binary subtasks clusters between 70.4 and 79.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest; under the tested configuration, GEPA-optimized prompts do not systematically outperform the hand-written ones, suggesting that the expert prompts provide a meaningful baseline.
著者のコメント
accepted at the ICML 2026 AI4Law Workshop; 32 pages (main text 9 pages + appendices 23 pages)
arXiv ID: 2609.26945 / 要約の誤りについて