法律文書の階層を使って多言語で根拠を追える要約を作る
LexLattice: Multilingual Extractive Summarization via Neural Cellular Automata on Document Hierarchies
この論文をやさしく読む
ひとことで言うと
法令の階層と離れた箇所の根拠をまとめ、原文から抜き出して多言語で要約する方法。
何に役立つ?
法律文書で、どの原文を使ったか追える要約を作る際の手法として役立つ。
この研究の面白いところ
学習可能な部分は180万パラメータでも、24言語の評価で大きな指示学習済みモデルを上回った。
どこまで分かった?
結果はEUR-Lex-SumのROUGE評価に基づく。法律実務での正確さや人による利用評価は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
法律文書の要約では原文に忠実であることが重要であり、原文にたどれる箇所をそのまま選ぶ抽出型の方法が使われる。しかし、こうした方法は段落などの構造単位を個別に順位付けすることが多く、文書中の離れた場所に分散し、重要性を共有する根拠をまとめることには十分注意を払わない。そこでLexLatticeを導入する。これは法令の階層を二次元の意味的な格子として明示し、マスク付き二次元ニューラルセルオートマトンで情報を統合してから選択する抽出型の要約器である。 LexLatticeはEUR-Lex-Sumの24言語すべてについて、多言語と異言語間の両設定でROUGEの最高水準を達成した。凍結した多言語エンコーダーの上に学習可能な180万パラメータの統合器だけを置く構成ながら、数十億パラメータの指示学習済み比較方法を上回った。資源が豊富な言語だけで学習した統合器も、未見の言語へ性能をほぼ落とさず移せた。保持率は0.99であり、表層的な言語表現ではなく言語に依存しにくい意味的な配置で働くことを示唆する。これらの結果は、文書構造に沿った明示的な情報統合が、モデルの大規模化に代わる、小規模で根拠を追える多言語法律要約の方法になることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Faithfulness is a central concern in legal text summarization, which motivates extractive approaches that select verbatim content traceable to its source. Such methods typically rank paragraphs or other structural units in isolation, yet give little attention to consolidating evidence that is distributed across, and shares salience between, distant parts of a document. We introduce LexLattice, an extractive summarizer that reifies a legal act's hierarchy as a two-dimensional semantic lattice and consolidates over it with a masked 2D neural cellular automata before selection. LexLattice attains state-of-the-art ROUGE across all 24 languages of EUR-Lex-Sum in both multilingual and cross-lingual settings, surpassing instruction-tuned baselines with billions of parameters, despite concentrating all trainable capacity in a 1.8M parameter consolidator over a frozen multilingual encoder. A consolidator trained only on high-resource languages further transfers to unseen languages with near-lossless retention (0.99), indicating that the model operates on language-agnostic semantic geometry rather than surface form. Our results position explicit consolidation over document structure as a compact and traceable alternative to scale for multilingual legal summarization.
著者のコメント
14 pages, 4 figures
arXiv ID: 2609.27032 / 要約の誤りについて