arXiv論文メモ
新着一覧
cs.AI / cs.IR · 査読状況未確認

国際法文書の固有表現認識データセットとモデル評価

IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

Genis Skura and Roland Bouffanais and Didier Wernli

この論文をやさしく読む

ひとことで言うと

国際法の判決や決議に出てくる固有表現を識別するための正解データを作り、複数のモデルを比較した研究です。

何に役立つ?

国際法文書内の参照先などを抽出する技術の評価に使うことが想定されています。人手確認を含むデータ作成方法の検討にも役立ちます。

この研究の面白いところ

境界が一致した表現だけを比べると人と機械の一致が非常に高く見えても、見落としや境界の誤りを含めると評価が変わることを数値で示しています。

どこまで分かった?

対象は3種類の機関文書と7種類の表現タイプです。κ、macro-F1、micro-F1は異なる集計であり、そのまま同一の正解率として比較できません。法的判断の正しさを評価した研究ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

国際法は、国家が行動を調整し、武力紛争を規制し、人権を保護するための規範的な枠組みを提供するが、その文書にはトークン単位の固有表現認識(NER)資源がまだない。本研究では、国際法の成文化された情報源を対象とするNERデータセット兼ベンチマークIntLawNERを導入する。国際司法裁判所(ICJ)の決定、国連安全保障理事会の決議、欧州人権裁判所(ECtHR)の判決から、正解注釈付き2,987文と8,094の固有表現範囲を収録し、機関に固有の7種類の表現タイプで注釈付けした。 IntLawNERは、候補検索、LLMによる審査、人間による確認を通じ、468,000の原文文を小規模な注釈集合に絞り込む、費用効率のよいアルゴリズムとエージェントの混合処理系で構築した。正解データの表現範囲の89.6%は、暫定注釈であるシルバー層から変更なしで採用された。しかし、シルバーから正解データへの分析により、専門分野のNERでは、人と機械の集計一致指標が誤解を招きうることが分かった。境界が一致する表現範囲でのCohenのκは0.964だが、見落とされた表現、境界の誤り、ラベル修正を含めるとmacro-F1は0.753であり、高いκの背後にこの差が隠れていた。 ベンチマークでは、ゼロショットの表現範囲ベースモデルGLiNERが、表層形よりも機関内の機能に依存する表現タイプで大きく性能を落とし、micro-F1は0.243だった。また、微調整したTransformerはまれなラベルに苦戦した。ラベルの対比を示す少数例を注意深く選ぶと、すべてのLLMでゼロショットの指示より性能が向上し、Claude Opus 4.6がmicro-F1 0.873で最高だった。国際法文書内の参照表現を抽出するためのベンチマークおよび再利用可能な資源として、IntLawNERを公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its texts remain without token-level named entity recognition (NER) resources. We introduce IntLawNER, a NER dataset and benchmark for codified sources of international law, covering 2,987 gold-annotated sentences and 8,094 entity spans from International Court of Justice (ICJ) decisions, UN Security Council resolutions, and European Court of Human Rights (ECtHR) judgments, annotated with seven institution-specific entity types. We construct IntLawNER with a cost-effective hybrid algorithmic-agentic pipeline that reduces 468k source sentences to a compact annotation set through candidate retrieval, LLM-based vetting, and human review, with 89.6% of gold spans accepted unchanged from the silver layer. However, the silver-to-gold analysis reveals that human-machine aggregate agreement metrics can be misleading in domain-specific NER: Cohen's kappa=0.964 on boundary-matched spans masks a macro-F1 of 0.753 when missing entities, boundary errors, and label corrections are included. The benchmark shows that zero-shot span-based GLiNER collapses on entity types dependent on institutional function rather than surface form (0.243 micro-F1), while fine-tuned transformers struggle on rare labels. Carefully selected few-shot examples that demonstrate label contrasts improve every LLM over zero-shot prompting, with Claude Opus 4.6 reaching the best score of 0.873 micro-F1. We release IntLawNER as a benchmark and reusable resource for extracting references in international legal texts.

arXiv ID: 2609.22529 / 要約の誤りについて