arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

尋問対話を属性別に要約するCASPER

Controlled Attribute-Specific Summarization of Interrogative Dialogues

A Aditya Bhardwaj, Arjit Singh Arora and Md Shad Akhtar

この論文をやさしく読む

ひとことで言うと

尋問者と証人の対話を、出来事や人物など指定した属性に沿って要約する手法とデータ集。

何に役立つ?

調査記録を整理する際に、必要な属性を落とさず要約する支援として使える可能性がある。実際の法的判断を自動化して安全に行えることは示していない。

この研究の面白いところ

6000組の発話対に注釈を付け、複数の役割による反復評価と構造化フィードバックを要約生成に組み込んでいる。

どこまで分かった?

要旨はROUGE、BERTScore、人手評価での優位を述べるが、改善幅の数値や実運用時の誤り率は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

尋問形式の対話を適切に要約することは、法科学や調査の場面で重要であり、事実の正確さ、一貫性、指定した属性との関連性が求められる。本研究は、尋問者と証人のやり取りを質の高い要約にするため、構造化されたプロンプトと反復的な修正を用いるCASPERという、思考の連鎖に基づく属性別の評価用要約枠組みを提案する。MINDコーパスを拡張したMINDSumデータセットも構築し、出来事の詳細、事実の陳述、人物描写、つなぎ言葉を注釈した6000組の発話対を収録する。 CASPERはRoleEvalという階層的な評価機構を用い、担当官、調査官、上級調査官という複数の役割が、事前に決めた基準に基づき要約を繰り返し評価する。固有表現抽出と構造化されたフィードバックの反復を組み合わせることで、従来の比較手法よりも事実の一貫性と文脈の網羅性を大きく改善する。実験では、標準的な要約モデルより語彙的指標ROUGEと意味的指標BERTScoreの両方で高い成績を示し、人手評価でも専門家の推論との整合性が確認された。結果は、厳密さが求められる分野で制御可能な要約を使える可能性を示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Effective summarization of interrogative dialogues is a critical task in forensic and investigative settings, requiring high factual accuracy, coherence, and attribute-specific relevance. In this work, we introduce CASPER, a novel Chain-of-Thought Attribute-Specific Prompting for Evaluative Summarization framework that leverages structured prompting and iterative refinement to generate high-quality summaries of interrogator-witness interactions. We construct MINDSum, a dataset extending the MIND corpus, comprising 6,000 utterance pairs annotated with event details, factual statements, character descriptions, and fillers. CASPER employs RoleEval, a hierarchical evaluation mechanism where multiple roles (officer, inspector, senior inspector) iteratively assess summaries based on predefined criteria. By integrating entity extraction and structured feedback loops, CASPER significantly improves factual consistency and contextual completeness compared to existing baselines. Experimental results demonstrate that our framework outperforms standard summarization models on both lexical (ROUGE) and semantic (BERTScore) metrics, while human evaluation confirms its alignment with expert reasoning. Our findings underscore the potential of controlled summarization in high-stakes domains, paving the way for AI-driven forensic intelligence.

arXiv ID: 2609.28004 / 要約の誤りについて