arXiv論文メモ
新着一覧
cs.SD / cs.CL · 査読状況未確認

似た通常電話を含めて音声詐欺検出を評価するデータセット

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen

この論文をやさしく読む

ひとことで言うと

詐欺電話と話題がよく似た正当な電話を含め、音声AIが本当に両者を見分けられるかを試す評価基盤です。

何に役立つ?

単に詐欺らしい話題を見つけるだけのモデルを高評価してしまう問題を調べられます。月ごとのテスト集合を固定することで、更新前の評価記録も保てます。

この研究の面白いところ

テキスト実験では、簡単な負例で完全だったMacro-F1が、似た文脈の負例に替えると0.65〜0.68へ下がります。評価データの作り方が結果を大きく左右する例です。

どこまで分かった?

電話は事例要約からシナリオ・対話・音声を生成して構築されます。実際の通話録音900件をそのまま集めたという説明ではありません。また0.65〜0.68は統制したテキスト実験の結果です。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

電話詐欺の台本は急速に変化し、通常のサービス会話に似せて作られることも多い。そのため、音声による電話詐欺検出の評価には2つの重要な要件がある。第一に、既に確立したテスト集合を上書きせず、新たに観測された詐欺パターンをベンチマークに取り込む必要がある。第二に、話題が異なる負例に頼らず、詐欺と、同じ領域に近い合法的な電話とを区別する必要がある。本研究では、Mixed-Tree Anti-Fraud Generation Pipelineで構築し、月ごとに固定した評価プロトコルで評価するTeleAntiFraud 2.0を提示する。この処理系は、オンラインの詐欺事例の要約をプロフィールに基づくシナリオへ変換し、混合木による生成で拡張し、共通の文脈の下で詐欺・非詐欺の対話経路を実現する。 検証済みの対話を役柄に合った音声へ変換し、得られた音声、ラベル、プロンプト、マニフェスト、来歴記録を月ごとの評価集合として固定する。各固定集合は中国語の電話900件からなり、詐欺600件と、同領域に近い非詐欺300件を含む。条件を統制したテキスト実験では、無関係または通常の負例で評価すると3つの分類器が完全なMacro-F1を達成する一方、同領域の兄弟的な負例では0.65〜0.68へ低下する。全集合での音声評価と、音声認識と大規模言語モデルを組み合わせたASR+LLM評価では、クラスの事前比率に依存した近道、予測の崩壊、スナップショットへの感度も明らかになった。これらの知見は、現実的に混同しやすい条件で音声電話詐欺モデルを評価するには、近接領域の例の構築と、予測の崩壊を考慮した報告が中核的な要件であることを示す。 付属する研究資料は構築コード、評価スクリプト、マニフェスト、文書を含む。データセットとコードは https://anonymous.4open.science/r/TeleAntiFraud-2_0-EEB2/ で公開している。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-17 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key requirements for audio-based telecom-fraud evaluation. First, benchmarks must incorporate newly observed scam patterns without overwriting previously established test sets. Second, they must distinguish fraud from lawful, near-domain calls rather than relying on topic-separated negative examples. We present TeleAntiFraud 2.0, constructed with our Mixed-Tree Anti-Fraud Generation Pipeline and evaluated under a monthly frozen evaluation protocol. The pipeline transforms online fraud-case abstracts into profile-grounded scenarios, expands them through mixed-tree generation, realizes fraud and non-fraud dialogue paths under shared contexts, renders validated dialogues as role-matched speech, and freezes the resulting audio, labels, prompts, manifests, and provenance records for each monthly evaluation set. Each frozen set contains 900 Chinese calls, comprising 600 fraud and 300 near-domain non-fraud cases. Controlled text experiments show that three classifiers achieve perfect macro-averaged F1 (Macro-F1) when evaluated against unrelated or ordinary negatives, but drop to 0.65-0.68 with near-domain sibling negatives. Full-set audio and automatic-speech-recognition plus large-language-model (ASR+LLM) evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity. Together, these findings establish near-domain construction and collapse-aware reporting as core requirements for evaluating audio-based telecom-fraud models under realistic confusable conditions. The accompanying research artifact includes the construction code, evaluation scripts, manifests, and documentation. Our dataset and code are available at https://anonymous.4open.science/r/TeleAntiFraud-2_0-EEB2/.

著者のコメント

12 pages, 4 figures, including supplementary material

arXiv ID: 2609.18748 / 要約の誤りについて