arXiv論文メモ
新着一覧
cs.CR · 査読状況未確認

聴力の違いによって映像・音声ディープフェイクの見分け方はどう変わるか

"What I See is What I Hear": Deepfake Detection Across Diverse Hearing Abilities

Magdalena Pasternak (1), Malvika Jadhav (1), Palavi V. Bhole (2), Aviva Smith (1), Elaina Trapatsos (2), Vincent Bindschaedler (1), Roshan Peiris (2), Ersin Uzun (2), Patrick Traynor (1), Matthew Wright (2), Kevin R. B. Butler (1) ((1) University of Florida, (2) Rochester Institute of Technology)

この論文をやさしく読む

ひとことで言うと

聴力が異なる人々が映像と音声のディープフェイクをどの程度見分けられるかを比べた調査。

何に役立つ?

ディープフェイク警告や検出支援を、聴覚に頼りすぎず利用できるよう設計する際の根拠になる。

この研究の面白いところ

全体の差は偽物を見逃すことだけでなく、本物を偽物と誤判定する率の違いにも大きく関係していた。

どこまで分かった?

結果は80人が各30本の動画を判定した対面調査に基づき、加工の種類ごとに正答率が異なった。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

映像と音声を組み合わせたディープフェイクの増加により、詐欺、なりすまし、誤情報の作成コストは下がったが、その成功は最終的には人間の知覚に依存する。検出には聴覚と視覚の手がかりを統合する必要がある一方、セキュリティとプライバシーの研究では、ろう者や難聴者がほとんど考慮されてこなかった。本研究はこの不足に対し、対面で行う混合手法の調査に80人を参加させた。内訳は聴者31人、難聴者15人、ろう者17人、人工内耳利用者17人である。各参加者は、音声合成、声の変換、口の動きの同期、顔の入れ替えのいずれかによる加工を含む30本の動画について、本物かどうかを判断した。ろう・難聴の参加者の全体の正答率は聴者より低く、76.4%対88.0%だった(p<0.001)。主因は本物の動画を加工済みと判定することが多かったためで、偽陽性率は29.7%対11.2%だった。違いは加工された情報の種類に大きく依存した。音声だけの加工では難聴者は聴者とほぼ同じ正答率で、90.0%対90.3%だった。人工内耳利用者は79.4%、ろう者は41.2%だった。映像と音声の両方が加工された動画では正答率が84~87%に集まったが、加工方法による違いはなお残った。本研究は、ディープフェイクが聴力の異なる人々へ与える影響を体系的に特徴づけ、映像・音声の加工がもたらすリスクの非対称性と、すべての利用者を支える利用しやすく個別に調整された防御策の必要性を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The proliferation of audiovisual deepfakes has lowered the cost of fraud, impersonation, and misinformation, but their success ultimately depends on human perception. Detection requires integrating auditory and visual cues, yet security and privacy research has largely overlooked d/Deaf and hard-of-hearing (DHH) populations. We address this gap with an in-person, mixed-methods study of 80 participants: 31 hearing persons (HPs), 15 hard-of-hearing (HoH) participants, 17 d/Deaf participants, and 17 cochlear implant (CI) users. Each participant judged the authenticity of 30 clips, where manipulations spanned text-to-speech, voice conversion, lip-sync, or face-swap. DHH participants were less accurate than HPs overall (76.4% vs. 88.0%, p<.001), primarily because they more often classified authentic clips as manipulated (FPR: 29.7% vs. 11.2%). Differences depended strongly on the manipulated channel. For audio-only manipulations, HoH participants matched HPs (90.0% vs. 90.3%), followed by CI users (79.4%) and d/Deaf participants (41.2%). When clips contained an audiovisual manipulation, accuracy clustered between 84% and 87%, although performance still varied by manipulation method. Our work systematically characterizes how deepfakes affect DHH populations, highlighting the asymmetric risks audiovisual manipulations may pose to groups with different hearing abilities and the need for accessible, tailored defenses that support all users.

著者のコメント

Proceedings of the Network and Distributed System Security (NDSS) Symposium 2027

arXiv ID: 2609.28659 / 要約の誤りについて