arXiv論文メモ
新着一覧
cs.AI / cs.CL / cs.LG · 査読状況未確認

合成音声の学習効果を論じる前に生成履歴を監査する

Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects

Sidi Chang, Peiying Zhu

この論文をやさしく読む

ひとことで言うと

合成音声がモデルに与えた影響を調べる前提として、音声の生成元や検査履歴をどこまで追跡できるかを監査しています。

何に役立つ?

合成データの研究で、後から検証できる記録を設計する助けになります。ケア引き継ぎは監査の事例であり、臨床的な有効性の検証ではありません。

この研究の面白いところ

波形とラベルだけでは不足するとし、生成設定、レビュー、版管理まで一つの監査単位に結び付けています。

どこまで分かった?

生成器の固定不足やクリップ単位の版情報欠落により、厳密な上流帰属はできていません。著者は因果的な学習効果、臨床的妥当性、無制限公開を明示的に主張していません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

モデルの振る舞いを合成学習データに帰属させるには、各学習項目が何を引き起こしたかを推定する前に、その項目を何が生成したかを知る必要がある。波形とラベルの組だけでは、この知識は保存されない。そこで、合成研究オブジェクトが、生成元の仕様、生成内容、波形、ターゲット、事実要件、品質信号、レビューの系譜、不変のマニフェスト識別情報を結び付ける、生成来歴の基盤を提案する。証拠としての意味を決めるのは生成主体と選択の仕組みであり、保存場所や変数名ではない。 この基盤を、非公開の日本語ケア引き継ぎパイプラインで監査する。レビュー対象の113アセットには、6系統のシナリオにわたる1.552時間の合成音声が含まれる。全項目に音声、書き起こし、候補メモ、事実チェックリストが関連付けられているが、人による証拠は選択的で、生成元に依存している。内容に忠実な項目だけを含む2つのマニフェストは、シナリオのシードが重ならず、不変の版として管理されている。一方、生成器の別名が固定されていないこと、クリップごとの音声合成およびコードの版情報が欠けていること、確認用プロンプトが版管理されていないことにより、上流への厳密な帰属はなお不可能である。 生成来歴は振る舞いの帰属に必要だが、それだけでは十分ではないと論じる。生成来歴は候補となる因果グラフと監査単位を定めるが、寄与の帰属には、固定された学習実行と、介入または影響に関する証拠が依然として必要となる。本論文は、簡潔な来歴契約、監査手順、範囲を限定した合成データ帰属の事例研究を提供する。管理された研究アクセスを提供する可能性はあるが、学習データの因果的帰属、臨床的妥当性、無制限の一般公開を主張するものではない。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.

著者のコメント

Accepted to the Third NeurIPS Workshop on Attributing Model Behavior at Scale: Data Attribution and Provenance. 4 pages, 0 figures, 1 table. An aggregate reproducibility package is available from the authors on request!

arXiv ID: 2610.01378 / 要約の誤りについて