当事者参加で音声データを整えるための枠組み
Towards participatory speech dataset curation: A queer case study and conceptual framework
この論文をやさしく読む
ひとことで言うと
クィア当事者を事例に、音声データを参加型で収集・整備する概念的な枠組みを提案する。
何に役立つ?
周縁化された集団の音声を扱う研究で、参加方法と本人の自律性を考える参考になる。
この研究の面白いところ
収集の慣習を見直し、コミュニティ定義から個人の自律性までを双方向の過程として捉える。
どこまで分かった?
提案は先行事例の検討に基づく概念的枠組みで、要旨には実施後の性能評価はない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
LGBTQIA+、すなわちクィアのコミュニティを事例として、当事者が参加して音声データセットを作る枠組みが必要な理由を示す。このコミュニティにはAIへの懸念と被害の報告があり、個人がクィアかどうかを特定できると称する「ゲイダー」技術の開発の試みも含まれる。著者らは一般的な音声データの収集方法、それがクィアの話者と関わる方法として適切でない可能性、クィアのコミュニティが関わった参加型AIの先行例、ほかの周縁化されたコミュニティの音声データ収集に特化した参加型の取り組みを検討する。この検討から、共同設計と知識共有の知見に基づき、周縁化されたコミュニティ自身によって、そのために、そしてともに行う参加型音声データの整備の概念的枠組みを作る。枠組みは、コミュニティの定義、プロジェクトの形成、参加の方法、個人の自律性という、重なり合い双方向に進む過程からなる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
In this paper, we motivate the need for a participatory speech dataset creation framework through a case study of the LGBTQIA+, or queer, community - a community with documented concerns about AI and reported harms, including attempts to develop 'gaydar' technologies that purportedly identify individuals as queer. We review common speech data collection practices, why these methods may be unsuitable for engaging with queer speakers, and discuss previous efforts in participatory AI with queer community engagement, as well as participatory endeavours specific to speech data collection for other marginalized communities. From this review, we develop a conceptual framework for participatory speech data curation by, for, and with marginalized communities drawing on insights from co-design and knowledge sharing. We propose a framework comprising overlapping and two-way processes of defining a community, project formulation, modes of participation, and personal autonomy.
著者のコメント
Accepted at Interspeech 2026
arXiv ID: 2609.25496 / 要約の誤りについて