arXiv論文メモ
新着一覧
cs.CR / cs.HC · 査読状況未確認

スマートフォンのデータからAIが推測する個人情報

Privacy Leakage Through AI-mediated Analysis of Smartphone Data

Sarah Radway, Zoe Robert, Matthew Soto, Julianna Cimillo, Sebastian Diaz, Meg Marco, James Mickens

この論文をやさしく読む

ひとことで言うと

スマートフォンのアプリが得たデータの一部からでも、AIが利用者の機微な情報を推測できることを参加者調査で示した。

何に役立つ?

アプリの権限表示や利用者の同意を設計するとき、元データへのアクセスだけでなく、そこから可能な推測を伝える必要性を考える材料になる。

この研究の面白いところ

465人が自分の電話で推測システムを使い、得られた推測の体験が、その後のデータ共有意欲にも影響したと報告している。

どこまで分かった?

Priva-Seeは広告技術企業の手法を著者らなりに再現したシステムであり、個別の企業が同じ推測を実施した証拠ではない。影響の具体的な大きさは要旨には記載されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

過去30年にわたり、オンライン広告業界は利用者の行動を追跡し、属性や関心を推測するために、大規模なデータ収集の仕組みを築いてきた。従来はIPアドレス、GPS座標、電子商取引の購入履歴、訪問URLといった、構造化された文字データを集約・分析していた。しかし近年の機械学習モデルは、構造化データだけでなく、写真、動画、受信箱、カレンダーなどのマルチメディアや非構造化テキストも自動分析できる。スマートフォンのアプリでは、このプライバシー上の危険が特に大きい。電話には機微な情報が集まりやすいが、利用者は、例えば写真へのアクセスを許可すると、画像のデータだけでなく、そこから推測できる自分自身の情報へのアクセスも与えることを理解していないかもしれない。 この危険を調べるため、アプリが集めた利用者データから推測を行う言語モデルベースのシステムPriva-Seeを構築した。これは、実際の広告技術企業が機械学習を使って利用者像を作る方法についての著者らの最善の理解を反映したものである。倫理審査委員会の承認を得た利用者研究では、465人の参加者が自分の電話にPriva-Seeを導入した。Priva-Seeは利用者データの一部にしかアクセスできなかったにもかかわらず、プライバシーを侵害する推測を行った。この体験は、その後に参加者がアクセス権限に関わるデータを共有しようとする意欲に大きな影響を与えた。観察されたプライバシー侵害を踏まえ、下流でデータをどう利用できるかを利用者によりよく伝えるため、スマートフォンOSがデータアクセスへの同意を得る方法の変更を提案する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Over the past thirty years, the online advertising industry built a large-scale data collection ecosystem, with the goal of tracking a user's online activity to infer their demographics and interests. Traditionally, the ecosystem relied upon the collation and analysis of highly-structured text data like user IP addresses, GPS coordinates, e-commerce purchase histories, and visited URLs. However, recent ML models can parse not only structured text, but also multimedia files and unstructured text inputs---meaning a user's photos, videos, inboxes, and calendars are now ripe for automated analysis. The privacy risks are particularly acute in the context of smartphone apps. A user's phone already acts as a natural collation point for sensitive user information, but users may not understand that permitting an app to, for example, access a user's photo does not just give the app access to the bytes in the photo: the app also receives access to inferences about the user that are enabled by the photo. To explore these privacy risks, we built Priva-See, an LLM-based inference system for app-collected user data; Priva-See reflects our best understanding of how real-life adtech companies would leverage machine learning to build user profiles. Through an IRB-approved user study, 465 participants deployed Priva-See on their phones; Priva-See made privacy-invasive inferences despite having access to only a subset of a user's data. We see the experience significantly impacted participant willingness to share permissions data moving forward. Based on the observed privacy violations, we suggest changes to how smartphone OSes should gather user consent for data access, to better inform users about downstream data usage capability.

arXiv ID: 2609.28537 / 要約の誤りについて