arXiv論文メモ
新着一覧
cs.CR · 査読状況未確認

子どものスマートフォン操作を後から調べる記録方式

GUIAuditor: Enabling Post-hoc Child Safety Forensics via Action-Guided GUI Provenance on Mobile Devices

Junlin Liu, Yifeng Cai, Shuai Wang, Zhineng Zhong, Shaofei Li, Jiacheng Liu, Yuanchun Li, Ziqi Zhang, Xiangqun Chen, Ding Li, Yao Guo

この論文をやさしく読む

ひとことで言うと

子どものアプリ操作を意味のある記録に変え、保護者が後から危険な出来事を調べられるようにする方式。

何に役立つ?

自動検出だけに頼らず、保護者が具体的な操作の証拠を探して確認する仕組みを検討する際に役立つ。

この研究の面白いところ

解析対象データを89.2%超削減しつつ、重要イベントの記録精度、質問からの証拠検索、端末上の負荷をまとめて評価した。

どこまで分かった?

評価は295件の操作クリップと3機種のスマートフォンに基づく。すべての危険を検出できるとは述べていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

スマート機器の普及により、子どもは正規のアプリ内に入り込んだ誘い出しや金銭詐欺などのオンライン上の危険にさらされる。自動の予防・検出は、ルール方式でもAI方式でも誤検出や見逃しを避けられないため、本研究は人が関わって事後に調べる補完的な方法を提案する。GUIAuditor は、子どもの操作の流れについて検索可能で意味を持つ記録「GUI Provenance」を作るシステムである。複数の形式を扱う大規模言語モデルを使い、画面操作イベントの時系列を人が理解できる説明へ変換する。モバイル端末で実用的にするため、証拠を絞り込む処理によって、業界標準で採用される定期サンプリング方式と比べ、解析が必要なデータを89.2%超減らし、精度への影響をわずかに抑えた。295件の操作クリップからなる新しいデータセットでは、重要なイベントの記録でマクロF1値95.23%を得た。また、二段階の事後調査用検索エンジンは、自然言語による質問の90.20%超で正しい証拠を検索結果の最上位に置いた。最新のスマートフォン3機種での一連の評価では、端末上のモデル推論を含む全処理によって消費電力が2.1W増え、イベント当たりの遅延は7.4秒、ピークメモリ使用量は約3.1GBだった。これらの結果は、モバイル端末上で事後の画面操作調査を実行し、保護者による安全確認に役立つ文脈を提供できることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDFDOI

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The proliferation of smart devices exposes children to online risks like grooming and financial scams that are deeply embedded within legitimate applications. Current approaches rely on automated prevention and detection, a paradigm that is fundamentally limited by its inherent fallibility. Whether rule-based or AI-driven, they inevitably produce false positives and negatives, failing to provide reliable protection. In this paper, we argue for a complementary, human-in-the-loop, post-hoc forensic paradigm. We present GUIAuditor, the first system designed to realize this vision by creating GUI Provenance: a queryable, semantic record of a child's interaction sequence. To generate this, GUIAuditor leverages a Multimodal Large Language Model (MLLM) to translate the temporal sequence of GUI events into a human-understandable narrative. To make this practical on mobile devices, a novel evidence distillation pipeline reduces the data requiring analysis by over 89.2% compared to periodic sampling approaches adopted by industry standards, with negligible impact on accuracy. On a new dataset of 295 interaction clips, GUIAuditor achieves a 95.23% Macro-F1 Score in logging significant events and, crucially, its two-stage forensic query engine successfully retrieves the correct evidence as the top result for over 90.20% of natural language questions. An end-to-end evaluation on three modern smartphones shows that the full pipeline, including on-device MLLM inference, adds 2.1W of power draw and 7.4s of per-event latency, with a peak memory footprint of ${\sim}$3.1GB. These results show that post-hoc GUI forensics can run on modern mobile devices and provide useful context for guardian-led safety review.

著者のコメント

Accepted by ACM IMWUT/Ubicomp 2026

arXiv ID: 2609.28205 / 要約の誤りについて