arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

動きから図形を見つけるCAPTCHAの評価

Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents

Zeyu Zhang, Dingyi Rong, Zijian Chen, Zicheng Zhang, Xiongkuo Min, Guangtao Zhai

この論文をやさしく読む

ひとことで言うと

静止画では分からず、背景の中で動く形を見分けると解けるCAPTCHAを作り、人間と画面操作エージェントで比較した。

何に役立つ?

現在の画面操作エージェントが、動きで定義された形をどこまで認識できるかを測る材料になる。実際の認証サービスで長期にわたり有効かを検証した結果とは区別が必要である。

この研究の面白いところ

600件のブラウザー上の問題で、人間99.6%、最高性能のエージェント16.8%という差が出た。前景だけの対照条件から、動く背景が難しさの主因だと分析している。

どこまで分かった?

この差は評価したエージェントとベンチマークでの結果であり、将来のエージェントにも同じ差が続くとは示していない。実運用での利用者負担や回避への耐性は要旨に記されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

既存の視覚的なCAPTCHAの多くは、静止画像の見た目、局所的な構造、画面の状態から必要な情報を読み取れる。マルチモーダル大規模言語モデルと画面操作エージェントの進歩により、強い視覚認識、推論、ブラウザー操作能力が得られ、この前提は弱まっている。 著者らは、動きを使う階層的なCAPTCHAの枠組みMotion Vision CAPTCHA(MVCAP)を提案する。答えの意味を持つ前景の形は動きによって定義され、時間とともに変化する背景から動きに基づいて分離して初めて見つけられる。この共通原理に基づき、知覚の難しさが段階的に高まる3つの水準、一貫した動き、構造的な動き、生物らしい動きを用意した。評価のために、ブラウザー上で動く600件のCAPTCHAからなるMVCAP-Benchと、対応する前景だけの対照ベンチマークMVCAP-Bench-FGを作成した。人間、ブラウザー操作エージェント、ネイティブのコンピューター操作エージェントに加え、同じ問題から作ったオフラインの視覚質問応答設定を評価した。 MVCAP-Bench全体では人間の正答率が99.6%だったのに対し、最も良い画面操作エージェントでも16.8%で、6択の偶然の正答率に近かった。前景だけを示す対照実験は、難しさの主因が答えの形式やブラウザー操作だけではなく、動的な背景による目くらましにあることをさらに示した。この結果は、人間と現在のエージェントの間に測定可能な動きの知覚能力の差があることを示し、MVCAP-Benchをその研究用のベンチマークとして位置づける。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Most existing visual CAPTCHAs remain spatially solvable: the required information is exposed by static appearance, local structure, and interface state. This assumption is weakened by advances in multimodal large language models (MLLMs) and Graphical User Interface (GUI) agents, which exhibit strong visual perception, reasoning, and browser interaction capabilities. We propose Motion Vision CAPTCHA (MVCAP), a hierarchical motion-based CAPTCHA framework in which target semantics are instantiated as motion-defined foreground structures and become recoverable only through temporal segregation from a dynamically evolving background. Built on this shared principle, MVCAP is instantiated in three perceptually progressive levels: coherent motion, structural motion, and biological motion. To evaluate this framework, we introduce MVCAP-Bench, a browser-based benchmark with 600 live CAPTCHA instances, together with a matched foreground-only control benchmark, MVCAP-Bench-FG. We evaluate humans, Browser Use agents, native computer use agents, and a supplementary offline VQA setting derived from the same instances. Results reveal a substantial human--agent gap: on the full MVCAP-Bench, human accuracy reaches 99.6%, whereas the best GUI agent achieves only 16.8%, close to the six-way chance level. The foreground-only control further shows that the key difficulty comes from dynamic background camouflage rather than answer format or browser interaction alone. These findings identify a measurable human--agent perception gap and position MVCAP-Bench as a benchmark for studying motion-defined perception in current agents.

著者のコメント

Accepted at ACM Multimedia 2026. 10 pages, 5 figures

arXiv ID: 2609.27461 / 要約の誤りについて