arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

データ作成から運用まで循環させるモバイル操作AI

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

Tingyu Qu, Weigao Sun, Yuecheng Liu, Yucheng Zhao, Yi Zhu, Yifeng Ding, Qiyi Wang, Sihan Cao, Pengkun Jiao, Hanlei Xie, Xiongwei Wu, Qichao Wang, Haodong Zhang, Jiajun Liu, Yuhao Wang, Yuqing Xie, Junpeng Zhao, Long Chen, Ming Ma, Sihan Yang, Ziwang Zhao, Yanhao Jia, Liangquan Gong, Feida Zhu, Yiran Zhong, Steven Hoi

この論文をやさしく読む

ひとことで言うと

スマートフォンなどを操作するAIを、データ作成、訓練、運用時の失敗記録を循環させて改善する仕組みです。

何に役立つ?

長い手順を要するモバイル操作エージェントの開発方法として考えられます。要旨ではMobilePA-Benchで比較システム中の総合性能が最も高いと報告しています。

この研究の面白いところ

人の確認を挟んだデータ作成と、実行の証拠を利用したモデル・ツール環境の共同改善を、一つの循環にまとめています。

どこまで分かった?

要旨には具体的な成功率や費用削減幅は示されていません。実機での長期運用の信頼性はこの要旨だけでは判断できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルの急速な進歩により、AIは受動的な文章生成から、工学や科学的発見の能動的な作業へ広がっている。本研究は、AIが次世代のAIシステムを開発する対象であると同時に、その開発に参加できるかを問う。拡張可能な開発と反復改善のため、閉ループのAI-for-AIの枠組みでQwen-Planner-Agentを構築する。モバイル機器上の計画はその厳しい試験となる。複雑で長い手順を要する課題はエージェントの信頼性を試し、実機とのやり取りには費用がかかって開発規模を制限する。枠組みは、行動、フィードバック、検証について共通の取り決めを使い、データ作成、モデル訓練、運用をつなぐ。 第一に、AI for Dataは人の確認を挟むエージェント型のデータ循環を作る。専門のエージェントが課題を構成し、操作の軌跡を収集し、訓練データを選別・均衡化し、訓練のフィードバックで次のデータ作成を導く。第二に、AI for Trainingは、教師ありの計画学習から始め、複合環境でオンラインのエージェント型強化学習を行う。そこでCAREという報酬と優位度の設計を導入し、課題性能を保ちつつ推論とツール利用の費用を減らす。第三に、AIはモデルとその実行枠組みを共進化させる。実行の証拠を使う循環が、運用時の記憶、技能、ツールを調整し、構造化した行動フィードバックと保存した失敗記録を、モデルと枠組みの協調的な修正へ戻す。Qwen-Planner-AgentはMobilePA-Benchで評価した全モデル・システム中、総合性能が最も高く、ツール利用、記憶、技能、下位エージェントの協調にわたり基盤モデルを改善した。追加の評価では、モバイル以外のエージェント課題でも改善し、一般能力をおおむね維持した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.

著者のコメント

https://tongyi-mai.github.io/Qwen-Planner-Agent/

arXiv ID: 2609.29892 / 要約の誤りについて