無人航空機の判断場面で29種類のAIを評価
PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI
この論文をやさしく読む
ひとことで言うと
ドローンの運用中に必要となる判断を記録から切り出し、AIがその時点の情報だけで適切な判断を選べるかを比較します。
何に役立つ?
自律システム向けAIの判断能力を、共通の条件で評価するためのデータになります。最高成績でも正解率は約半分で、評価上の課題が残っています。
この研究の面白いところ
センサー情報だけでなく、通信品質、ツール利用、他のエージェントとのやり取りも判断の文脈に含めます。将来の情報が解答へ漏れないようにしています。
どこまで分かった?
10,178件すべてで29モデルの最終評価をしたのではなく、設定選択は35件、固定評価は500件です。記録に基づく判断評価であり、実機の自律飛行の安全性や成功率を直接示すものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Physical AIの近年の進歩により、無人航空機(UAV)のように、動的環境で知覚、推論、計画、行動を行う自律システムへの基盤モデルの利用が加速している。既存のベンチマークは、物理的知覚、直感的物理理解、身体性を伴う移動、協調推論を評価するが、信頼できる自律性に必要なエージェントとしての意思決定を評価するものは少ない。本研究では、この能力を評価するPhysAI-Benchを導入する。 このベンチマークは、自律UAVミッションの会話記録から自動抽出した、標準化された10,178件の意思決定事例を含む。各事例には、ミッションの文脈、時間的依存関係、物理制約、Model Context Protocol(MCP)のツール呼出し、Agent-to-Agent(A2A)のやり取り、センサー観測に加え、遅延、パケット損失、スループット、エッジ負荷、ネットワークスライシングなど、AIネイティブな6Gネットワークの条件が保持される。各判断より前の情報だけを提示し、将来の事象の漏洩を防いで、オンライン意思決定に近付ける。 29種類の基盤モデルを2段階の手順で評価する。まず、人が検証した35事例の開発用集合について、0例・3例・5例を提示するプロンプトと4種類の温度設定を組み合わせた12条件を3回ずつ試し、モデルごとの設定を選ぶ。次に、選んだ設定を固定し、エピソードが重複しない固定の500事例の集合で3回評価する。GPT-5.3の正解率が52.00%で最も高く、GPT-5.2の49.40%、Grok 4.5の49.07%が続く。少数例を提示するプロンプトは全般に性能を改善する一方、温度の影響は限定的である。これらの結果は、Physical AIにおける信頼できるエージェントの意思決定が、依然として未解決の課題であることを示す。データセットは https://github.com/maferrag/physai-bench で公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynamic environments. Existing benchmarks assess physical perception, intuitive physics, embodied navigation, and collaborative reasoning, but rarely evaluate the agentic decision-making required for reliable autonomy. We introduce \textit{PhysAI-Bench}, a benchmark for evaluating this capability. It contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions. Each instance preserves mission context, temporal dependencies, physical constraints, Model Context Protocol (MCP) tool calls, Agent-to-Agent (A2A) interactions, sensor observations, and AI-native 6G network conditions, including latency, packet loss, throughput, edge load, and network slicing. We expose only information preceding each decision, preventing future-event leakage and approximating online decision-making. We evaluate 29 foundation models using a two-stage protocol. We select model-specific configurations from 12 combinations of zero-, three-, and five-shot prompting and four temperatures, tested in three runs on a 35-instance, human-verified development set. We then freeze each selected configuration and evaluate it in three runs on a fixed, episode-disjoint set of 500 instances. GPT-5.3 achieves the highest accuracy (52.00%), followed by GPT-5.2 (49.40%) and Grok~4.5 (49.07%). Few-shot prompting generally improves performance, while temperature has limited influence. The results demonstrate that reliable agentic decision-making in Physical AI remains an open challenge. The dataset is available at https://github.com/maferrag/physai-bench
arXiv ID: 2609.23695 / 要約の誤りについて