arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

23言語でAIエージェントの操作能力を比較する

BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

Peng Kuang, Yuchun Fan, Jiangnan Li, Minghao Wu, Jialong Tang, Hao-Ran Wei, Weixuan Wang, Jianhong Tu, Baosong Yang, Tong Xiao

この論文をやさしく読む

ひとことで言うと

同じ操作課題を23言語にそろえ、言語によってエージェントの成功や失敗の仕方がどう変わるかを測ります。

何に役立つ?

多言語サービスで、回答の自然さだけでなくツール操作の正確さやトークン消費を評価するために使えます。

この研究の面白いところ

低資源言語の問題を回答品質だけで説明せず、制御フローやツール使用の誤り、英語への切り替わりまで調べています。

どこまで分かった?

16,146件は702の原型課題を各言語に展開したインスタンス数です。最大約2倍は入力トークンの比較であり、会話回数や総費用が常に2倍という意味ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)エージェントは、ツールを使い、利用者や環境とやり取りしながら、複数段階の処理を行うことが増えている。しかし現在の評価は英語中心で、多言語環境での能力の理解を制限している。本研究ではBabelFlowを導入する。これは既存のエージェント用ベンチマークを新たな言語へ適応させる、ベンチマークに依存しないエージェント型の作業手順である。実行時の依存関係を分析し、構造を保つ翻訳を調整し、多層の検証と人間による確認を組み合わせて、課題と評価の意味を維持する。 BabelFlowを用いて、4つのベンチマーク系統、13領域、23言語にわたる702の原型課題から16,146インスタンスを作成し、課題を対応させたベンチマークBabelArenaを構築する。5つの先端モデルによる実験では、全ベンチマーク系統を通じて優位な単一モデルはなく、言語間の差は課題の成否にとどまらない。低資源言語には特有の失敗傾向があり、回答の質に関する誤りだけでなく、ツール使用と制御フローの誤りが占める割合が大きい。これは、言語の資源量の違いに伴う、信頼できる課題実行の隔たりを示している。 同じ課題でも、低資源言語のエージェントは英語より大幅に多くのトークンを消費し、入力は最大でおよそ2倍となるが、対話の長さがそれに比例して増えるわけではない。また、構造化出力を必要とする課題では使用言語の一貫性がさらに低下し、言語の切り替わり先は圧倒的に英語に偏る。BabelArenaは、信頼性と効率性に優れた多言語エージェントの研究を進める基盤になると考える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language model (LLM) agents increasingly execute multi-step workflows through tool use and interaction with users and environments. However, current agent evaluations are largely English-centric, limiting our understanding of agent capabilities in multilingual settings. We introduce BabelFlow, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-preserving translation, and combining multi-layer verification with human review to preserve task and evaluation semantics. Using BabelFlow, we construct BabelArena, a task-aligned benchmark comprising 16,146 instances derived from 702 canonical tasks across four benchmark families, 13 domains, and 23 languages. Experiments with five frontier models show that no single model dominates across benchmark families and that cross-language disparities extend well beyond task success. Lower-resource languages exhibit distinct failure patterns, with larger shares of tool-use and control-flow errors rather than answer-quality errors alone, pointing to gaps in reliable task execution across the resource levels of these languages. On the same tasks, agents in low-resource languages also consume substantially more tokens than in English (up to roughly twice the input) without proportional increases in interaction length, and language consistency degrades further on tasks requiring structured output, where switches are directed overwhelmingly toward English. We believe BabelArena provides a foundation for advancing research on reliable and efficient multilingual agents.

著者のコメント

20 pages, 11 tables, and 7 figures

arXiv ID: 2609.23490 / 要約の誤りについて