ツール使用エージェントの規則違反を試す課題を自動生成
EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation
この論文をやさしく読む
ひとことで言うと
仕様で禁止・制限されている操作を試す難しい課題を、実際のデータベースの状態に合わせて自動生成し、エージェントの改善に使います。
何に役立つ?
正常な依頼だけでは見つからない規則遵守の弱点を試験するためのデータ作成に役立ちます。生成データをモデルの微調整と周辺実行環境の最適化に使っています。
この研究の面白いところ
抽象的な難問を作るのではなく、仕様から規則を取り出し、データベースに存在する状態と結び付けて課題を作る点が特徴です。
どこまで分かった?
報告指標は平均進捗であり、全タスク成功率とは異なります。要旨の改善は%表記で、相対率かポイント差かの詳細はありません。ハーネスの数値比較は Gemma-4-e4b に関する結果です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ツールを呼び出す LLM エージェントは、企業向けアプリケーションへの導入が進んでいる。しかし、効果的な評価と最適化には、質が高く多様なタスクのデータセットが必要であり、プライバシーなどの制約から入手が難しいことが多い。既存の合成タスク生成法は、エージェントの基盤となる状態やデータベースを無視した一般的なタスクを作ることが多く、実世界の利用の多様性を反映できない。 本研究では、エージェントの仕様から遵守すべき規則を抽出し、それらに違反するよう設計した、データベースに根差す境界事例のタスクを生成する枠組み EdgeGen を提案する。既存の合成データ生成技術と組み合わせることで、EdgeGen は微調整とハーネス最適化を通じたエージェントの改善を可能にする。得られる処理系は、人手の注釈を必要としない完全自動の閉ループシステムを形成する。EdgeGen で生成したデータによる微調整は、tau2bench の航空会社領域で平均進捗を一貫して2〜42%改善する一方、ほかの基準手法では一部のモデルに性能低下がみられる。また、Gemma-4-e4b モデルのハーネス最適化では、人が整備したハーネスおよび基本ハーネスに対して、それぞれ10%と30%の平均進捗改善を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective evaluation and optimization require high-quality, diverse task datasets that are often difficult to obtain due to privacy and other constraints. Existing synthetic task generation methods often produce generic tasks that ignore an agent's underlying state or database and fail to reflect real-world usage diversity. We propose EdgeGen, a synthetic task generation framework that extracts compliance rules from an agent's specification and uses them to generate database-grounded edge-case tasks designed to violate these rules. When combined with existing synthetic data generation techniques, EdgeGen enables agent improvement through finetuning and harness optimization. The resulting pipeline forms a fully automated closed-loop system that requires no human annotation. Finetuning on data generated by EdgeGen yields a consistent mean progress improvement of 2 percent to 42 percent on tau2bench airline domain, while other baseline methods show degradation for some models. On the other hand, for harness optimization, our method shows a mean progress improvement of 10 percent and 30 percent over the human-curated and base harnesses, respectively, for the Gemma-4-e4b model.
著者のコメント
NA
arXiv ID: 2609.24115 / 要約の誤りについて