企業規模のデータ分析と業務判断を評価するArgo-Bench
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
この論文をやさしく読む
ひとことで言うと
大規模な企業データを調べ、分析結果に基づく業務措置まで行うAIを、模擬企業の中で評価するベンチマークです。
何に役立つ?
SQLが正しいかだけでなく、判断によってどのような結果が生じたかを測る評価に役立ちます。実際の機密データを公開せずに企業規模の課題を用意しています。
この研究の面白いところ
235表・75億行に分かれたデータから真実を復元し、予算配分や支払いなどの措置を出す必要があります。正解の文字列との一致ではなく、模擬環境での結果を採点します。
どこまで分かった?
企業は実データなどを根拠にしたシミュレーションであり、実企業での運用実績ではありません。最良モデルの34.8%は95点以上を得た課題の割合で、すべての意味での成功率とは区別が必要です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
実際の企業のデータサイエンスと分析では、何十もの表をまたいで推論し、統計分析を行い、その結果に基づいて行動する必要がある。既存のtext-to-SQLベンチマークはクエリ生成だけを評価し、監査では正解データがしばしば誤っていることも分かっている。実際の企業データウェアハウスは機密性が高く公開できないため、こうしたベンチマークは、1つの業務イベントが1つの表に収まる公開データセットで構成されている。 本研究は、210のデータサイエンス・分析課題からなる評価枠組みArgo-Benchを導入する。公開データ、査読付きの業界文献、規制当局への提出資料を基に、ニューヨーク市の料理配達プラットフォームを実際の規模でシミュレーションする。2024年の注文8,100万件、根拠に基づく経済条件、不正のパターン、市場の誘因を含む。この世界を、Oracle E-Business Suiteのスキーマに倣った235表・75億行のERPデータウェアハウスへ出力する。 シミュレータの真の状態は、エージェントに見えるウェアハウスからは隠される。このため、課題ではウェアハウスをたどって事実を再構成してから、それに基づき行動する必要がある。Argo-Benchはtext-to-SQLを超え、不正アカウントの停止、配達員への奨励金予算の配分、未払い分の支払いといった措置をエージェントが提出し、採点器はシミュレータ内での結果に基づいて採点する。すべての課題には、ウェアハウスだけを用いて解けることを示す、実行可能な参照解法がある。 14の最先端モデルおよび重み公開モデルのうち最も強いモデルでも、95点以上を得るのは課題の34.8%にとどまり、平均点は59.5点だった。Argo-Benchが、実際のデータ環境を理解し、たどり、その中で行動するエージェントの進歩を促すことを期待する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
著者のコメント
41 pages, 4 figures, 18 tables. Code: https://github.com/TextQLLabs/Argo-Bench. Data: https://huggingface.co/datasets/textql/Argo-Bench. Website: https://argo-bench.com
arXiv ID: 2610.02122 / 要約の誤りについて