業務の実演を見たAIが手順を理解したか測る評価資料
ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations
この論文をやさしく読む
ひとことで言うと
人が業務を実演しながら説明した後、AIが運用規則や例外を理解したかを測るための公開データと評価手順を作った。
何に役立つ?
業務自動化エージェントの学習画面や理解度を比較・改善する評価資料になる。予備試行は製品の優劣を確定するものではない。
この研究の面白いところ
50課題、502質問を公開し、回答ミスだけでなく、教える操作そのものを完了できない問題も評価対象にした。
どこまで分かった?
著者らは公開物の検証不足、予備試行の不均一な対象範囲、除外事項、採点の由来を明示している。218回の選別試行は統制された製品ランキングではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
私たちは仕事を実演し、その都度判断の理由を説明して同僚に教えることがある。同じ教え方でエージェントが何を理解したかを、どう確かめればよいだろうか。本研究は、説明付きの業務実演後の理解を評価するベンチマーク手順と公開データセット ShowTellArena を導入する。バージョン1.0には、録画、スクリーンショット、ナレーション、初期設定データ、502の質問を伴う50の業務ワークフロー課題が含まれる。課題は財務、採用、調達、顧客に関する判断、在庫、物流にわたる。評価手順では業務シナリオと質問を固定しつつ、各製品が独自の教育用画面を通して説明を取り込めるようにする。質問は運用規則、適用範囲、例外、提案された自動化の誤りを試す。39の業務事例にわたる、選別された218回の予備試行を分析した。このうち28事例は評価した3システムすべてが試行した。探索的な結果から、回答の誤りに加えて、教える体験を最後まで完了できない問題も見えてきた。公開物の検証上の不足や、予備試行における不均一な対象範囲、除外事項、採点の出所も説明する。貢献は、他者が拡張できる検査可能なデータセットと評価手順であり、選別された予備試行は製品を統制条件で順位付けしたものではない。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We often teach a colleague by showing the work and explaining the decisions as we go. How can we check what an agent understood from the same lesson? We introduce ShowTellArena, a benchmark protocol and public dataset for comprehension after narrated business demonstrations. The v1.0 release contains 50 business workflow tasks, with recordings, screenshots, narration, fixture seeds, and 502 questions. Tasks span finance, hiring, procurement, customer decisions, inventory, and logistics. The protocol holds the business scenario and quiz fixed while allowing each product to capture the lesson through its own teaching interface. Questions test operational rules, boundaries, exceptions, and errors in proposed automations. We analyze 218 selected pilot attempts across 39 workflow cases, including 28 cases attempted by all three evaluated systems. These exploratory results expose both answer errors and failures to complete the teaching experience. We describe the release's verification gaps and the pilot's uneven coverage, exclusions, and grading provenance. The contribution is an inspectable dataset and assessment workflow that others can extend; the selected pilot is not a controlled product ranking.
著者のコメント
6 pages, 1 figure, 2 tables. Code and dataset linked in the paper
arXiv ID: 2609.25467 / 要約の誤りについて