AIエージェントの職場課題評価で事前に文脈を選ぶ影響を測る
WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
この論文をやさしく読む
ひとことで言うと
職場課題のベンチマークで、課題に合わせて資料を事前に選ぶと評価がどれだけ易しくなるかを測る。
何に役立つ?
AIエージェントの情報探索能力を、課題解答能力と混同せずに評価する設計に役立つ。
この研究の面白いところ
組織の状態を先に固定して課題を後から出し、証拠へ到達するまでの作業を評価に残す。
どこまで分かった?
主な実証は架空の製薬会社の八課題・六役割と192件の対応評価。実際の企業環境で同じ差が出るとは記載されていない。
v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
多くの知識労働ベンチマークは個々の課題を中心に作られ、必要な文脈を課題が決まると同時か、その後に選ぶ。この設計では、課題のために組み立てた環境における職場に似た課題の性能を測ることになる。課題の内容に従って文脈を選ぶと、環境自体へ課題情報を埋め込み、実際の職場では必要な情報の所在を探す作業を、あらかじめ一部済ませてしまう可能性がある。 著者らは、組織の状態と課題の指定を分離する評価基盤WorkWorldsを導入する。まず改訂版、日付、従業員の役割を固定し、その従業員がアクセスできる組織の状態を具体化する。課題はその後で与える。主な実装は架空の製薬会社で、六つの従業員の役割にまたがる八つの測定課題を持ち、追加の組織世界も構築する。 対応をそろえた192件の評価では、課題ごとに文脈を選ぶと、必要な証拠にアクセスできた割合が72.8%から90.4%へ17.6パーセントポイント上がり、評価基準を満たした割合は68.0%から76.7%へ8.7ポイント上がった。一方、証拠にアクセスできた場合に限った合格率はほとんど変わらなかった。測定された差の大部分は、エージェントが十分な証拠に到達する前の段階で生じた。
v2の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-23 · v2
- 査読・掲載
- 査読状況未確認
更新履歴
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for each task selected together with or after the task has been specified. This design measures performance on workplace-like tasks in an environment assembled for the task. When task specification guides which context is selected, the evaluation can encode task information into the environment and pre-complete part of the information-localization work that workplace performance normally requires. We introduce WorkWorlds, an evaluation infrastructure that separates organizational state from task specification. A world first fixes a revision, date, and employee seat and materializes the organizational state that employee can access; tasks are introduced only afterward. We implement WorkWorlds in a primary synthetic pharmaceutical company with 8 measured tasks across 6 employee seats, and construct additional organizational worlds. Across 192 matched evaluations, task-level curation increased evidence access by 17.6 percentage points, from 72.8% to 90.4%, and criterion pass by 8.7 points, from 68.0% to 76.7%, while pass conditional on evidence access remained nearly unchanged; most of the measured difference occurred before the agent reached sufficient evidence.
著者のコメント
4 figures
arXiv ID: 2609.23806 / 要約の誤りについて