企業内の隠れた事実を調べるエージェント評価課題
Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge
この論文をやさしく読む
ひとことで言うと
企業内の記録を横断して隠れた事実を推論する質問を加え、エージェントの能力差を測り直した。
何に役立つ?
業務システム内の複数記録を調べるエージェントを評価する際、表面上の記録だけでは答えられない課題として使える。
この研究の面白いところ
見つかりやすい記録が誤解を誘い、別の記録から事実を推論しなければならない設計で、従来の高得点モデル間の差が大きく現れた。
どこまで分かった?
結果は生成した企業データ、八つの質問ひな型、12のモデルとプログラムの組み合わせでの評価に基づく。実企業の業務での性能は要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Era by Eonベンチマークでは、各質問が答えを求める規則を明示し、生成された企業のデータからコードで正解を計算する。エージェントがコードを実行できる場合、上位四つのモデルはいずれも27問中22~25問に正答し、モデル間の差がほとんど見えなくなった。本研究は、隠れた事実に依存する八つの質問ひな型を追加する。質問にも文書にも隠れた事実は直接記されず、それを持っていそうな記録には別の内容が書かれているが、他のデータから推論できる。例えば、営業システムには顧客が時期を理由に購入をやめたと記録されているのに、録音された通話では顧客が障害を理由に挙げている。 生成した企業ごとに、コードで各ひな型にデータを埋め込み、言語モデルを使わず正確な答えを計算する。12のエージェントを評価し、それぞれはモデルと、企業のシステムに接続するエージェントプログラムの組み合わせからなる。最良のエージェントは、各質問に3回ずつ挑む計24回のうち18回で正答した。六つのモデルのうち四つは、どのプログラムと組み合わせても24回中6回以下だった。最も難しい質問は、顧客が三つの更新提案のうちどれに署名したかなど、似た複数の記録から一つを選ぶ必要があった。これら二つの質問について、全エージェントの正答は計84回の試行でわずか1回だった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them. We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a purchase because of timing. On a recorded call, the customer blames an outage. For each generated company, code fills each template and computes an exact answer without a language model. We evaluate 12 agents. Each pairs a model with an agent program, which connects it to the company's systems. The best agent answers 18 of its 24 attempts, three per question, correctly. Four of the six models answer at most 6 of 24 with any program. The hardest questions require picking one of several similar records, such as which of three renewal offers a customer signed. All agents together answered two such questions correctly in only 1 of 84 attempts.
著者のコメント
9 pages
arXiv ID: 2609.30055 / 要約の誤りについて