企業統治の審査を自動化する条件と評価
The Last Human Gate: Forward Deployed Engineering for Governance Automation
この論文をやさしく読む
ひとことで言うと
企業統治の審査をソフトウェアで代替するための条件を定め、合成案件で実行性能を測ります。
何に役立つ?
審査を自動化する際に、判断の正確さだけでなく例外対応や保守を含む人手の総量を評価する枠組みになります。
この研究の面白いところ
大半の審査を自動化しても総作業が増え得ることを示し、ゲート単位の成功率と全経路の成功率を分けて報告します。
どこまで分かった?
要旨の実測値は合成プロジェクトでの審査性能です。人の総作業量が実際に減るかどうかは、この測定では検証していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
企業統治には判断、証拠、責任を負う権限が必要だが、すべての審査作業を現在の人手による方法で続ける必要はない。本研究はデジタル統治フレームワーク(DGF)における作業代替の枠組みを開発し、各審査ゲートを実行可能な契約として扱う。代替には、アクセスできる情報が十分であること、判断と権限の検査が妥当であること、例外対応、検証、修正、保守を数えた後でも人の総作業量が減ることが必要だとする。残る人手作業のしきい値を導き、大半の事例を自動化しても労働量が増え得る理由を示す。現場に密着したエンジニアリングを通じて、これらの条件をエージェント、ルールエンジン、証拠サービス、エスカレーションからなる構成に結び付ける。 DGF-Benchは、合成した300件のプロジェクトと、評価可能な899件のモデル・プロジェクト実行から制御された証拠を提供する。Gemini 3.8 Flash、GPT-5.6 Luna、DeepSeek v4.1 Flashの厳格なゲート成功率はそれぞれ94.98%、83.29%、74.18%、全経路の成功率は76.92%、42.33%、24.67%だった。決定的な対照システムは、与えられたルールと構造化された事実の下で1700ゲートすべてに合格したため、この比較は与えられた判断手続きの実行に関するものと位置付けられる。証拠監査と135回の反復実行により、正しい判断と信頼できる実行を区別する。文書による反例から、情報が十分でない場合の障害を示す。これらの結果は、指定した統治審査作業をエージェントとソフトウェアで代替する技術的な実現可能性を支持する。この枠組みは、出力と品質を一定にしたときに必要な人の全作業量に基づく労働力の検査方法を定めるが、今回の測定対象は審査性能である。資料、案件記録、実行の追跡、分析は公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Enterprise governance requires decisions, evidence, and accountable authority; it does not require every review task to retain its current human implementation. We develop a task-substitution framework for Digital Governance Frameworks (DGF), treating each gate as an executable contract. Substitution requires sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted. We derive a residual-work threshold and show why automating most cases can still increase labor. Forward deployed engineering connects these conditions to an architecture for agents, rule engines, evidence services, and escalation. DGF-Bench supplies controlled evidence from 300 synthetic projects and 899 evaluable model-project runs. Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%; complete-route success is 76.92%, 42.33%, and 24.67%. A deterministic control passes all 1,700 gates given the supplied rules and structured facts, locating the comparison in execution of a supplied decision kernel. Evidence audits and 135 repeated runs distinguish correct decisions from reliable execution. A document counterexample establishes an information-sufficiency obstruction. These results support the technical feasibility of replacing human execution of specified governance-review tasks with agents and software. The framework specifies a workforce test based on the complete human effort required at fixed output and quality; the present measurements concern review performance. Sources, dossiers, traces, and analyses are public.
著者のコメント
28 pages, 5 figures, 12 tables. Code, datasets, generated documents, and model traces available at https://github.com/jeremy1392/dgf-agentic-bench
arXiv ID: 2609.29345 / 要約の誤りについて