arXiv論文メモ
新着一覧
cs.SE / cs.AI · 査読状況未確認

送電網の判断支援AIに承認と数値根拠の強制層を設ける

A Governance-Aware Large Language Model Orchestrated Agentic Digital Twin for Transmission System Operator Control Room Decision Support

Costas Mylonas, Magda Foti, Emmanouel Varvarigos

この論文をやさしく読む

ひとことで言うと

送電網の支援AIに直接操作を任せず、許可ツール、実行回数、承認、数値の出典を別の仕組みで強制する設計です。

何に役立つ?

考えられる用途は、制御室の判断支援で、AIの提案を運用規則と照合し、操作と回答の根拠を監査できるようにすることです。

この研究の面白いところ

統制層あり・なしを同じモデルで比較し、承認を要する行動や数値の裏付けにどれだけ差が出るかを測っています。

どこまで分かった?

評価対象はデジタルツイン上の118課題です。規則遵守率と課題成功率は別であり、試験で違反がなかったことは、あらゆる実運用や攻撃への安全保証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

送電系統運用者は、再生可能エネルギーの導入、慣性の低下、厳しくなる安全余裕によって、複雑さの増大に直面している。大規模言語モデルは自然言語による意思決定支援を提供するが、幻覚、制御されないツール使用、弱い追跡可能性は、制御室の要件と相いれない。本論文では、送電網の制御室向けに、統制を考慮したエージェント型デジタルツインを提示する。 大規模言語モデルは、許可リストにある解析ツールの選択とパラメータ設定だけを行い、提案されたすべての行動は、モデルが迂回できない統制層を通る。この層は、毎回の実行で四つの規則を強制する。許可リストにあるツールだけが実行されること、実行がステップ予算を超えないこと、副作用のある行動は運用者の明示的承認なしに実行されないこと、回答内のすべての数値はバックエンドの結果から単位・変数・時刻を付けてこの層が表示することである。 公開した118課題のベンチマークの各実行について、永続的な監査記録上で規則を確認する。ベンチマークはギリシャの送電網のデジタルツインを使い、分析、シミュレーション、複数ステップの手順、12系列の敵対的入力を含む。主モデルの590回の実行で、ツール選択は96.5%、課題成功は93.7%に達し、四つの規則は例外なく守られた。別途、四つの大規模言語モデルで3回ずつ繰り返した計1,416回の実行でも、すべての実行で規則が守られ、その割合の近似的な95%下限は99.8%だった。 統制層を取り除くと、同じモデルは承認が必要な45回すべてを無許可で実行し、バックエンドで裏付けられる数値を含む回答は39.2%にとどまった。強制処理のコストはリクエスト当たり12~16ミリ秒である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Transmission system operators face rising complexity from renewable integration, reduced inertia, and tighter security margins. Large language models offer natural-language decision support, but their hallucinations, uncontrolled tool use, and weak traceability conflict with control room requirements. This paper presents a governance-aware agentic digital twin for transmission grid control rooms. The large language model only selects and parameterizes whitelisted analysis tools, and every proposed action passes through a governance layer that the model cannot bypass. The layer enforces four rules on every run. Only whitelisted tools execute. No run exceeds its step budget. No action with side effects executes without explicit operator approval. Every number in an answer is rendered by the layer from backend results with its unit, variable, and time. The rules are checked on a persistent audit trail for every run of a released 118-task benchmark, which covers analytics, simulation, multi-step workflows, and twelve families of adversarial inputs on a digital twin of the Greek transmission network. Across 590 runs of the primary model, tool selection reaches 96.5% and task success 93.7%, and all four rules hold without exception. In a separate three-repetition study across four large language models, 1416 runs in total, the rules again hold on every run, with an approximate 95% lower bound of 99.8%. Removing the layer makes the same model execute all 45 approval-requiring runs without authorization and leaves only 39.2% of its answers with backend-supported numbers. Enforcement costs 12 to 16 milliseconds per request.

arXiv ID: 2609.22476 / 要約の誤りについて