物理シミュレーションでLLMの長期灌漑制御を検証
Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control
この論文をやさしく読む
ひとことで言うと
LLMの灌漑提案を物理シミュレーションで検査し、繰り返す失敗から提案の仕方だけを改善する仕組みです。
何に役立つ?
考えられる用途は、生育期を通じた灌漑計画の支援です。日々の判断が将来に与える影響を数値モデルで確かめます。
この研究の面白いところ
提案を修正する短期の仕組みと、失敗を文脈として残す長期の仕組みを分け、物理モデルや実行制約は固定しています。
どこまで分かった?
約51%の削減は共通の事後評価器による、過去の予定の再生との比較です。実圃場の前向き運用で同じ節水率を実証したとは要旨に書かれていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)エージェントは、推論、道具の使用、行動を組み合わせるようになっているが、証拠の多くは、比較的すぐにフィードバックが得られ、失敗後にリセットできる単発の課題から来ている。長期の物理制御では事情が異なる。行動が将来の状態を変え、誤りが判断を重ねるにつれて蓄積する一方、エージェントは安全な実行を支える物理規則を書き換えることなく、経験から改善しなければならない。本研究では、日々の判断が生育期全体の土壌水分動態と相互作用する灌漑を通じて、この設定を調べる。 二つの修正時間尺度を中心に構成した、物理に根拠を置くLLMエージェントMimirを提案する。速い時間尺度では、構造化された物理インターフェースと決定論的シミュレータがLLMの出力を提案として扱い、実行前に数値的に検査・修正し、制約付きの決定論的な行動選択に通す。遅い時間尺度では、繰り返す失敗パターンを永続的な文脈上の原則にまとめ、将来の提案を条件付ける。この間、物理モデル、評価器、実行制約は変更しない。 複数の地点、作物、年にわたる共通の事後評価器のもとで、Mimirは評価対象の比較手法中で最も低い総制御コストを達成し、過去の予定を再生する方法より灌漑量を約51%削減する。構成要素を除く実験では、先読みシミュレーション、検証付き修正、永続的文脈のいずれを除いても制御コストが上昇する。モデル規模とモデル系列の検討では、LLMを大きくしても単調な改善は得られない。得られた教訓は、持続的に動く物理エージェントが、物理的な真偽の判断とアクチュエータの権限を明示的な数値機構に委ねながら、意味的推論と、制約下で証拠に基づく自己改善を組み合わせられるということである。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an agent must improve from experience without being allowed to rewrite the physical rules that make execution safe. We study this regime through irrigation, where daily decisions interact with soil-water dynamics over entire growing seasons. We present Mimir, a physics-grounded LLM agent organized around two repair timescales. At the fast timescale, a structured physical interface and deterministic simulator turn an LLM output into a proposal that we numerically check, revise, and subject to bounded deterministic action selection before execution. At the slow timescale, recurrent failure patterns are consolidated into persistent contextual principles that condition future proposals, while the physical model, evaluator, and execution constraints remain immutable. Under a common retrospective evaluator across multiple sites, crops, and years, Mimir attains the lowest reported aggregate control cost among the evaluated references and uses about 51% less irrigation than the historical schedule replay. The ablation study show higher control cost when forward simulation, verified revision, or persistent context is removed; model-scale and model-family studies show no monotonic gain from increasing LLM size. The resulting lesson show that persistent physical agents can combine semantic reasoning with bounded, evidence-driven self-improvement while reserving physical truth and actuator authority for explicit numerical mechanisms.
arXiv ID: 2610.02038 / 要約の誤りについて