観測行動から制約付き線形二次ガウスゲームを逆推定
Inverse Linear Quadratic Gaussian Games: Constrained Setting and Transferability
この論文をやさしく読む
ひとことで言うと
複数の主体が選んだ行動から、その背後にある費用の重みを推定し、条件が変わっても使えるか調べます。
何に役立つ?
観測した均衡行動を再現する費用モデルの同定や、近い力学条件への方策の移転可能性を評価する理論的な手段になります。
この研究の面白いところ
制約付きでは双対変数まで求め、制約なしでは力学や推定値のずれに応じた性能低下を評価します。交通シミュレーションと実ロボット実験も行っています。
どこまで分かった?
移転可能性の実験で示されたのは、十分に近い力学系への適用です。大きく異なる力学条件での性能は要旨からは分かりません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
本研究は、有限の時間範囲を持つ逆線形二次ガウスゲームを扱う。制約がある場合について、与えられた一般化ナッシュ均衡を生み出す費用関数のパラメータと最適な双対変数の値の集合を特徴付け、それらを計算するアルゴリズムを提案する。制約のない場合には、別の力学系へ結果を移せるかを扱う。具体的には、同定した費用パラメータによる方策と、熟練者のパラメータによる方策について、異なる力学条件での費用値の差を上から評価する。この費用値の差は、力学と同定した費用パラメータのずれに比例して増える。 数値シミュレーションでは、制約付きの場合に、提案アルゴリズムが、観測された一般化ナッシュ均衡に対応する方策と軌道を再現できる費用パラメータと双対変数の値を同定した。制約なしの場合には、交通シミュレーションと実ロボット実験により、同定した費用パラメータが十分に近い力学系の制御に利用でき、性能は力学と同定パラメータのずれに比例して低下することを示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
This work addresses finite-horizon inverse linear quadratic Gaussian games. In a constrained setting, we characterize the set of cost parameters and optimal dual values that generate a given generalized Nash equilibrium, and we propose an algorithm to compute these parameters. In an unconstrained setting, we address transferability, namely, we bound the cost value perturbation between two policies: one induced by the identified cost parameters, the other by the expert parameters, under a set of different dynamics. This cost value perturbation scales linearly with the deviations in the dynamics and the identified cost parameters. Through numerical simulations, we show that, in a constrained setting, our algorithm identifies the cost parameters and dual values that can reproduce the policy and trajectories corresponding to the observed generalized Nash equilibrium. In an unconstrained setting, we show with a traffic simulation and real-robot experiments that the identified cost parameters can be used to control sufficiently close dynamics, with performance degrading linearly with the deviations in the dynamics and the identified cost parameters.
arXiv ID: 2609.29321 / 要約の誤りについて