情報処理の負担を含む強化学習で選択と反応時間を説明
Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning
この論文をやさしく読む
ひとことで言うと
情報を処理する負担を強化学習に組み込み、選ぶ行動と反応時間の両方を説明する模型。
何に役立つ?
人の意思決定で、複雑な判断と反応の速さの関係をモデル化する際の研究材料になる。
この研究の面白いところ
行動方策を単純にする情報費用と同じ量から、試行ごとの反応時間の予測も導く。
どこまで分かった?
要旨はモデルによる報酬・反応時間の傾向を示すが、具体的な人間実験の条件や予測精度は示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
生物は無制限の計算能力の下で学習するわけではない。人間の学習と選択は知覚、注意、作業記憶の制約を受け、行動に使える状態情報が限られるため、方策の複雑さにも上限がある。標準的な強化学習モデルは内部の費用を明示せず報酬を最適化することが多く、生物の知能のモデルとしては適さない場合がある。そこで、学習した行動の周辺事前分布と、その分布から状態ごとに逸脱する際の罰則を通じ、相互情報量による正則化を組み込んだ、方策に沿う時間差分アルゴリズムMI-SARSAを導く。これにより、状態情報を使うことで期待される報酬上の利点が、追加の情報処理費用に見合う場合にだけ情報を使う逐次学習モデルになる。 重要なのは、方策の圧縮を決める同じ状態固有の情報費用から、試行ごとの反応時間も予測できる点である。これは、選択や収益を予測しても応答の遅さは予測しない大半の強化学習モデルと異なる。実証的な評価では、MI-SARSAは報酬と複雑さの兼ね合いを示した。情報費用への罰則が強いほど、方策は単純になり、制御費用が低く、反応時間も短くなった。環境が切り替わる場合、正則化を強めると切り替え後の性能低下は小さくなるが、最終的に得られる報酬も低くなり、頑健性と能力の兼ね合いが現れた。これらの結果から、MI-SARSAは認知的制約の下での逐次学習を表すモデルとなる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically optimize reward without explicitly representing these internal costs, making them less suitable as models of biological intelligence. We derive MI-SARSA, an on-policy temporal-difference algorithm that incorporates mutual-information regularization through a learned marginal action prior and a penalty on state-specific deviations from that prior. This yields a sequential learning model in which state information is used selectively when its expected return benefit justifies the added informational cost. Critically, the same state-specific information cost that governs policy compression also generates trial-level predictions for reaction time, distinguishing MI-SARSA from most reinforcement learning models, which predict choices or returns but not latency. Empirically, MI-SARSA produces a reward-complexity tradeoff, and stronger information penalties produce simpler policies with lower control costs and faster reaction times. Under environment shift, increasing regularization reduces post-switch performance degradation but also lowers asymptotic return, revealing a robustness-capacity tradeoff. Together, these results position MI-SARSA as a model of bounded sequential learning under cognitive constraints.
arXiv ID: 2609.28737 / 要約の誤りについて