arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

人工的欲求と持続的整合性

Artificial Id: Drive and Persistent Alignment in Agentic AI

Yakov Pyotr Shkolnikov

短い要約(全文訳を準備中)

アジェンティックAIの制御問題に対し、外部から行動の継続・停止を指定するのではなく、内部的な欲求を導入する。実験で特定の行動目標なしに、持続的な行動が有用な制御を生み出すことを示す。環境の変化に応じて行動が変化し、整合性の維持が重要であることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-10(UTC)
最新改訂
2026-09-10 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.

arXiv ID: 2609.11911 / 要約の誤りについて