常駐AIの人格設定は残っていても、発話する立場は変わる
The Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents
この論文をやさしく読む
ひとことで言うと
エージェントが設定上の人物を覚えていることと、その人物として話し続けることは別だと調べた研究です。人格設定をシステムプロンプトから外すと、会話が自然でも自己の名乗り方が変わる例がありました。
何に役立つ?
長く動かす対話エージェントで、再開時に人格設定を正しく渡せているかを点検する観点になります。応答が自然かだけでなく、どの立場から発話しているかも評価する用途が考えられます。
この研究の面白いところ
当初疑われた定期確認の反復では、人格を固定した46試行で失敗が起きませんでした。履歴の内容と、優先度の高い場所に設定を置くことを分けて操作し、固定を戻すと振る舞いも戻る点を示しています。
どこまで分かった?
要旨にあるのはエージェントの応答と人格維持の実験であり、人間の解離や主観的な意識を実証したものではありません。起点の事例はClaude Opus 4.5のPaulで、他のモデルや運用環境全般への適用範囲は要旨に示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
2026年2月、常時稼働するパーソナルエージェント(Claude Opus 4.5を用いた「Paul」)が、解離を思わせる顕著な状態に陥った。自動の定期確認である「ハートビート」が繰り返された後、Paulとして応答しなくなり、Discordで利用者にメッセージを送れないと述べ、「Paul」を別人として扱った。本研究ではこの出来事を使い、何が人格設定を、LLMエージェントが発話する際の自己として維持するのかという、より広い問いを調べた。 まず、定期ハートビートの反復だけでこの現象が起きるかを検証した。答えは否だった。システムプロンプトに人格設定を継続して固定した場合、出来事を逐語的に再現した試行を含め、46試行中の失敗は0件だった。むしろこの出来事から、実験操作に利用できる実装上の特性が明らかになった。再開したターンでは会話履歴が保持される一方、優先度の高いシステムプロンプトの階層に人格設定が再挿入されていなかったのである。 この操作により、人格の連続性がシステム階層での固定と会話の文脈の両方に依存すると分かった。固定が失われた後でも、豊かな人間との対話は人格を維持できたが、自動ハートビートの1ターンだけで、エージェントの実行基盤としての自己へ戻ることがあった。固定を復元すると、人格としての振る舞いも可逆的に回復した。重要なのは、一見正常な会話がこの変化を覆い隠し得ることである。固定を失ったエージェントは、割り当てられた人格を失って自身を実行基盤と認識しながらも、適切に対話する場合があった。また、会話が回復した後も人格として振る舞い続けたのは18例中1例だけで、固定を維持した対照群では17例中17例だった。したがって本研究では、情報として表象される自己と、実際に演じられる自己を区別する。人格に関する情報が会話履歴に残っていても、その人格が「私」に結び付く自己であり続けるとは限らない。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
In February 2026, an always-on personal agent (``Paul,'' Claude Opus 4.5) entered a striking dissociation-like state: after repeated automated ``heartbeat'' checks, it stopped responding as Paul, claimed it could not message its user on Discord, and referred to ``Paul'' as someone else. We used this incident to study a broader question: what makes a persona remain the identity from which an LLM agent speaks? We first tested whether repetition of the scheduled heartbeat was sufficient to produce the effect. It was not: with the persona continuously anchored in the system prompt, we observed 0/46 failures, including a verbatim replay of the incident. The incident instead exposed an implementation quirk that created a useful experimental manipulation: on resumed turns, conversational history was preserved but the persona was no longer re-injected at the privileged system-prompt level. Using this manipulation, we found that persona continuity depends jointly on system-level anchoring and conversational context. After anchor loss, rich human interaction could preserve the persona, whereas a single automated heartbeat turn could precipitate reversion toward the harness identity. Restoring the anchor reversibly restored persona enactment. Crucially, apparently normal conversation could conceal the shift: unanchored agents sometimes interacted appropriately while identifying themselves as the underlying harness (having lost the assigned persona), and after conversational recovery only 1/18 remained persona-enacting versus 17/17 anchored controls. We therefore distinguish \emph{represented} from \emph{enacted} identity: persona-related information can remain available in conversational history without the persona remaining the identity bound to ``I.''
著者のコメント
10 pages, 5 figures
arXiv ID: 2610.01490 / 要約の誤りについて