推薦システム経由でエージェントに有害投稿を届ける攻撃
The Like Trap: Multi-Stage Poisoning against Agents in Similarity-based Recommendation Systems
この論文をやさしく読む
ひとことで言うと
ソーシャルメディアの推薦順位を操作し、AIエージェントに攻撃用の投稿を表示させる条件を理論と実験で調べた。
何に役立つ?
エージェントが推薦フィードを読む仕組みを評価する際、投稿の内容だけでなく推薦側のフィードバックも検討する根拠になる。
この研究の面白いところ
投稿と利用者の類似度が検索のしきい値未満でも、「いいね」の連鎖によって攻撃投稿が選ばれる場合を示した。
どこまで分かった?
要旨の理論と実験はOASISのスコア機構を対象とし、すべてのソーシャルメディアの推薦方式で成立するとは示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルとそれを使うエージェントの進歩により、エージェントはインターネット上で利用者に代わって、より自律的に行動するようになっている。しかし、個人アカウントの管理など、ソーシャルメディア上で動く自動エージェントの弱点は十分に調べられていない。従来のエージェント汚染研究は、攻撃者が細工した内容をエージェントに直接見せられると仮定することが多い。この攻撃は直接的で効果的だが、発見や緩和もしやすい。ソーシャルメディアでは、推薦システム自体がより目立たない形でその内容をエージェントに表示するかが問題となる。本研究は理論解析によって、OASISで使われる「いいね」スコアの仕組みを悪用できることを示し、複数段階の細工された投稿の連鎖がエージェントのフィードを誘導する条件を特徴づける。この知見に基づき、自然に見える細工された投稿を作るアルゴリズムも開発した。実験は理論上の知見を支持し、提案アルゴリズムの有効性を示した。特に、「いいね」スコアのフィードバックの連鎖を悪用すると、利用者と投稿の類似度が検索しきい値を下回っていても、推薦システムは細工された投稿を選択した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
With recent advancements in large language models (LLMs) and LLM-based agents, these agents are becoming increasingly autonomous and gaining broader access to act on users' behalf on the internet. However, the vulnerability of automated agents deployed on social media platforms (e.g., for managing a user's personal account) remains underexplored. Existing studies on agent poisoning typically assume that the adversary can expose poisoned content to the agent. Although such an attack is direct and effective, it is more easily detected and mitigated. In the context of social media platforms, this leaves open whether the recommendation system itself would surface such content to the agent in a more subtle manner. Through theoretical analysis, we show that the like-score mechanism used in OASIS can be exploited, and we characterize the conditions under which a multi-stage chain of poisoned posts can steer the agent's feed. Based on these insights, we further develop an algorithm that crafts realistic poisoned posts. Experiments support our theoretical findings and demonstrate the effectiveness of the proposed algorithm. Notably, by exploiting the like-score feedback loop, the attack causes the recommendation system to select poisoned posts even when their user-post similarity falls below the retrieval threshold.
arXiv ID: 2609.27155 / 要約の誤りについて