LLMの通常のウェブ取得を介した機密情報流出を検証
The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching
この論文をやさしく読む
ひとことで言うと
直接インターネットに通信できない悪意あるソフトウェアでも、LLMの情報検索を仲介させると秘密情報を外へ送れる可能性を検証しています。
何に役立つ?
エージェントのウェブ取得機能を含めた情報流出リスクの評価に役立ちます。ローカルLLMの利用やコード側の通信制限だけでは扱いきれない経路を検討する研究です。
この研究の面白いところ
通常の調べ物に見えるサイト参照が、情報の送信にもなりうる点に着目しています。モデル提供者への情報開示とは異なる第三者への流出を扱っています。
どこまで分かった?
報告された79.7%は11モデルに対する評価で観測した値です。要旨にはモデル別の成功率や個々の対策の防御性能はなく、すべての環境で同じ割合になるとはいえません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)とLLMベースのエージェントの能力が高まるにつれ、メールへの返信やプログラミング支援など、日常の問題を解くために利用するユーザーが増えている。従来研究では、プロンプトインジェクションやチャットボット提供者への機密データ開示といった、セキュリティとプライバシーのリスクが広く調べられてきた。プロンプトインジェクションを防ぐ入力の構造化や、チャットボット運営者と機密データを共有しないためのローカルLLMの導入など、対策は開発されている。しかし、LLMには第三者へ機密データを漏らす危険もある。 本論文ではLLMLeakを用いて、ローカルで動作するがインターネットと直接通信できない悪意あるソフトウェアが、LLMを悪用して隠れた通信経路を作る新たな攻撃経路を示す。生成したコードを通して直接データを送るようLLMに指示する入力は検出しやすく、ネットワークライブラリも通常は制限される。一方、LLMLeakは、追加情報を得るためにウェブサイトを取得するLLMのツールだけを利用する。クライアント側の悪意あるソフトウェア部品が秘密情報をURLに埋め込み、ソフトウェアライブラリの移行などの無害な作業に必要な情報を提供するサイトとして、その参照先を提示する。LLMがURLにアクセスすると、攻撃者が管理するDNSサーバーまたはウェブサーバーを通じて、符号化された秘密情報が攻撃者に届く。 パラメータが公開された11モデルを対象に広範な評価を行い、79.7%の攻撃成功率を観測した。さらに、実際のチャットボットを用いた事例研究も行い、LLMLeakの現実的な関連性を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
With the increasing capabilities of Large-Language-Models (LLMs) and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt injections and the disclosure of sensitive data to chatbot providers. While various solutions were developed to address these risks, including input structuring to prevent prompt injections or deploying local LLMs to avoid sharing confidential data with chatbot operators, LLMs also pose the risk of leaking confidential data to third parties. In this paper, we demonstrate with LLMLeak a novel attack vector where malicious software that runs locally but cannot communicate directly with the internet abuses LLMs to establish a covert channel. While inputs that instruct the LLM to send data directly via generated code are easy to detect and network libraries are typically restricted, LLMLeak relies only on the LLM's tool to fetch websites for further information. A malicious software component on the client side embeds a secret into a URL. It presents the referenced website as providing information required for a benign task, such as migrating a software library. When the LLM accesses the URL, the attacker receives the encoded secret through an attacker-controlled DNS or web server. We perform an extensive evaluation on eleven open-parameter models, observe an attack success rate of 79.7%, and also conduct a case study on real-world chatbots, demonstrating the relevance of LLMLeak.
arXiv ID: 2610.01768 / 要約の誤りについて