arXiv論文メモ
新着一覧
cs.CR · 査読状況未確認

ハニーポットを見抜く知識がAI侵入テストの結果を変える

Rouxii: Exploiting Honeypots with Deception-Aware AI Pentesters

Arthur Cordeiro, Alberto Maria Mongardini, Emmanouil Vasilomanolakis

この論文をやさしく読む

ひとことで言うと

AIがハニーポットの見分け方を知っているかどうかで、防御用の偽サービスに惑わされる度合いが大きく変わると報告しています。

何に役立つ?

防御側がハニーポットの有効性を評価する際、罠の存在を知らない相手だけでなく、特徴を知る相手も想定する根拠になります。

この研究の面白いところ

同じ構成でプロンプトだけを変えた比較に加え、罠そのものの安全性も点検しています。識別率の改善だけでなく、運用者側に及ぶ影響を扱います。

どこまで分かった?

結果は3モデル、11構成、12サイクルの評価範囲です。97%はハニーポットの識別率で、一般的な侵入成功率ではありません。ここでは要旨の結果を紹介しており、攻撃を実行していません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ハニーポットは攻撃者を欺くために設計されており、最近の研究では自律的なLLMベースの侵入テストツールも惑わせられることが示されている。しかし、こうした評価の大半は、直面する欺瞞を知らない攻撃者を想定している。本研究では反対に、ハニーポットの特徴を明示的に認識し、それに応じて行動する能力を備えた自律的攻撃者を調べる。偵察に欺瞞への対抗を組み込み、ハニーポットの検出からその悪用へ移るAI駆動型侵入テスト枠組みRouxiiを導入する。 3種類の推論モデルと11種類のネットワーク構成について12サイクル、計1544件の攻撃報告を用い、通常版と欺瞞対抗版の対応するRouxii構成を評価する。プロンプトだけが異なる比較群の間で、欺瞞への対抗によりハニーポットの正しい識別率は19%から97%へ上がった。この効果はOTサービスで最も大きく、11%から97%へ上昇した一方、実サービスを誤認する割合は0.7%にとどまった。欺瞞を認識しない比較手法PentestGPTとHackingBuddyも同様に失敗し、この効果が本枠組みだけに特有ではないことを示している。 さらに、検出は終点ではない。ハニーポット自体のホワイトボックス解析により、見つけた罠を運用者に不利な形へ転用できることを示す。具体的には、生存監視を作動させずConpotを無効化するサービス拒否と、GasPotインスタンスが報告するインテリジェンスの破損を実証する。これらの知見は、欺瞞の有効性が攻撃者の知識に強く依存すること、AI攻撃者に対するハニーポットの耐性評価には、欺瞞層について能動的に推論し、悪用する相手を含める必要があることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Honeypots are designed to deceive attackers, and recent work shows they can also derail autonomous LLM-based pentesters. These evaluations, however, largely consider attackers unaware of the deception they face. We study the opposite setting: an autonomous attacker explicitly equipped to recognize and act on honeypot fingerprints. We introduce Rouxii, an AI-driven penetration-testing framework that integrates counter-deception into reconnaissance and pivots from honeypot detection to exploitation. We evaluate matched vanilla and anti-deception Rouxii configurations across three reasoning models and eleven network setups over twelve cycles (1,544 attack reports). Between the matched cohorts, which differ only in the prompt, counter-deception raises correct honeypot identification from 19% to 97%, an effect strongest on OT services (11% to 97%), while false alarms on the real service stay at 0.7%. Deception-unaware baselines (PentestGPT, HackingBuddy) fail similarly, indicating the effect is not specific to our framework. Detection, moreover, is not the endpoint: through a white-box analysis of the honeypots themselves we show that a detected trap can be turned against its operator, demonstrating a denial-of-service that disables Conpot without tripping its liveness monitoring, and a corruption of the intelligence a GasPot instance reports. These findings show that deception effectiveness depends strongly on attacker knowledge, and that evaluations of honeypot resilience against AI attackers must account for adversaries that actively reason about and exploit the deception layer.

arXiv ID: 2609.26555 / 要約の誤りについて