自律的な安全性テストで同じ操作を減らす実行順制御
The Fly That Stopped: Mushroom-Body-Inspired Habituation as a Reward-Free Scheduling Prior for Autonomous Penetration Testing
この論文をやさしく読む
ひとことで言うと
自律的な安全性テストで、同じツールとURLを繰り返し試す無駄を減らす順番制御の研究です。
何に役立つ?
限られた操作回数を重複操作に費やさないためのスケジューラー設計に役立ちます。
この研究の面白いところ
報酬を学習させず、失敗や問題なしの操作への『慣れ』を数えて、測定可能な対象では重複を減らしました。
どこまで分かった?
重複操作の削減を示した結果であり、脆弱性の発見数が増えた証拠ではありません。慣れの要素だけの効果も切り分けられていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
自律的な安全性テストのエージェントは、限られた操作回数の多くを、以前に選んだツールの繰り返しに費やすことがある。ショウジョウバエのキノコ体による新規性処理に着想を得た、報酬を使わない実行順の制御器を評価した。疎な状態表現と、構造によって分けたURLの種類およびツール群ごとに減衰する慣れのカウンターを組み合わせる。カウンターは、問題が見つからない結果やエラーを繰り返す操作に不利な値を与えるが、単一の報酬値から重みを更新しない。設計の動機となった条件を合わせた四つの試験では、報酬の集計誤りとツール障害によるループが見つかり、報酬を使う要素は、試した主要な結果について報酬なしの条件より改善しなかった。その後、事前登録した予備試験と二段階の確認試験で、同じツールとURLの組を繰り返し選ぶかを評価した。二段階目の確認試験では、選別した実験室の対象10件のうち、エラーの多い低速のクロスサイトスクリプティング試験2件を除いた8件で測定が可能だった。慣れを使う制御器は、同率の2組を除く比較対象6組すべてで操作の重複率を下げた。正確な片側検定のp値は0.015625で、最大の減少は60操作の上限内で重複51回から18回だった。これは測定可能だった予算固定の対象群における制御器全体の結果であり、慣れだけの効果を切り分けた結果でも、脆弱性発見が増えた結果でもない。補助的な研究では、1万3498個のニューロンと、重み5以上のシナプス結合50万1267本を持つMaleCNS由来の回路を調べた。固定した読み出しでは試した五つの局所可塑性手法のどれにも行動の選択性は見られなかった。読み出し側の可塑性では合成課題で条件付きの良い結果を得たが、生物学的な回路構造の優位性は示していない。対象群の範囲、残る入力の完全性への依存、内部でAIを支援に用いたレビュー手順も結果と併せて報告する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Autonomous security-testing agents can spend much of a fixed action budget repeating earlier tool selections. We evaluate a reward-free scheduler inspired by mushroom-body novelty processing in Drosophila. It combines sparse state encoding with decaying habituation counters over structural URL classes and tool families. The counters penalize repeated clean or error outcomes without updating weights from scalar reward. Four matched campaigns motivated this design by exposing reward-accounting errors and tool-failure loops; reward-driven components did not improve the tested primary outcomes over the reward-free MB condition. A pre-registered pilot and two confirmatory stages then evaluated repeated (tool, URL) selections. In the second confirmatory stage, 8 of 10 screened lab targets remained measurable after two error-heavy slow-XSS exclusions. The habituation-enabled scheduler lowered duplicate-action ratios in all 6 non-tied target pairs (exact one-sided p=0.015625), with two ties; the largest reduction was 51 to 18 duplicate steps within a 60-step budget. This is evidence for the complete scheduler on the measurable budget-hold population, not an isolated habituation ablation or a vulnerability-discovery gain. A complementary study on a 13,498-neuron MaleCNS-derived circuit (501,267 synaptic edges with weight at least 5) found no action selectivity from the five tested local-plasticity approaches under a fixed readout; readout plasticity produced qualified positive results in synthetic tasks without establishing a biological-topology advantage. We report the population bounds, remaining input-integrity dependencies, and an internal AI-assisted review protocol alongside the results.
著者のコメント
12 pages, 7 figures, 2 tables
arXiv ID: 2609.29126 / 要約の誤りについて