arXiv論文メモ
新着一覧
cs.CL / cs.AI / cs.CR · 査読状況未確認

Kali Linuxの操作指示を正しいコマンドへ変換できるか評価

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer

この論文をやさしく読む

ひとことで言うと

自然言語の依頼を、実際に実行できるセキュリティツールのコマンドへ正確に直せるかを測る評価セットです。ツール選択だけでなく、引数やオプションの組み立ても調べます。

何に役立つ?

セキュリティ作業支援LLMのコマンド生成能力の評価と学習に使うことが想定されています。コマンドの形式的な正しさを細かく評価できる点が目的です。

この研究の面白いところ

評価用データを作る際はサンドボックス実行も使う一方、学習時の報酬は実行なしで検証できるようにしています。大きなモデルだけでなく、小さなモデルの追加学習による改善も比較します。

どこまで分かった?

42%の上限は評価した重み公開モデルの制限なし設定に関する結果です。8Bモデルと685B MoEモデルの比較もこの評価の範囲であり、実際のセキュリティ作業全体の能力が同じとは示していません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

LLMはサイバーセキュリティの作業に使われることが増え、分析者の意図をツールの呼び出しに変換する役割が期待されている。しかし既存の評価は知識に基づく試験や、エージェントが最初から最後まで行うタスクに重点を置き、実際のセキュリティツールに対して実行可能なコマンドを生成する能力を直接測っていない。この隔たりは重要である。セキュリティ操作は厳密なコマンドラインインターフェース(CLI)に依存し、小さな構文誤り、オプションと値の対応の誤り、引数の順序違いでも実行が無効になるためである。 本研究では、Kali Linuxにおける自然言語からCLIへの変換を細かく評価するベンチマーク兼データセットKaliBenchを導入する。23の能力軸と5つのセキュリティ段階にわたる1,642のツールを対象に、8,504組の質問・コマンド対を含む。KaliBenchは文書を根拠とする処理手順で構築され、決定的な正規化と別名を考慮した評価により、ツール選択と引数構成を正確かつ再現可能に評価する。意味的な正しさと実際の実行可能性の両方を確保するため、LLMによる検証、サンドボックス内の端末実行、人間が関与する修正を組み合わせた多段階の検証手順を開発する。この細粒度で決定的な信号を基に、KaliBenchはさらに、学習のための実行不要な検証可能報酬を提供できる。 汎用およびセキュリティ特化の重み公開モデルについて、三つの評価モードと24の構成を調べたところ、制限のない設定でコマンド完全一致率42%を超える重み公開モデルはなかった。これは、明示的なツールのヒントなしでCLIベースのセキュリティツールを正確に使うことの難しさを示している。さらに、KaliBenchから得られる検証可能報酬を用いた教師ありファインチューニングと強化学習が、80億パラメーターのモデルを大幅に改善し、6,850億パラメーターのMoEモデルに匹敵する性能を達成することを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.

著者のコメント

Accepted at NeurIPS 2026 Evaluations and Datasets Track. Project page: https://risys-lab.github.io/KaliBench/ | Github: https://github.com/RISys-Lab/KaliBench

arXiv ID: 2610.02206 / 要約の誤りについて