Arm Cortex-M7で暗号計算の周辺システムを調整しML-KEMを高速化
System-Level Optimization Beyond Cryptographic Kernels: An ML-KEM Case Study on Arm Cortex-M7
この論文をやさしく読む
ひとことで言うと
暗号の演算コードを最適化した後でも、メモリ配置や公開データの再利用など、動作するシステム全体を調整すればさらに速くできるか調べています。
何に役立つ?
Arm Cortex-M7でML-KEMを動かす実装において、命令最適化の次にどこを見直すかを考える材料になります。
この研究の面白いところ
補助的な公開状態を持たない構成と、公開データを再利用する構成を分けて評価しています。最大74.6%という大きな削減は後者の特定構成に対応します。
どこまで分かった?
結果はCortex-M7と評価した構成についてのサイクル数です。要旨には公開データ再利用の詳細な前提や追加メモリ量はなく、全環境で同じ高速化が得られるとは限りません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
組み込み向け耐量子暗号の最近の研究は、主に命令レベルの最適化に注目してきた。これには算術カーネルの改善、アセンブリの調整、レジスタ割り当て、命令スケジューリングが含まれる。本研究では、Arm Cortex-M7上のモジュール格子に基づく鍵カプセル化機構(ML-KEM)を事例として、メモリ階層の活用、密結合メモリへの配置、周辺機能の統合、クロック設定、決定論的な公開データの再利用から得られる追加の改善を調べる。 評価は、SLOTHYで最適化された最先端の実装を出発点とし、ML-KEMの3つのパラメータ集合すべてを対象とする。暗号アルゴリズムや標準化された通信形式を変更せず、補助的な公開状態を持たない評価構成では、サイクル数を最大2.5%削減する。選択した公開データ再利用構成では、カプセル化とデカプセル化のサイクル数を、それぞれ最大74.6%と58.8%削減する。 これらの結果は、算術カーネルの最適化後にも配備時の大きな改善余地が残ることを示し、周囲の実行システムも調べる2段階の方法論を動機付ける。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent work on embedded post-quantum cryptography has focused primarily on instruction-level optimization, including arithmetic-kernel improvements, assembly tuning, register allocation, and instruction scheduling. Using the Module-Lattice-Based Key-Encapsulation Mechanism (ML-KEM) on an Arm Cortex-M7 as a case study, we examine the additional gains available from memory-hierarchy utilization, tightly coupled memory placement, peripheral integration, clock configuration, and deterministic public-data reuse. The evaluation starts from a state-of-the-art SLOTHY-optimized implementation and covers all three ML-KEM parameter sets. Without modifying the cryptographic algorithm or standardized wire formats, the evaluated profiles without auxiliary public state reduce cycles by up to 2.5%. A selected public-data-reuse profile reduces encapsulation and decapsulation cycles by up to 74.6% and 58.8%, respectively. These results demonstrate that substantial deployment gains remain after arithmetic-kernel optimization and motivate a two-stage methodology that also examines the surrounding execution system.
著者のコメント
22 pages and 4 figures. Extended version to the paper in the Proceedings of FPS-2026: 19th International Symposium on Foundations & Practice of Security
arXiv ID: 2610.01960 / 要約の誤りについて