arXiv論文メモ
新着一覧
astro-ph.HE / physics.comp-ph / physics.plasm-ph · 査読状況未確認

プラズマ粒子シミュレーションを複数GPUで高速化

Performance-portable GPU acceleration of the hybrid particle-in-cell code dHybridR

Bricker Ostler, Miha Cernetic, Damiano Caprioli

この論文をやさしく読む

ひとことで言うと

無衝突プラズマの計算コードdHybridRを複数メーカーのGPUで高速に動くよう実装し、大規模計算機で性能を測った。

何に役立つ?

考えられる用途は、従来は計算量が障害だった大規模3次元プラズマシミュレーションである。

この研究の面白いところ

OpenMPとSYCLを組み合わせ、共通コード基盤を保ったまま複数のGPUメーカーを対象にしている。

どこまで分かった?

AuroraとFrontierで最大49,152アクセラレータを用い、弱スケーリング効率86~97%を報告した。スループット最大32倍と消費エネルギー約90%減はセル当たり256粒子など記載の条件による。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ハイブリッド粒子・セル法のシミュレーションは、衝突の少ない天体物理・宇宙プラズマの運動論的過程を調べるため広く使われる。しかし、大規模な3次元計算には高い計算コストがかかり、実運用規模の研究の多くは2次元や限られた計算領域にとどまってきた。この課題に対し、ハイブリッド粒子・セル法コードdHybridRを、異なる環境でも性能を出せるようGPUへ実装した。 実装ではOpenMPのターゲット・オフロードと、特に計算負荷が大きい処理向けの専用SYCLカーネルを組み合わせ、Intel、AMD、NVIDIAのGPUを支えるCPU・GPU共通のコード基盤を保った。エクサスケールスーパーコンピュータAuroraとFrontierでは、計49,152個のアクセラレータにわたって86~97%の弱スケーリング効率を達成した。セル当たり256粒子の条件では、ノード全体のGPUスループットがベクトル化されたCPU実装の最大32倍となり、粒子更新当たりのエネルギー消費は約90%少なかった。 著者らの把握する限り、他のハイブリッド粒子・セル法コードでこの規模のGPU性能は報告されておらず、dHybridRはエクサスケールシステムを活用できる位置にある。この進展により、無衝突プラズマの大規模3次元ハイブリッド運動論シミュレーションに必要な計算資源の壁が大きく下がる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Hybrid particle-in-cell simulations are widely used to study kinetic processes in collisionless astrophysical and space plasmas, yet the high computational cost of large-scale three-dimensional runs has largely confined production studies to two dimensions or restricted domains. To address this challenge, we present a performance-portable GPU implementation of the hybrid particle-in-cell code dHybridR. The implementation combines OpenMP target offloading with specialized SYCL kernels for the most computationally expensive operations, while preserving a unified CPU-GPU codebase that supports Intel, AMD, and NVIDIA GPUs. On the exascale supercomputers Aurora and Frontier, dHybridR achieves weak scaling efficiencies of $86\%$ to $97\%$ across 49,152 accelerators, and at 256 particles per cell, its full-node GPU throughput is up to $32\times$ that of the vectorized CPU implementation at approximately $90\%$ less energy per particle-update. To our knowledge, no other hybrid particle-in-cell code has reported GPU performance at this scale, leaving dHybridR uniquely positioned to exploit exascale systems. These advances substantially lower the computational barrier to large-scale three-dimensional hybrid-kinetic simulations of collisionless plasmas.

arXiv ID: 2609.28422 / 要約の誤りについて