arXiv論文メモ
新着一覧
cs.AR · 査読状況未確認

半導体の検証成果物を評価する強化学習の枠組みを提案

Verification Reward Model for Reinforcement Learning in Chip Design Verification

Shashank Chaurasia

この論文をやさしく読む

ひとことで言うと

半導体のテストや検証を作るAIに、どんな証拠からどのように報酬を与えるべきかを設計した提案です。

何に役立つ?

コンパイルできるだけの成果物と、実際に故障を見つける成果物を区別する評価実験を設計する材料になります。

この研究の面白いところ

学習モデルが品質を予測しても、合否を決める検査は外部の決定論的な処理に残します。判定を監査できる契約と証拠を中心に据えています。

どこまで分かった?

著者が明記する通り、実験計画を示す構想論文です。学習済みモデルの性能、EDAツールでの成果、実チップの結果はいずれも報告されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

半導体チップの検証成果物を作成・修復する言語モデルを学習するために、検証報酬モデルVRMの枠組みを提案する。中心となるのは、要求、許容する入力刺激、観測境界、参照動作、評価予算、受け入れ基準を結び付ける、版管理された検証契約である。コンパイラ、シミュレータ、形式手法、ミューテーション、カバレッジ、および専門家レビューの証拠を、監査可能な記録に変換する。決定論的な受け入れ検査は、学習モデルの外部に残す。 学習された結果モデルは、仕様、生成された成果物、および明示的にマスクした部分的証拠から、取得コストの高い将来の証拠を予測する。意味的な批評器は証拠で裏付けられる弱点を特定し、決定論的な報酬合成器は、得られた品質ベクトルをタスクに応じた学習報酬へ変換する。拡張には、正常状態と故障状態を対にした介入、追加的な故障発見に対する報酬、不確実性を考慮した評価スケジューリング、および隔離された領域横断学習ループが含まれる。 評価では、独立に検証された故障検出、誤警報、設計族をまたぐ汎化、指定品質に到達するまでの総費用を重視する。軽量モデルと、公開ツールで扱える小規模なベンチマークを使う最初の実験を記述し、UVMの全面的な機能は機能ごとの適格性確認を通じて認める。本稿は立場を提示する論文である。枠組みと、それを検証するための実験を規定するが、学習、EDA、実チップに関する結果は一切報告しない。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We propose a Verification Reward Model (VRM) framework for training language models to create and repair chip verification artifacts. The central object is a versioned verification contract that binds requirements, permissible stimulus, observation boundaries, reference behavior, evaluation budgets, and acceptance criteria. Compiler, simulator, formal, mutation, coverage, and expert-review evidence are converted into auditable records. Deterministic acceptance checks remain outside the learned model. A learned outcome model predicts expensive future evidence from the specification, the generated artifact, and explicitly masked partial evidence; a semantic critic identifies evidence-supported weaknesses; a deterministic reward composer translates the resulting quality vector into task-conditioned training rewards. Extensions include paired clean/fault interventions, marginal fault-discovery rewards, uncertainty-aware evaluation scheduling, and a quarantined cross-domain learning loop. Evaluation emphasizes independently validated fault detection, false alarms, generalization across design families, and total cost to reach a specified quality level. We describe a first experiment on a small, open-tool-compatible benchmark with lightweight models, with full UVM capability admitted through feature-specific qualification. This is a position paper: we specify the framework and the experiments that would test it, and we report no training, EDA, or silicon results.

著者のコメント

Position paper; 1 figure

arXiv ID: 2609.22347 / 要約の誤りについて