arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

AIの運用条件に確率を付けて代表的なテストを設計

Probabilistic Modelling of Operational Design Domains, A New Approach for Testing AI Systems

Hans-Werner Wiesbrock

この論文をやさしく読む

ひとことで言うと

AIが使われる条件を分類するだけでなく、それぞれがどのくらい起きるかもモデルに入れ、実際の運用を代表するテストを作る方法です。

何に役立つ?

テストケースの抽出、テストを終える基準、既存結果の再評価、学習データの偏りの検討に使う枠組みです。自動列車運転をモデル化の例にしています。

この研究の面白いところ

大きな条件付き確率表を直接管理せず、周辺分布と関数による依存関係からBayesネットワークを補います。運用条件の記述を統計的なテスト設計につなげています。

どこまで分かった?

要旨はモデル化と評価手順を示しますが、実運行での安全性改善や事故率低減を報告していません。終了基準は与えた品質目標、有意水準、確率的なODDの枠組みに基づくものです。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

車両の障害物検知のような機械学習に基づくシステムに従来のテスト手順を適用すると、すぐに限界が生じる。テストで障害物を検出できなかった場合、古典的なバグ修正はできず、AIシステムには常に欠点が残るためである。したがってテスト結果は統計的に解釈するほかなく、そのためには、システムの運行設計領域(ODD)について網羅的なだけでなく、それを代表するテスト集合が必要になる。 この目的で、確率的に拡張されたオントロジー(PEON)を導入する。これはODDを記述するオントロジーに、そのオントロジーが誘導する区分上の確率分布を付加したものである。保守が困難な条件付き確率表の代わりに、周辺分布と関数で記述された依存関係だけを指定すればよい。結合と最適輸送に基づくアルゴリズムが、この指定をBayesネットワークへ補完する。 PEONから、代表的なテストケースのサンプリング、与えられた品質目標と有意水準に対する厳密なテスト終了基準、既存のテスト結果を再評価する方法、学習データのバランスを評価する方法を導く。自動列車運転を例として、複雑なODDの実用的なモデル化を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The conventional testing process quickly fails when applied to ML-based systems such as obstacle detection in vehicles: if an obstacle is not detected in a test, classical bug fixing is impossible and an AI system will always retain shortcomings. Test results can therefore only be interpreted statistically, which in turn requires test sets that are not only complete with respect to the operational design domain (ODD) of the system, but also representative of it. To this end, we introduce probabilistically extended ontologies (PEONs): ontologies describing the ODD, augmented with a probability distribution over the partitioning they induce. Instead of unmaintainable conditional probability tables, only marginal distributions and functionally described dependencies need to be specified; algorithms based on couplings and optimal transport complete this specification to a Bayesian network. From a PEON we derive the sampling of representative test cases, rigorous end-of-test criteria for given quality targets and significance levels, and methods for re-evaluating existing test results and for assessing the balance of training data. We demonstrate the practical modelling of a complex ODD using the example of automatic train operation.

著者のコメント

50 pages, 43 figures, Technical Report

arXiv ID: 2609.24397 / 要約の誤りについて