arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

安全のために存在すべき物が欠けた場面を言葉で説明する

Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint

Zhiyun Jiang, Hanyong Wang, Binbin Liang, Yu Xie, Menglong Yang, Wei Li

この論文をやさしく読む

ひとことで言うと

画像に写っている物だけでなく、安全上あるはずなのに欠けている物や状態を説明する方法です。

何に役立つ?

安全確認の画像理解で、保護具などの不在を意味のある欠落として扱うことが考えられます。

この研究の面白いところ

安全な状態を仮想的に再構成し、現実との差から否定的な説明を生成します。壊れた物の補完と、完全に不在の安全対象の推定を別の枝で処理します。

どこまで分かった?

実験で有効性を示すとしていますが、要旨にはデータ規模や具体的な精度値はありません。仮想的な安全状態を使う推論であり、現場の安全を保証するシステムと同一視できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

従来のシーン理解は、画像内に客観的に存在する肯定的な情報に注目する。しかし、安全が重要な分野では、本来あるべきなのに実際には欠けている重要情報を理解することが、リスクの軽減に不可欠である。この隔たりを埋めるため、安全性を認知上の制約とした、視覚シーンの否定的内容を説明するキャプション生成に取り組む。中心的な課題は、物理的な不在を意味的な否定事象へ変換することである。既存の視覚言語モデル(VLM)は、肯定バイアスが否定的推論を抑え、頭の中で欠落を補う能力の制約と表現バイアスも不在情報の推論を妨げるため、この処理を苦手とする。 これらの課題に対し、反実仮想再構成と対比的デコーディングに基づく否定キャプション生成フレームワークCRCDを提案する。人の認知に着想を得て、課題を反実仮想の潜在的変化の説明生成として定式化し直し、肯定バイアスを回避する。合成した安全な期待状態と現実を対比させ、意味的な欠落を特定する。 欠落を補う能力の制約に対しては、2分岐の反実仮想再構成アーキテクチャを設計する。アモーダル補完分岐は不完全な物体を復元し、機能的関連付け分岐は完全に欠けている安全関連物体を推論する。同時に、多条件表現学習機構を組み込み、汎用特徴を事前定義した安全基準の部分空間に射影して、より多くの次元にまたがる情報を捉えることで表現バイアスを緩和する。 再構成したシーンのプロトタイプと元の入力の間にある特徴レベルの意味的残差をデコードすることで、CRCDは不在の探索空間を限定し、デコーダーの否定論理を活性化する。広範な実験によってCRCDの有効性を検証し、この先駆的な課題に対する高性能なベースラインを確立した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Traditional scene understanding focuses on affirmative information objectively present in images. However, in safety-critical domains, comprehending key information that should exist but is actually absent is vital for risk mitigation. To bridge this gap, we focus on visual scene negative captioning with safety as the cognitive constraint. The core challenge is to convert physical absence into semantic negative events. Existing vision-language models (VLMs) struggle with this process because affirmation bias suppresses negative reasoning, while limited mental filling capability and representation bias further hinder the inference of absent information. To address these challenges, we propose a negative captioning framework based on counterfactual reconstruction and contrastive decoding (CRCD). Inspired by human cognition, CRCD reformulates the task as counterfactual latent change captioning to bypass affirmation bias. It contrasts a synthesized safe expectation with reality to identify semantic omissions. To address limited mental filling, we design a dual-branch counterfactual reconstruction architecture. The amodal completion branch restores defective objects, while the functional association branch infers completely absent safety objects. Concurrently, a multi-condition representation learning mechanism is integrated to mitigate representation bias by projecting universal features onto predefined safety criteria subspaces, thereby capturing information across more dimensions. By decoding feature-level semantic residuals between the reconstructed scene prototype and raw input, CRCD bounds the non-existence search space and activates the decoder's negative logic. Extensive experiments validate the effectiveness of CRCD, establishing a high-performance baseline for this pioneering task.

arXiv ID: 2609.19812 / 要約の誤りについて