画像内の物体位置を途中の候補領域から修正する
CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
この論文をやさしく読む
ひとことで言うと
文章で指定された物体を画像中から探す際に、途中の候補領域を座標として残し、その座標を後から修正する方法です。
何に役立つ?
物体の取り違えや境界のずれを、最終出力だけでなく途中の状態から点検・修正する仕組みになります。自然画像とリモートセンシング画像で評価しています。
この研究の面白いところ
推論の文章は固定し、空間座標を編集対象として扱います。大きく壊した位置推定からの回復能力を、通常の精度とは別に評価している点も特徴です。
どこまで分かった?
27ポイント超の改善は制御した破損条件でのボックス重なりの結果です。通常入力全体で同じ改善が得られるという意味ではありません。大規模モデルとの比較もグラウンディング精度についてです。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚的グラウンディングは、言語で指定された物体の位置を境界ボックスで特定する。多くのマルチモーダルなグラウンディングモデルは、対象の識別、空間推論、境界推定を1回の最終予測にまとめている。自由形式の根拠説明は推論を言語として明示するが、測定・編集可能な空間状態を必ずしも示さない。そのため、途中の位置推定の誤りを診断・修正することが難しく、誤った領域選択や不正確な境界が最終ボックスまで残る。 本研究では、グラウンディングを明示的な状態構築と状態編集に分ける、構築後に編集する枠組みCoEvolveを導入する。Region-Evolution Reinforcement(RER)は、グラウンディングの解析を意味と空間の両面で段階的に進む軌跡として組織し、各推論段階で明示的な候補領域を確定させる。Bidirectional Denoising Refiner(BDR)は、推論文を固定した意味的文脈として扱い、双方向の同一位置再構成によって軌跡の座標フィールドを精緻化する。幾何と振る舞いの各水準の目的関数が、対象形状と編集の選好に関する信号を与え、信頼できる候補を強化し、正確な入力を保持し、またはアノテーションへ向けて修正する。 自然画像とリモートセンシングのグラウンディングを評価する。90億パラメータの基盤モデルを用いるCoEvolveは、最大2410億パラメータのモデルに匹敵する位置特定精度を示す。制御した破損を加える条件では、BDRを1回適用するだけでボックスの平均重なりが27パーセントポイント超改善し、大きな位置推定誤差からの強い回復能力を示す。状態の生成元を比較した結果も、明示的な状態構築と生成元に合わせた編集の相補性を支持する。プロジェクトは https://sundongwei.github.io/CoEvolve_Project/ で公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region choices and imprecise boundaries to persist in the final box. We introduce CoEvolve, a construct-to-edit framework that separates grounding into explicit state construction and state editing. Region-Evolution Reinforcement (RER) organizes grounding analysis into a progressive semantic--spatial trajectory, with each reasoning step committing to an explicit candidate region. Bidirectional Denoising Refiner (BDR) treats the reasoning text as fixed semantic context and refines the trajectory's coordinate fields through bidirectional same-position reconstruction. Geometry- and behavior-level objectives provide target geometry and edit-preference signals for consolidating reliable candidates, preserving accurate inputs, or correcting toward annotations. Evaluations cover natural-image and remote-sensing grounding. With a 9B backbone, CoEvolve rivals models up to 241B parameters in grounding accuracy. Under controlled corruption, a single BDR pass improves mean box overlap by over 27 percentage points, demonstrating strong recovery from substantial localization errors. State-source comparisons further support the complementarity of explicit state construction and source-matched editing. The project is at https://sundongwei.github.io/CoEvolve_Project/.
arXiv ID: 2610.01710 / 要約の誤りについて