arXiv論文メモ
新着一覧
cs.HC · 査読状況未確認

複数の視覚条件を満たす画像を人の確認と少量ラベルで検索

ProbeScout: Visual Analytics for Attribute-Guided Image Search

Yifan Lv, Yiyun Chen, Daojun Ye, Haotian Yang, and Weikai Yang

この論文をやさしく読む

ひとことで言うと

『夕暮れ』『交差点』『信号機あり』を別々に確認し、すべての条件を満たす画像を、人のフィードバックで絞り込む検索システムです。

何に役立つ?

学習データの整理や、モデルが苦手な画像の収集など、細かい条件を変えながら画像を探す作業に役立ちます。属性の確認結果を次の検索に再利用できます。

この研究の面白いところ

一つの類似度だけに頼らず、属性ごとの判断を組み合わせます。少数のラベルと人の修正で、画像すべてに VQA をかける方法と比較しています。

どこまで分かった?

最大2%という数値は、別に実施した10タスク比較でのラベル付け対象の割合です。人のフィードバックを含む方式であり、完全自動で同じ結果が得られるとしたものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

分析者は、モデルの診断、データセットの整備、対象を絞った学習のために、『夕暮れの信号機のある交差点』など、複数の視覚条件を同時に満たす画像を特定する必要がある。埋め込みに基づく検索は大きな画像集合を効率よく順位付けできるが、視覚的に目立つ条件が弱い条件を覆い隠すことがあり、単一の類似度スコアでは必要な条件の論理積を強制できない。視覚質問応答(VQA)は各条件を明示的に検証できるが、画像集合全体への網羅的な適用は、特に検索条件を修正する場合に高価である。 このため、分析者が各条件の証拠を素早く作り、依頼が変わっても再利用できるよう、属性の水準で人を処理に参加させることを考える。本研究では、そのループを支援する視覚分析システム ProbeScout を提示する。まず少量の VQA ラベルから組み合わせ可能な属性プローブを作り、それらを融合して、条件の論理積を考慮した初期順位を得る。連動した表示により、素早い選別、条件を惜しくも満たさない例の診断、プローブ出力の絞り込みと組み合わせによる即時の部分集合構成を支援する。分析者が属性と検索要求の水準で軽いフィードバックを与えると、プローブ自体を固定したまま、融合の重みを段階的に改善する。検証された属性は将来の検索にも再利用できる。 3データセットの17検索タスクで評価し、埋め込みの基準手法に対する検索改善を示す。別の10タスクの比較では、画像集合の最大2%にラベルを付けるだけで、網羅的な VQA より高いタスク単位のマクロ平均 AP と F1 を達成する。さらに2つの事例研究により、現実的な作業の中で対話的な分析と改善を支援することを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Analysts often need to identify images that jointly satisfy multiple visual conditions, such as a crossroads with traffic lights at dusk, for model diagnosis, dataset curation, and targeted training. Embedding-based retrieval can rank the large gallery efficiently, but a visually dominant condition can obscure weaker conditions, and a single similarity score does not enforce the required conjunction. Visual question answering (VQA) can explicitly verify conditions, yet exhaustively applying it to the full gallery is costly, especially when analysts refine their query. These limitations motivate keeping humans in the loop at the attribute level, where analysts can quickly build evidence for each condition and reuse it when the request changes. We therefore present ProbeScout, a visual analytics system that supports this loop. It first builds composable attribute probes from sparse VQA labels and fuses them into a conjunction-aware initial ranking. Coordinated views support rapid screening, near-miss diagnosis, and on-the-fly subset construction by filtering and combining these probe outputs. Analysts provide lightweight attribute- and query-level feedback, which drives staged refinement of fusion weights while keeping the probes fixed. These verified attributes can be reused for future queries. We evaluate ProbeScout on 17 retrieval tasks across three datasets, showing improved retrieval over embedding baselines. A separate 10-task comparison achieves higher task-macro AP and F1 than exhaustive VQA while labeling at most 2% of the gallery images. Two case studies further demonstrate how ProbeScout supports interactive analysis and refinement in realistic workflows.

arXiv ID: 2609.24110 / 要約の誤りについて