分類階層を使い双曲空間の画像言語モデルを適応させる
Hierarchical Prompt Learning for Hyperbolic Vision-Language Models
この論文をやさしく読む
ひとことで言うと
「動物の中の犬」のような分類の親子関係を、画像と言葉を結び付けるモデルのプロンプト学習に取り入れます。
何に役立つ?
既存の分類体系に新しいクラスを追加する場面で、分類の整合性や転移を改善する方法として役立ちます。
この研究の面白いところ
クラスだけでなく親クラスのプロンプトも学び、その情報を予測に戻します。双曲空間の階層を表しやすい性質を活用しています。
どこまで分かった?
親クラス階層は事前に固定して与えます。ドメイン変化で常に改善するわけではなく、兄弟クラスの細かな識別や画像分布だけの変化では利点が小さいと報告しています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
双曲空間の視覚言語モデル(VLM)は、階層に自然に適した幾何構造で画像とテキストの特徴を表すが、下流タスクへの適応は主として固定プロンプトに依存してきた。一方、既存のプロンプト学習法はクラスラベルを平坦な集合として扱い、利用可能な分類階層を活用していない。本研究では、固定した双曲空間VLMに追加する階層的プロンプト学習プラグインによって、この隔たりに対処する。 事前にオフラインで固定した親クラス階層を与え、クラスのプロンプト学習器に、別の親クラス用プロンプト学習器、親レベルの教師情報、双曲空間の含意関係に基づく正則化、親からのフィードバックによるロジット融合を追加する。CoOp、CoCoOp、MaPLeを用いて具体化し、それぞれHyPLO、CoHyPLO、MaHyPLOを得る。標準的な11データセットのベンチマークで、すべての変種が既存クラスから新規クラスへの汎化とデータセット間転移を改善し、ドメイン変化の下では元のプロンプト学習ベースラインと同程度を保った。6つの階層的指標と埋め込み解析から、分類体系により整合する予測を生み、双曲空間で親、クラス、画像の埋め込みを階層と整合するように組織化することを示す。改善は、新規クラスを固定された分類体系の中に位置付ける必要があるときに最大で、兄弟クラス間の細かな混同や、画像分布だけに影響する変化では最小だった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Hyperbolic vision-language models (VLMs) represent image and text features in a geometry naturally suited to hierarchy, but their adaptation to downstream tasks has largely relied on fixed prompts. Existing prompt learning methods, meanwhile, treat class labels as a flat set and do not exploit available taxonomic structure. We address this gap with a hierarchical prompt learning plug-in for frozen hyperbolic VLMs. Given a fixed offline parent-class hierarchy, it augments a class prompt learner with a separate parent prompt learner, parent-level supervision, hyperbolic entailment regularization, and parent-feedback logit fusion. We instantiate the method with CoOp, CoCoOp and MaPLe, yielding HyPLO, CoHyPLO and MaHyPLO. Across the standard 11-dataset benchmark, all variants improve base-to-new generalization and cross-dataset transfer, and remain comparable to their prompt learning baselines under domain shift. Six hierarchical metrics and embedding analyses show that the method produces more taxonomically consistent predictions and induces a hierarchy-consistent organization of parent, class, and image embeddings in hyperbolic space. Its gains are largest when novel classes must be placed within a fixed taxonomy, and smallest for fine-grained confusions among sibling classes or shifts affecting only the image distribution.
arXiv ID: 2609.24276 / 要約の誤りについて