arXiv論文メモ
新着一覧
cs.CV / cs.LG · 査読状況未確認

少数例で病理画像と言語の複数モデルを融合

FFM-CP: Cross-Backbone Fusion of Vision-Language Foundation Models for Few-Shot Computational Pathology

Anh-Tien Nguyen, Trung DQ. Dang, Nghiem Tuong Diep, Bui Ngoc Han Nguyen, Tan-Ha Mai, Miriam Cindy Maurer, Phuong Hoa Nguyen, Thi Thuy Uyen Nguyen, Youngjun Park, Daniel Sonntag, Duy Minh Ho Nguyen, Anne-Christin Hauschild

この論文をやさしく読む

ひとことで言うと

少数の病理画像の正解例から、複数の視覚・言語モデルの特徴を合わせて分類する。

何に役立つ?

考えられる用途は、専門家の注釈が少ない病理画像分類の改善である。

この研究の面白いところ

追加の位置合わせネットワークなしに表現を揃え、文章原型と似た事例の二経路を使う。

どこまで分かった?

六つのデータセットで54比較中50比較に改善を示した。臨床現場での診断性能は要旨に記載がない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

病理分野の視覚・言語基盤モデルの性能は疾患や課題によって異なり、常に最良の単一モデルはない。専門家による病理の注釈には費用がかかり、課題ごとに適応するためのラベル付きデータも限られる。事前学習した表現を組み合わせる方法が考えられるが、少数のラベル付き例から効果的な融合を学ぶのは難しい。提案するFFM-CPは、少数例学習で複数の病理視覚・言語モデルを組み合わせる枠組みである。まず対応する支援画像から求めた閉形式の直交Procrustes変換で、異質な表現を位置合わせする。追加の位置合わせネットワークを学習せず、各モデル内の特徴の幾何を保つ。位置合わせ後の空間では、統一したグラフを通じてモデル間の情報交換を行い、支援画像の特徴と視覚・文章によるクラスの原型を一緒に改良する。この表現を用い、クラスの意味知識を捉える文章原型の経路と、クラス内の視覚的な変動を捉える事例検索の経路を組み合わせる。各経路は順序付きのすべてのモデル対の予測を統合するため、あるモデルで符号化した問い合わせに別モデルの証拠を使える。三種類のモデル組み合わせを六つの病理組織画像データセットで評価し、クラスごと4、8、16例を用いた。54比較中50比較で、FFM-CPの平均macro-F1は、融合対象のうち最も強い単独適応モデルを上回った。少ない注釈しかない場合に、相補的な事前学習表現の融合が病理組織分類を改善し得ることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Pathology vision-language foundation models vary in performance across diseases and tasks, with no single model consistently performing best. The high cost of expert pathology annotation can also limit the labeled data available for task-specific adaptation. Combining complementary pretrained representations is a potential approach to these limitations, yet learning an effective fusion from few labeled examples remains challenging. We introduce Few-shot Fusion Foundation Models of Computational Pathology (FFM-CP), which is a framework that combines multiple pathology vision-language models in the few-shot learning setting. The framework first aligns heterogeneous representations using a closed-form Orthogonal Procrustes transformation estimated from corresponding support images. This alignment preserves within-model feature geometry without training an additional alignment network. Within the aligned space, a unified graph enables information exchange across backbones by jointly refining support-image features and visual and textual class prototypes. These refined representations support complementary text-prototype and case-retrieval branches that capture semantic class knowledge and within-class visual variation, respectively. Each branch learns to combine predictions from all ordered backbone pairs, allowing queries encoded by one model to draw on evidence represented by another. We evaluate three backbone combinations on six histopathology datasets at 4, 8, and 16 shots per class. FFM-CP achieves higher mean macro-F1 than the strongest individually adapted member of each fused set in 50 of 54 comparisons. These findings suggest that combining complementary pretrained representations can improve histopathological classification when annotations are limited.

arXiv ID: 2609.27710 / 要約の誤りについて