追加学習なしで分野をまたぐ少数画像分類を改善
Training-Free Spectral Transductive Refinement for Cross-Domain Few-Shot Classification
この論文をやさしく読む
ひとことで言うと
各クラスの見本が少ない画像分類で、判定対象の画像群もまとめて使い、学習済み特徴の配置からクラス代表を改善します。
何に役立つ?
エンコーダを再学習できない条件で、異なる分野の少数画像分類を改善する用途が考えられます。
この研究の面白いところ
訓練を追加せず、テスト時に画像群の関係をスペクトル座標で捉えて反復更新する点です。
どこまで分かった?
判定対象群をまとめて利用するトランスダクティブ設定です。画像を1枚ずつ独立に分類する条件と同一ではなく、要旨には具体的な正解率はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
固定した視覚特徴を用いる少数例認識は、分野の違いがあり、各クラスのラベル付き画像が1枚だけの場合に特に不安定になる。1枚の画像ではクラスを信頼性高く推定できないためである。本研究では、エンコーダを再学習せず、元の分野のデータ拡張も行わずに、テスト時の処理だけでこの不安定さをどこまで減らせるかを問う。ラベル付き支援例と判定対象からなるエピソード全体の幾何を利用する、学習不要のトランスダクティブな推論規則Spectral Transductive Refinement(STR)を提示する。 固定した埋め込みを入力として、STRは両者をまとめたk近傍グラフを構築し、エピソードを正規化Laplacianのスペクトル座標系に写す。ラベル付き支援例からクラス代表を初期化し、擬似ラベルを付けた判定対象を使って反復的に改良する。STRを2つの評価手順で検証する。固定したResNet-18の特徴を使う、条件を制御した構成要素の検討では、スペクトルによる改良は、分野が異なる5領域で単一プロトタイプのスペクトル初期化より一貫して改善した。改善が最大となったのは、支援例からの推定が最も弱い1-shot設定だった。 続いて、miniImageNetで事前学習した標準的なResNet-10を用い、広く使われる8つの対象領域で、最近の分野間少数例学習(CD-FSL)手法と比較した。STRは推論時だけに動作し、比較手法の中で1-shotの平均成績が最も高く、5-shotでも競争力を保ち、元の分野で大規模なメタ学習用拡張を行う手法に匹敵した。STRはトランスダクティブな手法なので、その設定を明示する。診断分析では、改善はスペクトル座標での反復的な改良に由来し、本設定では働いていないプロトタイプ容量の追加に由来するものではなかった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Few-shot recognition with frozen visual features is especially fragile under domain shift and one-shot supervision, where a single labelled image is an unreliable estimate of its class. We ask how far this fragility can be reduced purely at test time, without retraining the encoder or augmenting the source domain. We present Spectral Transductive Refinement (STR), a training-free transductive inference rule that exploits the geometry of the complete support-query episode. Given frozen embeddings, STR builds a joint k-nearest-neighbour graph, maps the episode into a normalized-Laplacian spectral coordinate system, initializes class representatives from the labelled support, and iteratively refines them using pseudo-labelled queries. We evaluate STR under two protocols. A controlled component study with frozen ResNet-18 features shows that spectral refinement consistently improves over single-prototype spectral initialization across five shifted domains, with the largest gains in the one-shot regime where support estimates are weakest. We then benchmark STR against recent Cross-Domain Few-Shot Learning (CD-FSL) methods using the standard miniImageNet-pretrained ResNet-10 backbone over eight established target domains. Operating entirely at inference time, STR attains the highest 1-shot average among compared methods and remains competitive at 5-shot, rivalling approaches relying on heavy source-domain meta-training augmentations. Because STR is transductive, we report its setting explicitly. Diagnostics attribute its gains to iterative refinement in spectral coordinates rather than added prototype capacity, which remains inactive in our configuration.
著者のコメント
9 pages (excl. references), 3 figures, 6 tables
arXiv ID: 2609.23758 / 要約の誤りについて