arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

周波数情報と複数課題の連携で超音波画像を解析

FreqDINO++: A Frequency-Guided Multi-Task Routing Vision Foundation Model for Universal Ultrasound Analysis

Qing Xu, Yixuan Zhang, Yue Li, Xiangjian He, Qian Zhang, Mainul Haque, Rong Qu, Wenting Duan, Jieyun Bai, Zhen Chen

この論文をやさしく読む

ひとことで言うと

超音波画像の病変領域の分割や良悪性分類などを、一つの基盤モデルで連携させて扱います。

何に役立つ?

複数の超音波解析課題を効率よく学習する方法の検討に役立ちます。27の臨床課題シナリオに対応するベンチマークで評価しています。

この研究の面白いところ

課題共通と課題固有の情報を分けるアダプター、周波数情報の強化、局所領域と画像全体の予測を結ぶデコーダーを組み合わせます。

どこまで分かった?

ベンチマーク上での優位性と未見データへの一般化が報告されていますが、要旨には具体的な差の数値や前向き臨床試験の記載はありません。診療での安全性まで確立したとはいえません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

超音波画像解析は、がん検診や出生前診断で重要な役割を果たすが、包括的な評価には病変の領域分割や良悪性分類などの課題を同時に扱う必要がある。近年の視覚基盤モデルは優れた汎用表現を示しているものの、自然画像との大きなドメイン差が、超音波分野でその能力を引き出す際の障害となっている。既存手法は通常、個別の課題ごとに大規模な視覚エンコーダを微調整するため、多大な計算負荷が生じ、異種の課題に共通する性質も見落としている。 本研究では、汎用的な超音波解析のために、周波数情報に基づくマルチタスクルーティング視覚基盤モデルFreqDINO++を提案する。まず、課題共通の知識と課題固有の知識をパラメータ効率よく統合するMulti-task Routing Adapter(MR-Adapter)を導入する。続いて、超音波画像の豊かなマルチスケール周波数特性を捉えるFrequency-aware Feature Enhancer(F²-Enhancer)を設計する。さらに、グローバルなトークンとローカルなトークンの相互作用を通じて、密な予測と画像全体の予測を行う課題の協調を促すTask-aligned Collaborative Decoder(TC-Decoder)を考案する。 大規模なマルチタスク超音波ベンチマークと外部の単一課題ベンチマークで広範に実験した結果、FreqDINO++は27種類の多様な臨床課題シナリオにわたり、有力なベースラインや近年の基盤モデルを一貫して上回り、未知のデータへの汎化にも有望な結果を示した。コード:https://github.com/MingLang-FD/FreqDINO-Plus。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Ultrasound image analysis plays a crucial role in cancer screening and prenatal diagnosis, yet comprehensive assessment requires jointly addressing tasks such as lesion segmentation and benign-malignant classification. While recent vision foundation models have shown remarkable universal representations, unlocking their potential for ultrasound is bottlenecked by the considerable domain gap from natural images. Existing methods typically fine-tune heavy vision encoders for isolated tasks, incurring substantial computational overhead while overlooking the underlying commonalities across heterogeneous tasks. In this work, we propose FreqDINO++, a frequency-guided multi-task routing vision foundation model for universal ultrasound analysis. We first introduce a Multi-task Routing Adapter (MR-Adapter) to support parameter-efficient integration of task-common and task-specific knowledge, a Frequency-aware Feature Enhancer (F$^2$-Enhancer) is then designed to capture the rich multi-scale frequency characteristics of ultrasound images, and a Task-aligned Collaborative Decoder (TC-Decoder) is devised to promote collaboration between dense and global prediction tasks through global-local token interaction. Extensive experiments on large-scale multi-task and external single-task ultrasound benchmarks demonstrate that FreqDINO++ consistently outperforms strong baselines and recent foundation models across 27 diverse clinical task scenarios, while also showing promising generalization to unseen data. The code is at https://github.com/MingLang-FD/FreqDINO-Plus.

著者のコメント

Accepted by TBME

arXiv ID: 2609.20340 / 要約の誤りについて