arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

推論文を生成せず候補から選んで画像検索を高速化

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang

この論文をやさしく読む

ひとことで言うと

小型の画像・文章検索エージェントで、毎回長い推論を書く代わりに、用意した推論候補から適切なものを選びます。

何に役立つ?

小規模モデルを使うマルチモーダル検索で、成功率を維持しながら応答待ち時間を減らす方法として評価されています。

この研究の面白いところ

候補の尤度を並列に計算し、共通の文脈の計算結果も再利用します。追加の分類用ヘッドを設けず、元のモデルの尤度を使います。

どこまで分かった?

評価対象は20億・40億規模のモデルと7つの検索ベンチマークです。90%超は推論部分だけの短縮で、質問全体のモデル計算の短縮は28〜54%です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

マルチモーダルエージェントは一般に、各行動の前に自由形式の推論を生成する。小規模モデルでは、モデルの能力が限られているため、行動生成にほとんど役立たない長い推論を作り、大きな推論計算コストを生じる場合がある。この課題に対し、推論を自由な生成ではなく選択として定式化し直す、選択に基づく構造化推論(SSR)を導入する。SSRは繰り返し現れる高水準の推論を、事前に指定した再利用可能な自然言語の候補として表現する。各ターンでは、現在の文脈を条件とした尤度に基づき、追加のタスク用ヘッドを必要とせず、モデルがこれらの推論候補から選択する。 事前に指定した推論過程を使うことで、並列のスコア計算が可能となる。教師強制を用いるプリフィル処理が、共有された文脈のKVキャッシュを使い、候補内および候補間のトークン尤度を同時に計算する。私たちは20億・40億パラメータのモデルを用い、7つのマルチモーダル検索ベンチマークでSSRを評価する。複数の強化学習目的関数と教師あり微調整にわたり、SSRはタスク性能を犠牲にせず、大きな効率改善を実現する。同規模の主要な検索エージェントに匹敵する平均成功率を達成しつつ、1ターン当たりの推論部分の待ち時間を90%超、1問当たりのモデル推論の総待ち時間を28〜54%削減する。プロジェクトページ:https://zfy0314.github.io/ssr-webpage/。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.

arXiv ID: 2610.01892 / 要約の誤りについて