arXiv論文メモ
新着一覧
cs.AI / cs.CV · 査読状況未確認

蒸留と強化学習を交互に使うマルチモーダル言語モデル

Pistis Technical Report

Heyun Chen, Xiaohan Lan, Jiaxi Li, Zhilin Lu, Qi She, Weiwen Xu, Fei Yu, Yujie Zhong, Jinghuan Chen, Zijian Feng, Siyu Jiao, Yiheng Lin, Xinhao Wang, Sihan Yang, Jieyu You, Changbin Zhang, Hengyu Zhang, Xudong Zhang, Yunqing Zhao, Shuai Zheng

この論文をやさしく読む

ひとことで言うと

画像なども扱う言語モデルを、蒸留と強化学習を交互に使って事後学習し、実行環境も改善する技術報告。

何に役立つ?

マルチモーダル推論やツール利用を行うモデルの学習方法、実行環境の改善方法を検討する材料になる。

この研究の面白いところ

二つの学習目標を固定した比率で混ぜずに交互に使い、モデルを変えず実行環境だけを調整する方法も試した。

どこまで分かった?

要旨は基盤モデルより良いと述べるが、個別ベンチマークの数値は示していない。報告された効果はこのモデル群と実験設定に関するもの。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

本稿は、Qwen3.6を基にした270億パラメータとQwen3.5を基にした90億パラメータのマルチモーダル大規模言語モデル群Pistisを紹介する。モデルは汎用的で拡張可能な事後学習の枠組みで開発した。まず大規模なマルチモーダル教師あり微調整(SFT)によって基盤を作る。その上で、方策に沿った蒸留と強化学習を一つの学習サイクルに密接に組み込むInterleaved Distillation and Reinforcement Learning(IDRL)を提案する。どちらか一方だけを最適化したり、固定した損失で同時に組み合わせたりする代わりに、二つの目的を交互に使う。これによって、知識の伝達、最適化の安定性、長いエージェント実行経路への正確な貢献度の割当てが改善し、能力間でよく見られる相反を抑えながら性能を高めるとする。 両方のモデル規模で、深いマルチモーダル推論を強めるPistis-Thinkingと、長期計画、反復的な推論、ツール利用を支えるためエージェントの実行経路データも加えたPistis-Agenticを作った。後者は特にマルチモーダル検索に強い。両規模とも対応する基盤モデルを上回った。モデルのパラメータ最適化に加え、エージェントの推論用実行環境を反復的な最適化で自動改善するPistis-Auto-Harnessing(PAH)も提案する。実験では、PAHがモデルのパラメータを更新せず、やり取りの予算も増やさずに性能を高めた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.

arXiv ID: 2609.28554 / 要約の誤りについて