FASA:VLAモデルのフィードバック適応サンプリング
FASA: Feedback-Aware Sampling Adaptation for Efficient Diffusion-Based VLA Models
この論文をやさしく読む
ひとことで言うと
拡散型のロボット行動モデルで、毎回同じ回数の生成計算をする代わりに、画像や把持力などのフィードバックで必要なステップ数を変えます。追加学習なしの実行時適応です。
何に役立つ?
計算資源の限られた機器で、状況に応じて生成処理の負担を調整するために役立ちます。実時間の視覚・力覚・自己状態を、行動を作る計算量の制御にも使います。
この研究の面白いところ
相互作用の情報で大まかなステップ予算を変え、自己受容情報でその範囲内のステップを絞ります。固定的な枝刈りやキャッシュでは捉えにくい動作段階ごとの負荷差に対応します。
どこまで分かった?
複数ベンチマークで最大1.45倍の推論高速化を示し、成功率は競争力を維持したと報告しています。全条件で成功率が同一という意味ではなく、要旨には課題別の値や実機構成の詳細はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
拡散型視覚・言語・行動(VLA)モデルは身体タスクで高性能だが、反復サンプリングの計算量とメモリアクセスがエッジでのリアルタイム展開を妨げる。既存の高速化は高コストな学習を要するか、ロボット相互作用の仕事量の変化を無視した静的枝刈り・キャッシュで知覚性能を下げる。 FASAは、リアルタイムのマルチモーダルフィードバックを脱ノイズの制御信号にする学習不要の実行時枠組みである。相互作用駆動の範囲アダプタが視覚とグリッパー力から全体のサンプリングステップ予算を調整し、固有感覚対応ステップアダプタが調整範囲内の最適ステップを特定する。比較評価では、競争力のある成功率を保ちながら推論速度を最大1.45倍にできた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Diffusion-based Vision-Language-Action (VLA) models achieve strong performance in embodied tasks, but their iterative sampling imposes heavy computational and memory-access cost, blocking real-time deployment on edge platforms. Existing acceleration methods either require expensive training (e.g., distillation, flow matching) or degrade perception via statically scheduled pruning and caching, ignoring the dynamic workload variance of robotic interactions. This paper presents FASA (Feedback-Aware Sampling Adaptation), a training-free runtime framework that treats real-time multimodal feedback as a control signal for the denoising pipeline: an interaction-driven range adaptor modulates the global sampling-step budget based on visual and gripper-force feedback, and a proprioception-aware step adaptor pinpoints the optimized step within the adapted range. This co-designed framework allows the underlying hardware architecture to adaptively match the workload demands of different execution phases. Comparative evaluations across several benchmarks show that the inference speed can be increased by up to 1.45$\times$ while maintaining competitive success rates, providing a novel dynamic runtime architecture paradigm for deploying heavy generative embodied AI workloads onto resource-constrained computing platforms.
著者のコメント
Accepted by the 18th International Conference on Networking, Architecture, and Storage (NAS 2026)
arXiv ID: 2609.19475 / 要約の誤りについて