arXiv論文メモ
新着一覧
cs.SD / cs.LG / eess.AS · 査読状況未確認

音の特徴を再利用して軽い反復処理で認識する

LAST: Looped Audio Spectrogram Transformer

Haider Al-Tahan, Sean O'Brien, Anastasia Razdaibiedina, N. Apurva Ratan Murty

この論文をやさしく読む

ひとことで言うと

音の特徴を毎回計算し直す代わりに、分類に使うトークンだけを繰り返し更新して認識精度を高める方法です。

何に役立つ?

音響分類モデルのパラメーター数や演算量を減らしつつ、認識性能を保つ設計に役立ちます。実測スループットの改善も報告されています。

この研究の面白いところ

層を深くする代わりに同じブロックを再利用し、反復の対象をクラストークンに絞っています。2回から10回への比較では、精度向上に対する計算量増加が小さい点が特徴です。

どこまで分かった?

2回と10回の結果は別々に学習したモデルの比較で、同じモデルの推論回数を自由に変えた結果とは区別されます。2.1%は相対的な改善であり、2.1ポイントではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

Transformerモデルを深くすると認識性能は向上するが、大きな費用を伴う。層を追加するたびにパラメーターが増え、計算効率が悪くなる。本研究では、追加処理を、すでに計算した特徴の統合に集中させられるかを問う。Looped Audio Spectrogram Transformer(LAST)は、最初にすべてのトークンを処理し、その後は同じブロックを再利用して、固定した音響特徴に対してクラストークンだけを洗練する。これにより、後続の処理を低コストにする。 AudioSetでは、10回通過するLASTが平均適合率の平均0.345を達成し、12層の逐次型Transformerを相対値で2.1%上回った。同時にパラメーター数は49.4%、積和演算数は42%少なく、実測スループットは9.8%高かった。個別に学習したモデル同士で比較すると、通過回数を2回から10回へ増やすことで、計算量の増加をわずか1.2%にとどめながら精度が向上する。追加の評価では、時間方向のマスキングやその他のさまざまな音響拡張に対する頑健性が向上し、音楽、環境音、イベント音の分類課題でも一般化が改善することが示された。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Increasing depth of transformer models improves recognition, but it comes at a substantial cost. Each additional layer requires more parameters, which makes the process computationally inefficient. We ask whether additional processing can focus on integrating features already computed. Looped Audio Spectrogram Transformer (LAST) first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, thereby making later passes inexpensive. On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. Further evaluations show improved robustness to temporal masking and various other auditory augmentations, with better generalization on classification tasks with music, environmental, and event sounds.

著者のコメント

6 pages, 4 figures, 1 table

arXiv ID: 2610.01926 / 要約の誤りについて