TART:ギター音声からタブ譜生成のフレームワーク
TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription
短い要約(全文訳を準備中)
ギター音声からタブ譜を生成するための新しい手法TARTを提案。表現技術や弦・フレットの割り当てを改善し、ノイズのある音声でも性能を向上させた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-10(UTC)
- 最新改訂
- 2026-09-10 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-10 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Automatic Music Transcription (AMT) for guitar remains limited by three challenges: existing systems often fail to capture expressive techniques such as slides, bends, and percussive hits; they often assign notes to incorrect string-fret combinations; and they are typically trained on clean recordings, limiting their generalization to noisy real-world audio. To address these challenges, we propose TART, a modular four-stage audio-to-tablature pipeline consisting of (1) an audio-to-MIDI transcription model, (2) an expressive technique classifier, (3) an audio-conditioned T5 encoder-decoder for string-fret assignment, and (4) an automated tablature generator. We evaluate TART in a zero-shot setting on GuitarSet, EGDB, and two augmented benchmarks, Noisy GuitarSet and Noisy EGDB. Averaged across these four benchmarks, TART achieves 81.35% audio-to-MIDI F50 (+6.67 points over the best prior baseline), 71.8% string-fret Tab F1 (+8.5 points over the best prior baseline), and 54.08% end-to-end Tab F1. To our knowledge, TART is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio.
著者のコメント
ISMIR 2026
arXiv ID: 2609.11904 / 要約の誤りについて