arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

表形式基盤モデルの予測能力を軽量モデルへ蒸留する

Distillation of Tabular Foundation Models into Efficient Predictors

Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo

この論文をやさしく読む

ひとことで言うと

表形式の基盤モデルが出した予測を使って、データセットごとの軽いモデルを訓練し、繰り返しの推論を速くする研究です。

何に役立つ?

同じ表データの課題で予測を何度も実行する場面で、教師モデルを毎回呼ぶ費用を抑える用途が考えられます。要旨では性能と推論速度の両方を比較しています。

この研究の面白いところ

教師に見せるラベル付き文脈には訓練集合全体を使い、生徒には実データと合成データへの教師予測を学習させるという手順を、別のベンチマークにも変更なしで適用しています。

どこまで分かった?

結果は二つのTFMと検討した生徒モデルでの評価です。速度は教師との比較で、Eloの比較は調整・アンサンブル済みの教師あり生徒との比較、TALENTの比較はデフォルトの生徒との比較であり、比較対象を混同できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

表形式基盤モデル(TFM)は、文脈内学習によって高い予測性能を達成するが、ラベル付きデータを条件として毎回取り込むため、推論コストが高い。知識蒸留により、その予測能力を軽量でデータセット専用の生徒モデルへ移せば、このコストを削減できる。ただし、TFMの予測はラベル付き文脈とクエリの両方に依存するため、教師の指導信号をどう構成するか、そしてクエリの範囲を広げると蒸留が改善するかという二つの設計上の問いが生じる。 二つのTFMと、ニューラル型・木構造型の生徒モデルの双方について、これらの問いを調べ、有効な蒸留手順を導く。この手順では、ラベル付き訓練集合の全体を教師の文脈とし、観測されたクエリと合成クエリに対する教師の予測だけで生徒を訓練する。TabArenaでは、得られた生徒モデルは、教師あり学習を行い、調整とアンサンブル化を施した対応モデルを57〜98 Eloポイント上回る。同じ手順を変更せずTALENTへ適用すると、300データセット中236〜258データセットで対応するデフォルト設定の生徒モデルを改善し、主要誤差の中央値を4.0〜6.4%減らす。 蒸留した生徒は、教師に対して中央値3.0〜21.6倍の推論高速化も達成し、予測性能と繰り返し推論のコストとの間で実用的な折り合いを提供する。コードは https://github.com/nums-ai/TFM_Distillation で公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Tabular foundation models (TFMs) achieve strong predictive performance through in-context learning, yet repeatedly conditioning on labeled data makes inference expensive. Knowledge distillation can reduce this cost by transferring their predictive ability to lightweight, dataset-specific students. However, the dependence of TFM predictions on both a labeled context and a query introduces two design questions: how to construct teacher supervision and whether expanding query coverage improves distillation. We examine these questions across two TFMs and both neural and tree-based students, and derive an effective distillation recipe. The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. On TabArena, the resulting students outperform their supervised trained tuned-and-ensembled counterparts by 57-98 Elo points. Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets and reduces median primary error by 4.0-6.4%. The distilled students also achieve median inference speedups of 3.0-21.6 times over their teachers, offering a practical trade-off between predictive performance and repeated inference cost. Code is available at https://github.com/nums-ai/TFM_Distillation .

arXiv ID: 2610.01435 / 要約の誤りについて