転写因子の結合配列を使うゲノム言語モデルの分割法
Motif-Vocab: StatisticallyCalibrated Transcription-Factor-Identity Tokenization forGenomic Language Models
この論文をやさしく読む
ひとことで言うと
DNA配列を言語モデルに入力する際、転写因子の結合モチーフを意味のある単位として切り出す方法を提案した。
何に役立つ?
考えられる用途は、転写因子の結合配列が重要なゲノム予測課題で、モデルの入力表現と解釈性を改善すること。
この研究の面白いところ
異なるモチーフを共通の統計的な尺度で判定し、転写因子ごとのトークンにする。ランダム化モチーフや一般的なモチーフトークンなど、複数の対照と比較している。
どこまで分かった?
著者は汎用の精度向上策とは位置付けていない。モチーフなしの密なトークン化も強い基準で、5課題のBERT-base評価ではほぼ同等。報告された利点は、特にモチーフに関わる課題でのもの。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
トークン化はゲノム言語モデルの中心的な設計選択だが、DNA向けの多くの方法は、文字、固定長のk-mer、または出現頻度から作った部分語を用い、DNA結合性の制御因子が持つ配列特異性の事前知識を明示的には利用しない。本研究は、生物学的知識を取り入れたトークン化手法Motif-Vocabを提案する。DNAの両鎖を走査し、統計的に校正されたモチーフとの一致から転写因子(TF)の種類を表すトークンを出力し、一致しない配列には塩基、k-mer、またはバイト対符号化(BPE)を適用する。モチーフごとの帰無分布により、長さや縮退度が異なる位置重み行列(PWM)を共通の有意性尺度に置き、決定的な重複処理規則で表現の再現性を保つ。 20億塩基対を使った条件をそろえたBERTの事前学習では、実際のモチーフライブラリが、対象範囲内の下流課題55件中54件でランダム化モチーフの対照を上回った。DART-Eval Task 2から作った、モチーフが重複しない認識課題では、TF固有のトークンは、位置をそろえた汎用モチーフトークンよりmacro-F1が0.040、対応をそろえたモチーフなしのトークン化より0.033高かった。後者の差の95%ブートストラップ信頼区間は0.027〜0.038だった。モチーフトークンは、シャッフルした対照より強い寄与度の割り当てを受け、隠した際の影響も大きかった。一方、密なモチーフなしトークン化は汎用の強い基準手法であり、5課題のBERT-baseパネルではほぼ同等だった。したがってMotif-Vocabは、あらゆる課題の精度を置き換える方法ではなく、モチーフに敏感なゲノムモデルに対象を絞った、解釈しやすい帰納バイアスである。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Tokenization is a central design choice in genomic language models, yet most deoxyribonucleic acid (DNA) tokenizers use characters, fixed-length k-mers, or frequency-derived subwords without explicitly using prior information about the specificity of DNA-binding regulatory factors. We introduce Motif-Vocab, a biologically informed tokenizer that scans both DNA strands for statistically calibrated motif matches, emits transcription-factor (TF) identity tokens, and applies nucleotide, $k$-mer, or byte-pair encoding (BPE) to unmatched sequence. Motif-specific null distributions put position-weight matrices (PWMs) of different lengths and degeneracy on a common significance scale; deterministic overlap rules make the representation reproducible. In controlled Bidirectional Encoder Representations from Transformers (BERT) pretraining on two billion base pairs, real motif libraries outperform randomized-motif controls on 54 of 55 in-scope downstream tasks. On a motif-disjoint recognition task derived from DART-Eval Task 2, TF-specific tokens improve macro-F1 by 0.040 over a position-matched generic motif token and by 0.033 over a matched no-motif tokenizer (95\% bootstrap confidence interval: 0.027--0.038). Motif tokens also receive stronger attribution and produce larger occlusion effects than shuffled controls. Dense no-motif tokenizers remain strong general-purpose baselines, including a near-tie on the five-task BERT-base panel. Thus, Motif-Vocab is not a universal accuracy replacement; it is a targeted, interpretable inductive bias for motif-sensitive genomic modeling.
著者のコメント
9 pages, 4 figures, submitted to the Forty-First AAAI Conference on Artificial Intelligence (AAAI-27)
arXiv ID: 2609.28386 / 要約の誤りについて