生命科学の配列予測に根拠を求める評価と学習
Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks
この論文をやさしく読む
ひとことで言うと
DNA・RNA・タンパク質の配列予測で、答えの正しさと生物学的な根拠の妥当さを別々に評価した。
何に役立つ?
生命科学向けモデルの推論の弱点を見つけ、根拠を示す配列予測モデルの訓練に役立つ。
この研究の面白いところ
17モデルの評価で分類精度と根拠の質が逆の傾向を示し、その課題に対応する訓練法を複数規模で試した。
どこまで分かった?
近道学習は考えられる解釈として述べられており、原因が確定したわけではない。結果は要旨の6課題と評価モデルの範囲である。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
複数のオミクス配列には複雑な生物学的パターンが含まれるが、その仕組みを解読して科学的発見を自動化するのは難しい。大規模言語モデルが配列を解釈するときは、予測だけでなく科学的な推論も評価する必要がある。既存のベンチマークは分類や回帰の指標に頼り、モデルが根拠となる生物学的証拠を理解しているかを見落としている。そこでDNAの制御、RNAの処理、タンパク質の機能にまたがる6課題、専門家が検証した1,160問からなるOmicsBenchを提案する。事例ごとに専門家と作った基準で、追跡可能な証拠の連鎖を評価する。17モデルの評価では、科学特化型モデルは汎用モデルより配列分類の正確さに優れる一方、予測を支える妥当な証拠を示せないという逆の関係が見られた。一つの解釈として、生物学的機構でなく統計的パターンに頼る近道学習が考えられる。この結果を受け、ツールを使うオンポリシー蒸留TA-OPDを提案し、配列予測と証拠に根ざした推論を合わせる。0.8Bから27Bまでの五つのQwen3.5モデルで、生物学的証拠への根拠付けが一貫して強まり、大半の課題で予測性能も向上した。改善はモデル規模をまたいで続き、推論の強さは容量を増やすだけでなく証拠を重視する訓練でも改善できることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Multi-omics sequences contain complex biological patterns, yet deciphering their mechanisms for automated scientific discovery remains challenging. As large language models (LLMs) interpret these sequences, evaluating both predictions and scientific reasoning is critical. However, existing benchmarks for multi-omics sequence tasks rely on classification and regression metrics, neglecting whether models grasp the underlying biological evidence. We introduce OmicsBench, the first reasoning benchmark for multi-omics sequences, comprising 1,160 expert-validated questions across six tasks spanning DNA regulation, RNA processing, and protein function. OmicsBench requires traceable evidence chains, evaluated using instance-specific rubrics developed with domain experts. Evaluating 17 LLMs reveals an inverse relationship: while scientific LLMs outperform general-purpose LLMs in sequence classification accuracy, they fail to provide valid evidence to support their predictions. One plausible interpretation is shortcut learning: specialized models may rely on statistical patterns rather than the biological mechanisms needed for scientific discovery. Motivated by this finding, we introduce tool-augmented on-policy distillation (TA-OPD), a post-training method to align sequence prediction with evidence-grounded biological reasoning. Across five Qwen3.5 models spanning 0.8B to 27B parameters, TA-OPD consistently strengthens biological evidence grounding while improving predictive performance on most tasks. These gains persist across model scales, indicating that stronger sequence reasoning does not arise solely from increased model capacity, but can be improved through evidence-aware training. Together, OmicsBench and TA-OPD provide a framework for diagnosing reasoning failures in multi-omics LLMs and a path toward models whose predictions are better grounded in biologically meaningful evidence.
著者のコメント
18 pages, 5 figures
arXiv ID: 2609.23435 / 要約の誤りについて