arXiv論文メモ
新着一覧
cs.LG / q-bio.QM · 査読状況未確認

細胞予測モデルが入力情報を本当に使うか監査

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

Mengran Li, Bo Li, Chengyang Zhang, Yang Yan, Jinfeng Xu, Zhenchao Tang

この論文をやさしく読む

ひとことで言うと

細胞の反応を予測するAIが、入力された化合物や用量の情報を実際に使い、予測改善につなげているか調べる研究。

何に役立つ?

考えられる用途は、エージェントが自動で見つけた細胞モデルの主張を監査すること。単なる予測スコアでは分からない入力の寄与を評価できる。

この研究の面白いところ

予測の変化と改善を別々に測る。48候補中47は化合物を替えると予測が変わったが、両分割で改善が確かだったのは20候補。

どこまで分かった?

sci-Plexでの改善の対応比較区間はゼロをまたぐ。独立集団では用量の寄与が残る一方、化合物の種類の寄与は裏付けられなかった。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

AIによる仮想細胞は、指定した介入に対する細胞の反応を予測しようとする。しかし、未使用データで予測が良好でも、与えた摂動情報を実際に使ったとは言えない。この予測と主張の隔たりは、言語モデルのエージェントがスコアによるフィードバックを受けて予測器を生成・修正するモデル探索で重要になる。本研究はCELLAUDITを導入し、ある入力が指定された計算経路に入れるか、学習後の予測がその入力に依存するか、その依存が観測された反応の予測を改善するかを調べる。 形態画像とトランスクリプトームを対応づけた摂動ベンチマークBBBC047では、エージェントが選んだ予測器の未使用データにおけるGlobal Pearson相関係数の平均は0.3153だったが、化合物を入れ替えても予測は変わらなかった。対照条件のプロファイルだけを使う予測器は0.3142だった。ソースの点検により、単一要素のキー・バリュー注意機構が化合物情報の問い合わせ経路を遮っていると分かり、対照ウェルを分離して再学習しても予測が変わらない性質は続いた。 関連する二つの課題にまたがる48候補を層別に監査すると、47候補は未使用データの両分割で化合物の入れ替えによって予測が変わったが、目標損失の改善の区間が両分割でゼロを上回ったのは20候補だけだった。BBBC047では、反証に導かれた修正によって、対照プロファイルだけの基準より良い性能を維持しつつ、化合物情報の平均寄与を正にできた。対応づけたsci-Plexの探索では、監査を加えたフィードバックは5回の探索経路で未使用データの性能と化合物・用量の平均寄与を高めたが、対応比較の区間はゼロをまたいだ。独立に取得した集団で固定した設計を再学習すると、予測の一般化は入力利用の主張の一般化を意味しないことが分かった。用量の寄与は残ったが、化合物の種類の寄与を裏付ける証拠は残らなかった。CELLAUDITは、生成・採点・修正のモデル探索に、反証の段階を加える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

AI virtual cells aim to predict cellular responses to specified interventions, yet held-out predictive performance alone does not establish use of the supplied perturbation information. This prediction-claim gap matters in agentic model discovery, where language-model agents generate and revise predictors using score-based feedback. We introduce CELLAUDIT, which audits input-use claims by asking whether an input can enter the cited computation, whether fitted predictions depend on it, and whether that dependence improves prediction of observed response. On a paired morphology-transcriptomics perturbation benchmark (BBBC047), an agent-selected predictor attains a mean held-out Global Pearson correlation coefficient (PCC) of 0.3153 but remains invariant to compound replacement; a control-profile-only predictor reaches 0.3142. Source inspection identifies a compound-query pathway blocked by singleton key-value attention, and the invariance persists after refitting with disjoint control wells. In a stratified audit of 48 candidates across two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 show target-loss gains with intervals above zero on both folds. On BBBC047, falsification-guided revisions recover positive mean compound contributions while retaining gains over the control-profile-only baseline. In matched sci-Plex searches, audit-enriched feedback yields higher held-out performance and larger mean compound and dose contributions across five trajectories, although paired intervals span zero. Refitting fixed designs on an independently acquired cohort shows predictive generalization need not imply generalization of input-use claims: dose contribution persists, whereas support for compound identity does not. CELLAUDIT adds a falsification layer to agentic model discovery, moving from generate-score-revise toward discover-falsify-revise.

arXiv ID: 2609.27234 / 要約の誤りについて