引用による影響を手掛かりに研究案を生成
Learning to Ideate for Scientific Impact
この論文をやさしく読む
ひとことで言うと
過去の論文の引用に基づく指標を使い、より影響が大きいと推定される研究案を生成する方法です。
何に役立つ?
研究案を比較・生成する手掛かりとしての用途が考えられます。実際に将来の引用や研究成果が増えたことを示すものではありません。
この研究の面白いところ
10万件超の論文を使って、研究目標と着想から引用影響ラベルを予測する報酬を学習し、生成モデルを調整します。
どこまで分かった?
引用数は研究の受け入れを表すノイズのある代替指標です。結果は生成案の推定影響度に関する評価であり、将来の実際の影響を測ったものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
科学的な着想に大規模言語モデルが使われる機会が増えているが、従来の着想システムは、目新しさ、明確さ、実現可能性など、すぐに判定できる代替指標で訓練・評価されることが多い。そこで本研究は、後から分かる研究の受け入れを示す信号を使い、期待される影響の大きい研究方向へモデルを導けるか調べる。学術界での受け入れを測る、ノイズはあるが大量に得られる代替指標として、引用数を正規化した影響度を用いる。 計算機科学の論文10万件超から、研究目標を条件とする着想の説明を抽出し、各論文へ年ごとに正規化した引用の順序ラベルを付けて、大規模なデータセットを作る。次に、研究目標と着想の組から引用に基づく影響ラベルを予測する報酬モデルを訓練し、その報酬で、教師あり微調整に続く強化学習を通じて着想生成器を調整する。循環的な評価を減らすため、生成案を、同じ研究目標の下で過去の着想と比較し、参照した着想の引用影響ラベルで判定に重みを付ける、保留データに基づく評価手順を使う。実験では、強化学習で調整したモデルが、元のモデルおよび教師あり微調整だけの比較手法より、一貫して推定影響度の高い着想を生成した。この結果は、研究成果に基づく影響度が、開かれた科学的発見において言語モデルを調整する実用的なフィードバック信号となり得ることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}. We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea's citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery.
著者のコメント
RLxF Workshop ICML 2026
arXiv ID: 2609.29802 / 要約の誤りについて