元データなしで学習用データセットを生成する評価
Invent a Dataset: Measuring dataset generation abilities with zero seed
この論文をやさしく読む
ひとことで言うと
例となるデータを一つも渡さず、欲しいデータセットの説明だけから学習用サンプルを生成するシステムの評価です。
何に役立つ?
学習させたい能力の実データがない場合に、事後学習用データを準備する用途が考えられます。生成データの評価に加え、それを使ったモデルの学習結果も比較しています。
この研究の面白いところ
データ量を増やすほど多様性の優位性が大きくなったと報告し、生成品質だけでなく後続モデルの性能とのつながりを評価しています。
どこまで分かった?
報告された改善率は相対値です。要旨には品質・多様性指標の具体的な定義や後続学習の数値がなく、あらゆる分野で同じ改善が得られることまでは示していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
データセットの構築は、AI開発の中でも特に手作業が多く、不安定な部分であり続けている。本技術報告では、現実の実務者が直面する、最も極端であると同時に最もよくある設定、すなわちデータが全くない状況に焦点を当てる。この状況では、学習させたい能力について実務者はデータを一切持っていない。 本研究では、データセットの説明から、現実的で大規模な事後学習用データセットを作るプロンプトベースのシステムInvent-A-Datasetを導入する。Anthropic、Google、OpenAI、DeepSeek、Zaiの5つの最先端モデルAPIに対してInvent-A-Datasetを評価する。8種類のタスクと最大2万サンプルのデータセットにわたり、Invent-A-Datasetは大きく上回り、最高の品質(相対的に17%向上)と、最も高いサンプルの多様性(相対的に19%向上)を同時に実現した。 多様性での優位性は学習データセットの規模とともに拡大し、200サンプル時の同等水準から、2万サンプル時には相対的に37%の改善となる。これは後続の学習における大幅な改善につながり、事後学習済みモデルの性能を大きく高める。異なる事後学習モデルのアーキテクチャにわたり、Invent-A-Datasetで微調整したモデルは、他の生成器のデータで微調整したモデルより一貫して上位に位置する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don't have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier model APIs including Anthropic, Google, Open AI, DeepSeek, Zai. Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset significantly outperforms with both the highest quality (17% relative gains) while simultaneously producing the most diverse samples (19% relative gains). Its diversity advantage widens with scale of training dataset size (from parity at 200 samples to 37% relative gains at 20K samples). This translates into considerable downstream training gains, resulting in far more performant post-trained models. Invent-A-Dataset fine-tune consistently ranks higher compared to other generator fine-tunes across different post-trained model architectures.
arXiv ID: 2610.01674 / 要約の誤りについて