arXiv論文メモ
新着一覧
cs.AI / cs.CV / cs.IR / cs.LG · 査読状況未確認

文書・映像・センサーデータを共通基盤で検索するUniK

UniK: Universal Knowledge Perception for Digital and Physical AI

Nirmit Desai, Kunal Sawarkar, Aditya Mahakali, Dongkon Lee, Kevin Park, Eric Song

この論文をやさしく読む

ひとことで言うと

社内文書、映像、分子情報、センサー記録などを取り込み、整理して検索する仕組みを、対話AIとロボット学習で共通化する提案です。

何に役立つ?

形式の異なる知識をAIから利用しやすくし、検索に基づく回答を改善する用途があります。要旨の具体的な正解率は政府データや医学質問応答などのデジタルAI評価です。

この研究の面白いところ

課題ごとのモデル微調整ではなく、情報を拡充した複数の索引を融合する検索側の仕組みを中心に据えています。700億パラメータモデルと組み合わせた成績を報告しています。

どこまで分かった?

比較モデルの規模や優越性は著者らの記述として扱う必要があります。要旨にはロボットの制御成功率や世界モデル学習の改善量はなく、デジタル領域の数値をフィジカルAIの実証結果とはみなせません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

組織の活動を変える二つのAIシステムが広がっている。一つは、企業の知識について推論し、チャットボットやエージェントの業務フローを動かすデジタルAIである。もう一つは、映像、ゲームプレイ、センサーのテレメトリーから、ロボットや自律システムの制御を学ぶフィジカルAIである。両者は共通の基礎的な障害に直面する。異なるモダリティにまたがる大量の生の知識が非公開コーパスに閉じ込められ、既存のAI基盤からは信頼性と効率を保ってアクセスできないことである。 両者に共通する基盤としてUniversal Knowledge Perception(UniK)を提案する。これは、リッチテキストや映像から分子データ、センサーテレメトリーまで、さまざまなモダリティを通じて、取り込み、情報拡充、索引化、検索、継続的評価という知識のライフサイクル全体を扱う。UniKは、自動的に情報を拡充した索引群を融合するPolymath Retrievalを基礎とし、課題固有の微調整を必要としない。 デジタルAIの5領域、すなわち医学文献、オープンドメイン質問応答、化学、法的手続きの映像、政府公開データにおいて、700億パラメータのオープンソースモデルと組み合わせたUniKは、桁違いに大きい最先端の独自開発LLMに一貫して匹敵するか、それを上回ると報告する。政府データでの検索拡張生成(RAG)の正解率はGPT-5の47%に対して76%であり、医学質問応答では微調整なしで77.9%となり、化学ではすべてのオープンソースのパイプラインを上回った。同じ基盤は、フィジカルAIの世界モデル学習におけるデータ整備、索引化、検索の課題にも直接対応することを示す。そこでの知識の問題はより難しいものの、構造的には同じである。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Two transformative classes of AI systems are reshaping how organizations operate: \textit{digital AI}, which reasons over enterprise knowledge to power chatbots and agent workflows; and \textit{physical AI}, which learns to control robots and autonomous systems from video, gameplay, and sensor telemetry. Both face the same foundational bottleneck: raw knowledge at scale, spanning heterogeneous modalities, locked in private corpora that existing AI infrastructure cannot access reliably or efficiently. We propose \textit{Universal Knowledge Perception (UniK)} as a common platform for both classes, covering the full knowledge lifecycle (ingestion, enrichment, indexing, retrieval, and continuous evaluation) across modalities from rich text and video to molecular data and sensor telemetry. We present UniK, built on Polymath Retrieval (multi-index fusion over automatically enriched indices) with no task-specific fine-tuning. Across five digital AI domains (medical literature, open-domain QA, chemistry, legal video proceedings, and government open data) UniK combined with an open-source 70-billion-parameter model consistently matches or outperforms frontier proprietary LLMs that are orders of magnitude larger: 76\% RAG accuracy on government data versus 47\% for GPT-5; 77.9\% on medical QA without fine-tuning; topping all open-source chemistry pipelines. We show that the same infrastructure directly addresses the data curation, indexing, and retrieval challenges facing physical AI world model training, where the knowledge problem is harder but structurally identical.

著者のコメント

17 pages

arXiv ID: 2609.23971 / 要約の誤りについて