長い特許文書から発明の構造を抽出し先行技術を探す
From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search
この論文をやさしく読む
ひとことで言うと
特許の全文から発明の関係図を作り、先行技術の検索に使う解析器です。短い文章で学習しても長い文書を処理できるようにしています。
何に役立つ?
数万トークンの特許を途中で切り捨てず、構造を使って検索する用途に役立ちます。報告された検索改善は引用再現率による評価です。
この研究の面白いところ
規則ベースの解析器を教師として学習しながら、検索結果と推論コストでは教師を上回ります。近くの要素だけを採点する工夫で計算量を減らしています。
どこまで分かった?
引用再現率の改善0.5%と1.1%が相対比かポイント差かは要旨に明示されていません。法的な新規性判断の正確さを実証したという内容でもありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
特許検索では、数万トークンを超える文書を日常的に処理する必要がある。多くのニューラル検索手法は入力を切り詰めて扱うため、有効性が制限される。グラフに基づく検索は、各特許を構造化された発明グラフで表すことで対処するが、グラフの構築は頑健性に乏しい規則ベースの解析器に依存している。 本研究では、依存構造解析の双アフィン注意を応用し、特許本文から直接発明グラフを予測するニューラル解析器を提案する。局所双アフィン注意は、要素対の採点を移動窓の内部に限定し、計算量をO(n²)からO(n・w)へ減らす。局所的な採点と大域的な採点は同じ重みを共有するため、短い系列で学習したモデルを再学習することなく、4万トークンを超える文書に適用できる。規則で解析した100万文書から蒸留して学習したモデルは、推論コストを3分の1に抑えながら教師を上回る。下流のGraph Transformer検索システムでは、ニューラルグラフによって引用再現率が短い問い合わせで0.5%、全文書で1.1%改善する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Patent search requires processing documents routinely exceeding tens of thousands of tokens. Most neural retrieval approaches operate on truncated inputs, limiting their effectiveness. Graph-based retrieval addresses this by representing each patent as a structured invention graph, but constructing these graphs relies on brittle rule-based parsers. We present the neural parser, which adapts biaffine attention from dependency parsing to predict invention graphs directly from patent text. Our local biaffine attention restricts pairwise scoring to a sliding window, reducing complexity from $O(n^2)$ to $O(n \cdot w)$. Since local and global scoring share the same weights, the model trains on short sequences and deploys on documents exceeding 40,000 tokens without retraining. Distilled from 1 million rule-parsed documents, it surpasses its teacher at 3$\times$ lower inference cost: neural graphs improve citation recall by 0.5% on short queries and 1.1% on full documents in a downstream Graph Transformer retrieval system.
著者のコメント
Accepted for publication at the ECML PKDD 2026 conference (Applied Data Science track)
arXiv ID: 2610.01553 / 要約の誤りについて