arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

水処理文献に特化したWaterBERTで情報抽出と検索を改善

Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining

Mudi Zhai (1), Ruihong Qiu (2), Qingyun Zeng (3,4), T. David Waite (1), Bing-Jie Ni (1), Haoran Duan (1,5) ((1) UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia (2) School of Electrical Engineering and Computer Science, The University of Queensland, Brisbane, QLD 4072, Australia (3) Microsoft Copilot Studio AI, Redmond, WA 98052, United States (4) Departments of Mathematics & Department of Computer and Information Science, University of Pennsylvania, Philadelphia, PA 19104, United States (5) Department of Civil Engineering, The University of Hong Kong, Pokfulam, Hong Kong SAR, China)

この論文をやさしく読む

ひとことで言うと

水処理の専門文献で追加学習したモデルを作り、処理法の分類、物質などの抽出、関係の整理、文献検索に使っています。

何に役立つ?

大量の水処理文献から必要な研究を探し、どの処理法や対象が扱われているかを整理する支援になります。69万件超の要旨で知識グラフを構築したと報告されています。

この研究の面白いところ

個別の抽出課題だけでなく、トピック発見と知識グラフを使った検索までつなげて評価しています。

どこまで分かった?

F1スコアと検索の関連度スコアは異なる尺度です。関連度77.7は正答率77.7%という意味ではなく、要旨には尺度の定義や費用の具体額はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

水処理研究は急速に拡大しているが、得られた知識の多くは非構造化された文献に分散したままである。この分野には、大規模な文献マイニングのために、水処理に固有の意味を効率的に捉えられる専用言語モデルが依然として不足している。そこで、水処理の文章からの意味表現と構造化情報抽出を目的とした、分野適応型のエンコーダモデルWaterBERTを開発する。WaterBERTは、約29億7000万トークンからなる大規模な水処理コーパスで継続事前学習して開発した。WaterBERTに基づく三つの微調整モデルを下流タスクで系統的に評価した結果、汎用および分野特化のBERTモデルの中で全体として最良の性能を達成した。F1スコアは、処理プロセスの多クラス分類で90.12%、固有表現認識で79.50%、関係抽出で74.04%だった。これらのベンチマーク課題に加えて、大規模な文献処理でもWaterBERTの利点を示した。Environmental Science & Technologyの5144本の論文に適用したWaterBERT-BERTopicは、あらかじめ分類項目を定めず、一貫性と多様性があり、分野に即した研究トピックを同定した。さらにWaterBERTを用い、商用LLMより大幅に低い費用で、競争力のある抽出性能を保ちながら69万3211件の要旨を処理し、水処理の構造化知識グラフを構築した。この知識グラフを語彙検索および密ベクトル検索と統合し、水知識強化検索システムWaterKERSを開発した。同システムの関連度スコアは77.7で、文章に基づく検索の比較手法の54.7〜64.5を大きく上回った。本研究はWaterBERTを通じて、水処理研究における大規模情報処理とエビデンスの整理に向けた、コンパクトで拡張可能な意味表現の基盤を提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Water treatment research is expanding rapidly, but much of the knowledge acquired from this research remains scattered across unstructured literature. The field still lacks a dedicated language model that can efficiently capture water treatment-specific domain semantics for large-scale literature mining. Here, we address this by developing WaterBERT, a domain-adapted encoder model designed for semantic representation and structured information extraction from water treatment texts. WaterBERT was developed by continual pretraining on a large-scale water treatment corpus comprising about 2.97 billion tokens. Three fine-tuned models based on WaterBERT were systematically evaluated on downstream tasks, achieving the best overall performance among general-purpose and domain-specific BERT models, with F1 scores of 90.12% for multiclass treatment process classification, 79.50% for named entity recognition, and 74.04% for relation extraction. Beyond these benchmark tasks, we further demonstrated WaterBERT's advantages for large-scale literature processing. Applied to 5,144 Environmental Science & Technology articles, WaterBERT-BERTopic identified coherent, diverse, and domain-specific research topics without predefined categories. Building on WaterBERT, we processed 693,211 abstracts at substantially lower cost than commercial LLMs while retaining competitive extraction performance to construct a structured water treatment knowledge graph. The knowledge graph was then integrated with lexical and dense retrieval to develop a Water Knowledge-Enhanced Retrieval System (WaterKERS), which achieved a relevance score of 77.7, substantially outperforming text-based retrieval baselines (54.7-64.5). Through WaterBERT, this study provides a compact and scalable semantic foundation for large-scale information processing and evidence mapping in water treatment research.

arXiv ID: 2609.26034 / 要約の誤りについて