arXiv論文メモ
新着一覧
cs.CL / cs.AI / cs.IR / cs.SE · 査読状況未確認

生テレメトリから業務意味層を誘導する枠組み

Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction

Yuanzhe Jia, Ali Anaissi

この論文をやさしく読む

ひとことで言うと

雑多なアプリのログから、業務上の機能や指標へつなぐ意味の層を自動で組み立てる方法です。

何に役立つ?

手作業のパーサーや意味対応の保守を減らし、生ログを業務分析へ結びつける用途です。

この研究の面白いところ

業界知識を加えた大きな機能分類の後、検索・フィルタリング・クラスタリング・統一命名で細かな業務ノードを作ります。人の意味品質評価は100点中50から80超へ上がったと報告します。

どこまで分かった?

保守負担80%削減や雑音74%除去は評価した本番規模データの結果です。要旨には評価組織やサンプル数の詳細がなく、LLM評価のκ=0.87もすべての場面の品質保証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

アプリケーションの大量で異質な生テレメトリを業務上の洞察へ変換するのは難しく、データエンジニアは意味の不一致、解析ロジック、KPIとの対応を保守する。本論文は生ログから業務意味層を自動構築する枠組みを示す。まず業界知識を加えたLLM推論で高水準の業務特徴を特定し、次にデータ精錬、ハイブリッド検索、多段階フィルタリング、意味クラスタリング、標準名付けで細粒度ノードを導出する。実運用規模テレメトリで、人手評価の意味品質を100点満点で50点から80点以上へ高め、保守作業を80%削減し、ノイズの74%を除去し、LLM-as-JudgeでCohenのκ=0.87を達成した。ラベル付き訓練データや手作業ルールなしに生テレメトリから意味層を誘導する問題を扱う。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Modern applications generate massive volumes of raw telemetry data, but translating those noisy, heterogeneous event streams into actionable business insights remains a fundamental challenge. Data engineers and analysts expend substantial effort reconciling semantic discrepancies, hand-crafting parsing logics, and maintaining fragile mappings between raw data and business KPIs. In this paper, we present an end-to-end framework that fully automates the construction of a business semantic layer from application raw logs. Our approach introduces a two-stage semantic abstraction: first, high-level business features are identified via LLM inference augmented with domain-specific industry knowledge; second, fine-grained business nodes are derived through a structured pipeline comprising data refinement, hybrid retrieval, multi-stage filtering, semantic clustering, and canonical naming. Evaluation on production-scale telemetry demonstrates that our system improves human-assessed semantic quality from 50 to 80+ on a 100-point scale, reduces maintenance effort by 80%, filters out 74% of noise, and achieves 0.87 Cohen's kappa via an integrated LLM-as-Judge evaluation, enabling continuous, scalable quality assurance. Overall, our work distinguishes itself from prior work by addressing the novel problem of business semantic layer induction from raw telemetry, operating without labeled training data or manual rule engineering.

arXiv ID: 2609.19615 / 要約の誤りについて