トルコ語の専門文書で公開重み言語モデルを評価
Evaluating Open-Weight LLMs for Turkish Domain Documents Under Retrieval and Hardware Constraints
この論文をやさしく読む
ひとことで言うと
限られたGPUでトルコ語の長い専門報告書を質問応答に使うとき、モデルと検索の失敗を分けて評価した研究。
何に役立つ?
トルコ語の専門文書向けローカル質問応答システムを選ぶ際、検索方法とモデル性能を別々に確認する設計に役立つ。
この研究の面白いところ
証拠注釈を使い追加のモデル呼び出しなしで失敗箇所を分離した。7種類の検索法を比較したが、文字単位TF-IDFを有意に上回るものはなかった。
どこまで分かった?
5モデル、2文書、各100問、6 GB VRAMのGPUを用いた評価である。正答率49~75%は主ベンチマークの値であり、他のトルコ語文書全般の性能を示すものではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
トルコ語に対応する大規模言語モデル(LLM)の多くは、長く構造の複雑な専門文書ではなく、汎用ベンチマークで評価されている。本研究は、計算資源が限られたローカル環境でのトルコ語文書質問応答について、重みが公開された7B~8B規模のモデル5種類を評価する。主なベンチマークは109ページの産業研究開発報告書から作成し、系統的に検証した100問で構成する。評価手順は、112ページの公共部門報告書と、独立に作成した別の100問でも繰り返す。すべてのモデルを、VRAM 6 GBのNVIDIA RTX 3050ノートPC用GPUで、プロンプト、生成方法、4ビット量子化を統制してローカルに評価する。 方法上の主な貢献は、追加のモデル呼び出しなしで検索の失敗と後段のモデル推論の失敗を分けられる、証拠注釈付きの評価手順である。主ベンチマークでの一連の処理全体の正答率は49~75%だった。さらに、語彙ベース、密ベクトル、混合型の7種類の検索設定を、95%Wilson区間と対応のある厳密McNemar検定で比較したところ、どちらの文書でも文字単位のTF-IDFを有意に上回る設定はなかった。証拠を検索で拾える割合は2つの報告書で異なる形で頭打ちとなり、検索性能と有効な文脈容量が、文書によって制約になる場合とならない場合があることを示す。したがって、トルコ語の専門文書に公開重みLLMを導入する際には、モデルの選択、検索の振る舞い、ハードウェア上の制限を分けて評価する必要がある。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B models for Turkish document question answering under a resource-constrained local deployment setting. The primary benchmark contains 100 systematically validated questions derived from a 109-page industrial R&D report, and the evaluation protocol is replicated using a second 112-page public-sector report and an independently constructed 100-question set. All models are evaluated locally on an NVIDIA RTX 3050 laptop GPU with 6 GB VRAM using controlled prompting, decoding, and 4-bit quantisation. The principal methodological contribution is an evidence-annotated evaluation protocol that separates retrieval failure from downstream model reasoning failure without requiring additional model calls. On the primary benchmark, end-to-end accuracy ranges from 49% to 75%. Seven lexical, dense, and hybrid retrieval configurations are additionally compared using 95% Wilson intervals and exact paired McNemar tests; none significantly outperforms the character TF-IDF baseline on either document. Evidence recall saturates differently across the two reports, showing that retrieval and effective context capacity can be binding constraints for some documents but not others. These results demonstrate that model selection, retrieval behaviour, and hardware limits must be evaluated separately when deploying open-weight LLMs for Turkish domain documents.
著者のコメント
6
arXiv ID: 2609.28007 / 要約の誤りについて