arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

LLMを評価者として使う設計と報告をLLJ Cardsで整理

LLJ Cards: Best practices for the Use of LLMs as Judges

Khaoula Chehbouni, Melina Medjdoub, Florian Carichon, Golnoosh Farnadi, Jackie Chi Kit Cheung

この論文をやさしく読む

ひとことで言うと

LLMに採点や比較をさせるとき、何を測り、どのように設計し、どう報告するかを整理する枠組みです。

何に役立つ?

LLMによる評価の方法を他者が確認・再現できる形に整えるための指針として利用できます。単にプロンプトを調整する以外の評価実務を扱います。

この研究の面白いところ

測定理論の妥当性や信頼性の考え方を、言語生成と機械学習の評価実務に接続しています。評価者の性能だけでなく、設計と報告の透明性を対象にしています。

どこまで分かった?

要旨は枠組みと目的の紹介で、LLJ Cardsを使った際の精度改善や再現性向上を数値で示していません。導入すれば評価が必ず正しくなるという実証結果ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

近年、大規模言語モデル(LLM)は評価を行う代替手段として広く用いられるようになった。LLMによる評価者(LLJ)と呼ばれることの多いこれらのシステムは、人の判断と比べた性能、拡張性、費用対効果を背景に、幅広い測定課題で研究者や実務家に採用されている。しかし、LLJを評価者として使う際の妥当性と信頼性への懸念を示す研究も増えている。既存の対策は主に、偏りを軽減する技術の開発やプロンプト戦略の改善に重点を置いてきた。これらは重要な前進ではあるものの、主として技術的な修正を提供するにとどまり、標準化され、透明で、再現可能な評価実務が欠けているという、より根本的な課題を残している。 本論文では、測定理論、自然言語生成、機械学習の文献にある実践上の知見を、LLJによる評価のための実用的な指針へまとめた枠組みLLJ Cardsを導入する。LLJは評価を大規模に実施する有望な方法だが、有効に利用するには、妥当性、信頼性、再現可能性を確保する厳密な評価原則が必要である。LLJ Cardsは、自動評価の設計と報告にこれらの原則を適用するための体系的な枠組みを提供することで、この必要に応える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-effectiveness relative to human judgment. However, a growing body of work has shown that the use of LLJs raise concerns about their validity and reliability as evaluators. Existing efforts to address these challenges have largely focused on developing bias-mitigation techniques and refining prompting strategies. While these approaches represent an important step forward, they primarily offer technical fixes and leave a more fundamental challenge unaddressed: the lack of standardized, transparent, and reproducible evaluation practices. In this paper, we introduce LLJ Cards, a framework that synthesizes best practices from measurement theory, natural language generation, and machine learning literature into practical guidelines for LLJ-based evaluations. While LLJs offer a promising path toward scalable evaluation, their effective use requires grounding in rigorous evaluation principles to ensure validity, reliability, and reproducibility. LLJ Cards addresses this need by providing a structured framework for applying these principles in the design and reporting of automated evaluations.

著者のコメント

Prepared for conference submission

arXiv ID: 2609.24516 / 要約の誤りについて