arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

中国語の事実性推論で正解と近さを両立する学習法

UO-FIE: Combining Exact-Label Supervision with Graded Utility for Factivity Inference

Xinchen Xiao

この論文をやさしく読む

ひとことで言うと

中国語の事実性を九段階で判定する課題で、完全一致と正解に近い予測を両立する学習法を提案した。

何に役立つ?

クラスが偏り、正解との距離も評価される分類課題で、教師信号や復号法を設計する参考になる。

この研究の面白いところ

566件の学習例の64.1%が一クラスに集中する条件で、微調整部門の1位を得た。

どこまで分かった?

成績はFIE2026の各部門での評価であり、二つの部門のマクロ効用は直接同じ条件の比較ではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

Factivity Inference Evaluation 2026(FIE2026)は、中国語の文脈と仮説の組を、順序のある九つの事実性区間へ分類する課題である。評価指標は、正しい区間をぴたりと当てることと、正解に近い区間を選ぶことの両方を評価する。一方、学習例566件の64.1%は一つのクラスに集中している。予備実験では、複数のmDeBERTa分類模型が主にその最多クラスを予測し、Huber回帰の基準法は正解に近い区間をより多く予測するが、完全一致は少なかった。 本研究は、正解ラベルの教師信号と段階的な効用を組み合わせる、パラメータ効率の良いUtility-Oriented Factivity Inference(UO-FIE)を導入する。九クラスの確率分布を予測し、確定ラベルによる教師信号、効用に基づくソフトターゲット、段階的に変えるクラス重み、順序損失を組み合わせる。期待効用による復号を条件をそろえて比較し、提出したシステムには、分割外予測から選んだ順序校正を用いた。Qwen3.5-9BにLoRAを適用したUO-FIEは、微調整部門でマクロ効用0.8316を得て1位だった。別のプロンプトベースのアンサンブルは、微調整なし部門でマクロ効用0.8450を得て3位だった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The Factivity Inference Evaluation 2026 (FIE2026) classifies Chinese context-hypothesis pairs into nine ordered factivity intervals. Its evaluation metric rewards both exact predictions and proximity to the correct interval, while 64.1% of the 566 training examples belong to a single class. In preliminary experiments, several mDeBERTa classification models predominantly predict the dominant class, whereas a Huber-regression baseline produces more predictions near the correct interval but fewer exact matches. We introduce Utility-Oriented Factivity Inference (UO-FIE), a parameter-efficient system that combines exact-label supervision with graded utility. UO-FIE predicts a distribution over the nine classes and combines hard-label supervision, utility-based soft targets, scheduled class weights, and an ordinal loss. We evaluate expected-utility decoding in controlled comparisons and use ordinal calibration selected on out-of-fold predictions for the submitted system. Based on Qwen3.5-9B with LoRA, UO-FIE ranks first in the fine-tuning track with a macro utility of 0.8316. A separate prompt-based ensemble ranks third in the non-fine-tuning track with a macro utility of 0.8450.

著者のコメント

11 pages, 6 figures. Accepted as oral presentation at CCL26-Eval

arXiv ID: 2609.28605 / 要約の誤りについて