arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

英語と中国語の面接記録を使ううつ症状重症度の尺度間転移

Cross-Scale Transfer Learning for Depression Severity Prediction: From PHQ-8 to HAMD-17 Across Languages and Clinical Paradigms

Wenjie Feng, Sahba Zojaji, Satoshi Nakamura

この論文をやさしく読む

ひとことで言うと

英語の面接データで学んだモデルを中国語の臨床相談へ転移し、うつ症状スコアの予測を試した。

何に役立つ?

考えられる用途は、少ない面接データでの尺度間転移の研究である。要旨は診断やスクリーニングで使えると証明したものではないと明記する。

この研究の面白いところ

PHQ-8からHAMD-17へ言語と面接方式もまたいで転移し、患者単位の交差検証で複数の対照と比較した。

どこまで分かった?

単一施設の探索的な内部評価であり、尺度・言語・面接方式の寄与は分離できていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

本研究は、データが少ない状況で、臨床面接の文字記録からうつ症状の重症度スコアを連続値として予測する課題を扱う。尺度をまたぐ転移のため、低ランク適応(LoRA)を順に行う手順を提案する。まず、範囲を制限した回帰ヘッドを持つQwen3本体を、英語のDAIC-WOZデータセット、アバターを介した189面接とPHQ-8の得点で追加学習する。次に、そのアダプターを初期値として、中国語のPDCHデータセット、実際の臨床相談100件とHAMD-17で追加学習する。この段階では尺度ごとに初期化し直したヘッドで臨床家が付けた得点を予測する。すべての設定で、患者単位に層別化した5分割を2回繰り返す交差検証を使う。 データの少ないHAMD-17への転移では、0.6Bと1.7Bの両モデルで、逐次方式が平均絶対誤差(MAE)、二乗平均平方根誤差(RMSE)、マクロF₁の点推定値で最良となり、転移先データだけでの学習や非LLMの基準手法を上回った。0.6Bでは順に4.96、6.59、0.36、1.7Bでは4.38、5.62、0.46だった。要因を除いた比較から、正しく対応した転移元の教師信号が点推定値では最良だが、教師なしの事前接触やラベルを入れ替えた対照でも一部の改善があることが示唆された。また、中国語の原文入力は英語への機械翻訳入力より良く、学習順序を逆にしても実行ごとの変動を超える明確な改善はなかった。 これは単一施設での探索的な内部評価であり、スクリーニングや診断への有用性を確立するものではない。また、尺度、言語、面接方式の変化それぞれの寄与を分離して特定していない。著者らの知る限り、DAIC-WOZからPDCHへのこの逐次転移設定を評価した先行研究はない。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

This work addresses continuous depression-severity score prediction from clinical interview transcripts under data scarcity. We propose a sequential low-rank adaptation (LoRA) protocol for cross-scale transfer: a Qwen3 backbone with a bounded regression head is first fine-tuned on the English DAIC-WOZ dataset (189 avatar-mediated sessions, PHQ-8), and the adapter then initializes fine-tuning on the Chinese PDCH dataset (100 real clinical consultations, HAMD-17), where a reinitialised, scale-specific head predicts the clinician-assigned score. All configurations use patient-level stratified 5-fold, 2-repeat cross-validation. On the data-scarce HAMD-17 target, the sequential protocol attains the best point-estimate MAE , RMSE, and macro-$F_1$ on both 0.6B and 1.7B backbones, outperforming target-only training and non-LLM baselines---4.96/6.59/0.36 with Qwen3-0.6B and 4.38/5.62/0.46 with Qwen3-1.7B. Ablations suggest that correctly aligned source supervision gives the best point estimates (unsupervised exposure and shuffled-label controls also show partial gains), that native-Chinese target input outperforms machine-translated English input, and that the reversed order yields no clear gain within run-to-run variance. The study is an exploratory, single-site internal evaluation: it does not establish screening or diagnostic utility, nor separately identify the contribution of the scale, language, or paradigm shifts. To our knowledge, no prior study evaluates this specific DAIC-WOZ-to-PDCH sequential transfer setting.

著者のコメント

preprint to ICASSP 2027

arXiv ID: 2609.28430 / 要約の誤りについて