arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

施設ごとに未収集の項目も補完する連合学習

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, and Stavroula Georgia Mougiakakou

この論文をやさしく読む

ひとことで言うと

施設ごとに測っていない項目が違う場合にも、患者の生データを集約せずに欠損項目を補う方法を評価した。

何に役立つ?

施設間で取得項目が異なる共同研究における、表形式データの補完方法の選択に役立つ。臨床での予測改善そのものを実証した結果ではない。

この研究の面白いところ

通常の個々の値の欠損に加え、施設全体で一度も測られていない特徴量の欠損を明示的に評価している。

どこまで分かった?

要旨は合成・実世界データを使う複数シナリオでの補完誤差を報告する。患者の診療成果や、データ保護の法的適合性までは示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

複数施設の臨床研究や生物医学研究では、単一施設を超えて一般化するモデルを作るため、施設をまたぐデータの活用が求められている。その際、患者の生データを施設間で共有することを規制が制限し得る一方、施設ごとに収集する特徴量の集合が、異なる手順の下で部分的にしか重ならないという、2つの課題がある。連合学習なら、生データを一か所に集めずに共同でモデルを学習できる。しかし既存の連合学習による欠損値補完法は、ある施設では特徴量そのものが一度も観測されない、特徴量単位の欠損をほとんど評価していない。この状況に対応するため、マスク付きオートエンコーダReMaskerを連合学習向けに適応させたFed-ReMaskerを提案する。共同参加施設から学んだ知識を使い、各施設で一度も観測していない特徴量も補完できる。線形・非線形の関係を持つ合成データと、臨床データを含む実世界の表形式データを対象とするベンチマークで評価した。施設数、欠損率、参加施設間の異質性を変化させた。同質な設定のベンチマークでは、値単位の欠損シナリオの93.2%、特徴量単位の欠損シナリオの96.7%で補完誤差が最小となった。単純な連合平均を用いても施設間の異質性に頑健で、値単位の36シナリオすべてで全比較手法を上回り、特徴量単位では各比較手法に対して36シナリオ中少なくとも35で優れた。また、統合データで学習した中央集約モデルとの差は平均3.0%以内だった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Multi-center clinical studies and biomedical research collaborations increasingly seek to utilize data across centers to build models that generalize beyond any single center. This creates two distinct challenges: data protection regulations may restrict the sharing of raw patient data across institutions, while centers may collect only partially overlapping sets of features under different protocols. Federated learning enables collaborative model training without centralizing raw data. However, existing federated imputation methods rarely evaluate feature-level missingness, in which entire features are unobserved at some centers. To address this setting, we adapt the ReMasker masked autoencoder to federated learning (Fed-ReMasker), enabling centers to impute features never observed locally by leveraging knowledge learned across collaborating centers. We evaluate Fed-ReMasker in a benchmark spanning synthetic datasets with linear and nonlinear relationships and real-world tabular datasets, including clinical data. The benchmark varies the number of centers, the missingness ratios, and client heterogeneity. Fed-ReMasker achieves the lowest imputation error in 93.2% of value-level and 96.7% of feature-level scenarios in the homogeneous benchmark. It also remains robust to client heterogeneity using simple federated averaging, outperforming all baselines in all 36 value-level scenarios and each baseline in at least 35 of 36 feature-level scenarios, and comes within 3.0% on average of a centralized model trained on the pooled data.

arXiv ID: 2609.28105 / 要約の誤りについて