健康状態の五段階回答を二値化すると地域差の判断が変わる
The risks of dichotomising ordinal outcomes: A spatial analysis of self-rated health in Western Europe
この論文をやさしく読む
ひとことで言うと
五段階の健康状態の回答を二つにまとめる区切り方で、健康状態が悪いと判定される地域が変わることを示した研究です。
何に役立つ?
公衆衛生の地域比較で、二値指標の区切り方に結論が左右されないか確認する材料になります。
この研究の面白いところ
年齢や教育との基本的な関連は保たれた一方、地域の判断は「普通」の回答をどちらに入れるかで変わりました。
どこまで分かった?
分析対象は欧州社会調査第11回の西欧の男性回答者です。自己評価による健康状態であり、臨床診断の結果ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
調査の回答は順序のある段階で測ることが多い。欧州社会調査の自己評価による健康状態は「非常に良い」から「非常に悪い」までの五段階だが、分析では「良い」と「悪い」の二値にまとめることが一般的である。二値指標は割合を伝えやすい一方で情報を減らし、健康状態の変動を捉えにくくする可能性があり、区切り方が推定値や実質的な結論を左右し得る。地域を扱う分析でその影響を体系的に示した証拠はまだ少ない。本研究は2023~2024年の欧州社会調査第11回を用い、西欧の男性回答者の自己評価による健康状態を、年齢、教育、地域に基づいてモデル化する。事後層化を伴うBayes空間個人レベルモデルで、人口を代表する推定値を得る。五段階の回答に対する順序累積logitモデルと、「普通」の健康状態をどちらへ分類するかが異なる二つの二値化によるBernoulliロジスティック回帰モデルを比較する。 高年齢と低学歴は、どの設定でも自己評価による健康状態の悪さと関連し、基本的な年齢・教育の傾向は比較的頑健だった。一方、関連の大きさと不確実性、および一部の地理的な結論は、回答をどうモデル化するかに敏感だった。「普通」を二値のどちら側に入れるかで、平均より健康状態が不利な地域として特定される場所が変わった。順序モデルは各段階の情報を保ちながら、まとめることで従来型の二値の割合も出せる。二値指標は、とくに監視や情報伝達には引き続き有用である。ただし、区切り方を正当化し、特に地域比較では別の区切り方や順序モデルに対する感度を検討すべきである。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Survey responses are often measured using ordered response categories. In the European Social Survey, self-rated health is measured on a five-point scale from very good to very bad, yet analyses commonly dichotomise responses into binary categories of "good" and "poor" health. Binary indicators provide prevalence measures that are straightforward to communicate, but dichotomisation reduces information and may limit captured health variation, while the cut-off may influence estimates and substantive conclusions. Systematic evidence on these effects in spatial settings remains limited. We address this gap using European Social Survey round 11 (2023/2024) for Western Europe. We model self-rated health among male respondents by age, education and region. Bayesian spatial individual-level models with poststratification provide population-representative estimates. We compare an ordinal cumulative logit model for the five-category outcome with Bernoulli logistic regression models using two dichotomisations differing in the classification of "fair" health. Older age and lower education are consistently associated with worse self-rated health across specifications, suggesting relatively robust fundamental age and educational gradients. However, their magnitude and uncertainty, and some geographical conclusions, are sensitive to how the outcome is modelled. Assigning "fair" health to either side of a binary cut-off changes the regions identified as having above-average levels of less favourable health. The ordinal model retains category-specific information and can produce familiar binary prevalence estimates through aggregation. Binary indicators remain useful, particularly for monitoring and communication. Nevertheless, the selected cut-off should be justified and sensitivity to alternative cut-offs or an ordinal modelling approach should be considered, especially for geographical comparisons.
arXiv ID: 2609.29893 / 要約の誤りについて