arXiv論文メモ
新着一覧
cs.SI / stat.ME · 査読状況未確認

二部ネットワークの小構造を解釈する前に次数の影響を点検

Auditing bipartite motif interpretations: a worked example with conservation checks and open-path decomposition

Tengfei Shao

この論文をやさしく読む

ひとことで言うと

ネットワークに多い小さな形を見て意味を付ける前に、各点のつながりの本数だけで決まる部分を差し引くべきだと検証する研究です。

何に役立つ?

顧客行動や訪問行動のネットワークで、構造の頻度を過剰解釈しないための点検に役立ちます。元データの個数の整合性も確認対象にしています。

この研究の面白いところ

統計的に珍しい不足が頑健に出ても、その不足を2つの要因へ数式で分けただけでは原因が分かったことにならないと示しています。削除感度と成分間の相関まで調べています。

どこまで分かった?

推論に用いた具体例は観光ネットワーク1件で、高級品の集計は来歴と辺数の整合性に問題があるため記述的に扱っています。64.4%という分解比率も不安定で、因果的な寄与率ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

観光客と訪問先、顧客と商品取引のような、主体と対象からなる二部ネットワークのモチーフ分布は、構造上の役割やネットワーク間の違いの証拠として読まれるが、2種類の次数列だけで何がすでに固定されているかを問わないことが多い。単純二部グラフでは、一方のノード型における誘導kファンの個数は次数の組み合わせの和であり、両方の次数列を保存する帰無モデルの下で分散はゼロになる。 この既知の結果を、観光客17人、観光地80か所、辺637本からなる再構成した観光評価ネットワークに適用する。これを推論に使う唯一の具体例とし、さらに、来歴に制約のある例示として、36か月分の高級品顧客・商品ネットワークについて公表されたモチーフ個数の集計にも適用する。4種類のファンは次数列の厳密な関数である。観光ネットワークの生のファン個数と、サイズ3の2種類のファンの比率である84.7%のファンアウトは、それらの次数列を言い換えているにすぎない。高級品の公表個数には少なくとも89502本の顧客・商品間の辺が必要だが、報告された取引数は26451件であり、99.8%のファンインは記述値としてのみ報告する。 次数を厳密に保存する二部配置帰無モデルと比べると、4サイクル数は次数と整合的でzは約+1.0、開いたパスは帰無平均より2243個、2.4%少なく、zは約−5.8となる。この不足は観光客を1人ずつ除いたすべての再解析でも残り、zは−4.5~−8.1である。厳密な恒等式により点推定での不足を混合64.4%、4サイクル35.6%に分けられるが、この分解は原因の帰属ではない。観測された混合項は500個の帰無標本すべてより低く、2成分は帰無モデルの下でほぼ共線的(r=0.968)で、観光客を1人ずつ除くと混合の割合は36.6~116.5%に変動する。 この不足は標本化した帰無モデルに比べて極端である一方、クラス単位の解釈は未確定で不安定である。解釈前に行う4段階の点検と参照実装を提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Motif profiles of bipartite agent-object networks, such as tourist-site visits and customer-item transactions, are read as evidence about structural roles and about differences between networks, often without asking what the two degree sequences already fix. In a simple bipartite graph the induced k-fan count on one node type is a sum of degree combinations, so it has zero variance under a null that preserves both degree sequences. We apply this known result to a reconstructed tourism rating network of 17 tourists, 80 sites and 637 edges, the sole inferential worked example, and, as a provenance-limited illustration, to published motif-instance aggregates over 36 monthly luxury customer-item networks. The four fan classes are exact functions of the degree sequences: in the tourism network the raw fan counts and the size-3 two-fan ratio (84.7% fan-out) restate those sequences. The published luxury counts require at least 89,502 customer-item edges against 26,451 reported transactions, so their 99.8% fan-in is reported as a descriptive value only. Against a hard bipartite configuration null, the four-cycle count is degree-consistent (z about +1.0) and the open path is deficient (z about -5.8) by 2,243 instances, 2.4% of the null mean; the deficit survives every leave-one-tourist-out re-run (z -4.5 to -8.1). An exact identity splits it at the point estimate into 64.4% mixing and 35.6% four-cycle, but that split is not an attribution: the observed mixing term lies below all 500 null samples, the two components are almost collinear under the null (r = 0.968), and the mixing share ranges from 36.6% to 116.5% under leave-one-tourist-out deletion. The deficit is extreme relative to the sampled null, while its class-level interpretation is undetermined and unstable. We give a four-step pre-interpretation check and a reference implementation.

著者のコメント

52 pages, 8 figures, 10 tables. Submitted to PeerJ Computer Science. Analysis code and cached null ensembles: https://doi.org/10.5281/zenodo.22308052 ; tourism rating matrix: https://doi.org/10.5281/zenodo.22299150

arXiv ID: 2609.22014 / 要約の誤りについて