2023/10/28 by Taha Ameen, Ameen, Taha, Bruce Hajek +1 · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Physics and Astronomy · #Applications (stat.AP) #Bayesian Modeling and Causal Inference #Complex Network Analysis Techniques #FOS: Computer and information sciences #FOS: Mathematics #Gene Regulatory Network Analysis #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.2310.18543
openalex publication_date 2023/10/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Two models are introduced to investigate graph matching in the presence of corrupt nodes. The weak model, inspired by biological networks, allows one or both networks to have a positive fraction of molecular entities interact randomly with their network. For this model, it is shown that no estimator can correctly recover a positive fraction of the corrupt nodes. Necessary conditions for any estimator to correctly identify and match all the uncorrupt nodes are derived, and it is shown that these conditions are also sufficient for the k-core estimator. The strong model, inspired by social networks, permits one or both networks to have a positive fraction of users connect arbitrarily. For this model, detection of corrupt nodes is impossible. Even so, we show that if only one of the networks is compromised, then under appropriate conditions, the maximum overlap estimator can correctly match a positive fraction of nodes albeit without explicitly identifying them.