Papers for

data robustness teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Post-hoc method improves frozen graph node classification accuracy

Propagate, Then Sharpen: Post-Hoc Refinement of Frozen Node Classifiers

Abstract: We study post-hoc refinement of frozen node classifiers: given only the graph $G$ and class distributions $Q$ predicted by a frozen model, can we improve accuracy without access to node features, model parameters, or gradients? APPNP answers this by propagating logits with a restart towards the initial predictions, minimizing the anchored Dirichlet energy. Instead, we consider the Potts energy, and decompose it into a Dirichlet term, which penalizes disagreement between neighbouring nodes, and a Gini term, which penalizes indecision within each node. This decomposition motivates Propagate, Then Sharpen (PtS), which alternates between propagation of class probabilities and node-wise, mass-preserving sharpening, with only one additional hyperparameter selected using labelled validation nodes. Across nine homophilic graphs, with a frozen MLP backbone, PtS improves mean test accuracy over independently tuned APPNP by $1.71$ percentage points on clean inputs and $3.90$ under severe Gaussian feature corruption. Gains over APPNP become smaller, but remain positive with frozen GCN and GraphSAGE backbones. Sharpening also removes most of the accuracy loss of deep propagation: on clean inputs without restart, accuracy falls by $2.2$ points between $2$ and $100$ propagation steps under PtS, compared with $33.8$ for APPNP.

Mon 28 SeptMachine Learning
The gist
Sometimes models predict categories for items in a network, but their initial guesses can be improved. The authors found that by spreading predictions across the network and then refining them repeatedly, they could make more accurate guesses without needing extra information or retraining. Their approach worked better than previous methods on several test cases, even when the original data was noisy. This helps improve classifications using only the network structure and frozen model outputs.
Open → 2609.35080v1