Quantum-Grassmann-Plucker Token Mixing for Deep Learning-Based Post-Disaster Damage Assessment
2026-08-31 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionMachine Learning
AI summaryⓘ
The authors tackle the problem of assessing building damage from satellite images after disasters, which is hard because damage can be subtle and varies across events. They applied a method called Grassmann-Plucker token mixing, which captures relationships between image parts without relying on traditional attention mechanisms. They also introduced two new approaches inspired by quantum computing to improve image classification. Testing on tornado damage data showed that their Quantum-inspired Grassmann-Plucker method performed best overall, especially on new, unseen disaster data. Their work suggests this approach can be a useful alternative in this field.
Grassmann-Plucker token mixingVision Transformersatellite imagerypost-disaster damage assessmentquantum-inspired machine learningimage classificationtoken mixingtornado damagemacro-F1 scorecross-event transferability
Authors
Kooroush Farahkhah, Umut Lagap, Taha Rezaei, Saman Ghaffarian
Abstract
Timely post-disaster building damage assessment from satellite imagery is a critical engineering decision support task, yet it remains constrained by class imbalance, ambiguous intermediate damage states, and limited cross-event transferability. This study presents, to our knowledge, the first application of Grassmann-Plucker (GP) token mixing to computer vision and introduces two extensions for image classification: the Quantum-inspired Grassmann-Plucker (QGP) head and the Hybrid Quantum Machine Learning Grassmann-Plucker (HQML-GP) head. The GP head represents multiscale relationships among image patch tokens by encoding subspaces formed by token pairs with Plucker coordinates; QGP enriches these coordinates with amplitude-derived probability features, whereas HQML-GP incorporates expectation values generated by a simulated quantum circuit into the geometric token representation. Paired pre- and post-event image patches from the xBD tornado dataset were processed using a frozen six-channel Vision Transformer base encoder with 16 x 16-pixel patches. The three GP-based heads were compared with multilayer perceptron and Transformer baselines under identical training, checkpoint selection, and evaluation protocols. Joplin and Moore tornado samples were used for model development and seen-event testing, while Tuscaloosa was reserved for unseen-event evaluation. QGP led both test sets in accuracy and macro-F1: 83.46% and 64.50% for the seen events, and 66.45% and 52.70% for the unseen event. Although HQML-GP obtained the highest validation macro-F1 of 65.63%, it did not surpass QGP on either test set and required substantially more training time per epoch. These results establish GP token mixing as a competitive attention-free alternative to conventional Transformer-based token mixing for paired satellite image damage classification.