Vision transformer improves sewer defect classification with lightweight models
Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification
Computer Vision and Pattern Recognition
Summary
Sewer systems sometimes have different types of damage that need to be found to keep them working well. Existing computer programs find it hard to quickly and accurately recognize multiple types of sewer problems from images. The authors created a new way to analyze sewer photos using a vision Transformer that combines details at different levels, making it more accurate. They also built smaller models that work well with limited computing power, suitable for on-site inspections. Their tests showed these models can reliably detect sewer defects and work better than earlier methods, even when there is less training data.
What this means in practice
- •For infrastructure maintenance teams: Automatically classify multiple sewer defects from inspection images to support condition assessment and prioritize maintenance tasks.
- •For embedded system developers: Deploy lightweight sewer defect classification models on resource-limited devices used in mobile inspection robots or cameras.$Commercial implications: Enables compact, efficient AI products for automated sewer inspection that can be sold to utilities or inspection service providers.
Authors
Xu Fang, Zhuoran Wang, Qing Li, Shengyu Zhang, Guanzhi Deng, Jianbiao He, Qingquan Li
Abstract
Automated classification of sewer defects is essential for infrastructure condition assessment and maintenance decision-making, but existing deep learning methods struggle to balance classification accuracy and computational complexity in large-scale multi-label scenarios. This study develops Sewer-Transformer-ML, a hierarchical vision Transformer with multi-level feature fusion, together with two lightweight architectures, Sewer-MobileNet-ML and Sewer-Mobile-TransNet, for resource-constrained inspection scenarios. On the Sewer-ML test set, Sewer-Transformer-ML-Base achieved an $F2_{\text{CIW}}$ of 65.68% and an $F1_{\text{Normal}}$ of 92.68%, ranking first on the public leaderboard and exceeding the second-ranked method by 7.6 percentage points in $F2_{\text{CIW}}$. Sewer-MobileNet-ML achieved an $F2_{\text{CIW}}$ of 65.73% with only 17 M parameters, representing an approximately 95% parameter reduction relative to the base model. Under the standard Sewer-Capsule data split, Sewer-Mobile-TransNet achieved 96.43% classification accuracy. When the training set was reduced to 1,177 images, pretraining on Sewer-ML consistently improved model performance. Ablation experiments further showed that direct concatenation was more effective for Transformer features, whereas attention-based fusion better supported multiscale CNN features. These findings provide a computational basis for automated sewer inspection, lightweight model design, and adaptation across civil infrastructure inspection platforms.