Polyp image synthesis improves colonoscopy data with adaptive mucosal context

Beyond the Foreground: FOV-Aware Polyp Image Synthesis via Lesion-Guided Adaptive Mucosal Context Propagation

Computer Vision and Pattern Recognition

Summary

Finding enough annotated colonoscopy images with polyps is hard, which slows progress in medical imaging. The authors present a new method called LAMP that creates realistic synthetic colonoscopy images with polyps by carefully copying details from the lesion and surrounding mucosal tissue. This technique avoids mistakes caused by black non-tissue areas and produces better textures and lighting. Their approach helps train segmentation models that can better detect polyps, potentially improving colonoscopy image analysis.

What this means in practice

  • For medical image analysts: Generate more realistic polyp images to improve training of colonoscopy image segmentation models.
  • For computer vision engineers: Enhance synthetic image generation techniques for images with complex foreground and background regions using lesion-guided context propagation.

Authors

Tong Wang, Yuting He, Bin Ren, Yutong Xie, Guanyu Yang

Abstract

Synthetic image and mask pairs can alleviate scarce colonoscopy annotations, but realistic synthesis requires preserving the supplied lesion while generating compatible mucosa. Existing foreground-guided methods treat all non-foreground pixels as background and rely mainly on local integration. Directly applying them to colonoscopy causes two problems: non-mucosal black regions contaminate generated tissue, and local reasoning produces inconsistent mucosal texture and illumination. We propose LAMP, the first foreground-guided framework for polyp image synthesis based on lesion-guided adaptive mucosal context propagation. LAMP explicitly separates the lesion, valid mucosa, and camera exterior using a field-of-view (FOV) mask. Lesion-to-Mucosa cross-attention extracts lesion appearance conditions for valid-mucosa locations, while FOV-constrained multidirectional Vision Receptance Weighted Key Value propagates them over legal tissue support. An adaptive gate then controls their residual fusion into the diffusion U-Net. Extensive experiments on five polyp datasets demonstrate that LAMP substantially outperforms existing methods in overall generation quality and consistently improves five downstream segmentation models. Our code will be released at https://github.com/wangtong627/LAMP.