Multimodal tokens unify images and text for better retrieval and generation
FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
Computer Vision and Pattern Recognition
Summary
Combining images and text for tasks like captioning or image generation is hard because the two are usually processed separately. The authors developed a way to turn both images and text into a shared 1D sequence of tokens that work well for both recognizing and creating content. This helps a single model handle searching for images from text, generating images from text, and describing images with text, all with a flexible number of tokens. Their approach improves the quality of these tasks and supports interesting features like blending two images or texts smoothly.
What this means in practice
- •For multimedia software developers: Create applications that search and generate images and captions using a single unified model with flexible-length input tokens.
- •For digital marketing teams: Generate tailored images and captions from text descriptions for advertising with improved consistency and quality by using one adaptable model.$Commercial implications: Enables tools selling automated image and caption creation for marketing campaigns, offering better integration and generation quality.
Authors
Guangyu Sun, Shlok Kumar Mishra, Wentao Bao, Robert Zhenheng Yang, Xiao Wang, Xiyuan Wang, Yujunrong Ma, Chen Yuan, Max Xiangjun Fan, Jun Xiao, Jianpeng Cheng
Abstract
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks generative performance behind frozen embeddings. To bridge this gap, we revisit joint multimodal representation learning and generation to produce linearly interpolatable embeddings that are directly consumable by generative decoders. We present FLAT (Flexible-Length Aligned Transmodal representations), a representation pre-training framework that jointly optimizes a shared multimodal encoder alongside downstream text-to-image (T2I) and image-to-text (I2T) decoders. By combining contrastive alignment with bidirectional cross-modal generative objectives, FLAT ensures its representations function as both discriminative semantic descriptors and generative conditions. Architecturally, FLAT maps visual and textual inputs into a unified continuous 1D sequence space, applying nested dropout over prefix-K tokens to enable dynamic output lengths. A single pre-training stage allows FLAT to perform cross-modal retrieval and generation across variable prefix K, achieving a T2I GenEval score of 71.1. Task-specific fine-tuning aligns model performance with state-of-the-art baselines: 83.1 GenEval on T2I generation; 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO image captioning; and Recall@5 scores of 86.8 (I2T) / 75.8 (T2I) on MS-COCO alongside 98.3 (I2T) / 93.6 (T2I) on Flickr30K. Finally, qualitative evaluations demonstrate that FLAT representations natively support linear interpolation, latent space arithmetic, and zero-shot composed retrieval.