COLIP-2: Olfaction-Vision-Language Embeddings

2026-07-20Robotics

RoboticsArtificial IntelligenceEmerging TechnologiesMachine Learning
AI summary

The authors created a model called COLIP-2 that combines smell, vision, and language into one shared system. It learns from molecules, sensor data, describing words, and images so a robot can guess where a smell is coming from in a scene. Because there aren’t big datasets linking smells and images, they collected some to train the model. Their goal is to show what’s possible with open smell data and push for better methods and datasets to improve smell-based machine sensing, especially for robots. Although made for robots, the model could help in any area that needs understanding of smells along with images and words.

multimodal embeddingsolfactioncontrastive learningmolecular structuregas sensorsodor descriptorsshared representation spacerobotic perceptionreal-time edge computingdataset collection
Authors
Kordel Kade France
Abstract
The Contrastive Olfaction-Language-Image Pre-training 2 (COLIP-2) model is a multimodal embeddings space that places olfaction as a first-class citizen among vision and language. Molecular structure, gas-sensor readings, odor-descriptor language, and images are all trained into a single shared representation space, so that a robot can localize a detected aroma to objects in a scene probabilistically. No ImageNet-scale datasets of paired image-scent examples exists which warrants the need for their collection. Our intent with the release of COLIP-2 is to demonstrate the limit of what can be built for robotics with open-sourced olfactory data in order to ground the argument for why new methodologies and datasets are necessary in order to enable advanced olfactory-oriented perception capabilities. We enumerate results from internal testing of the COLIP-2 architecture and make necessary optimizations to run the model at the edge for real-time robotics applications. While developed with robotics in mind, the design of COLIP-2 has been influenced by experts across many disciplines of science in academia and industry, and we hope that the model can be useful in any multimodal domain requiring olfactory intelligence.