Atomizer IO enables flexible processing beyond image grids
Atomizer-IO: Beyond Pixels, Patches and Grids
Computer Vision and Pattern Recognition
Summary
Most vision systems assume data comes in neat grids like pixels in a photo, which works well for pictures but not for many other types of sensor data that vary in shape or timing. The authors propose Atomizer-IO, a system that treats each data point as an individual unit with its own information and relates them through spatial connections rather than fixed grids. This approach works well even when input data changes formats or comes from unordered 3D point clouds, showing flexibility in handling various sensing setups. It performs comparably with specialized systems designed for specific earth observation tasks and allows control over computing resources during use.
What this means in practice
- •For earth observation teams: Process satellite and sensor data with varying spatial and channel configurations without redesigning models for each new setup.
- •For 3d imaging developers: Apply atomic data processing methods to unordered 3D point cloud data for flexible and efficient inference.
Authors
Hugo Riffaud de Turckheim, Sylvain Lobry, Nicolas Houdré, Damien Robert, Roberto Interdonato, Diego Marcos
Abstract
Most vision architectures assume that observations lie on a regular grid, an effective abstraction for natural images but a restrictive one for sensing data whose channels, temporal sampling, spatial resolution, and geometry can vary. Generic set-based architectures remove the grid, but also remove useful spatial inductive biases. We introduce Atomizer-IO, an architecture that places observations first and derives structure from their physical relationships. Building on top of an atomic representation of the data, each observation is described by its measurement and acquisition metadata, while local cross-attention maps observations to anchor points that can be arbitrarily placed. We evaluate this design by progressively relaxing the grid assumption, from varying input raster configurations and incomplete channel sets to flexible output density and, ultimately, inputs without a raster grid. Atomizer-IO is competitive with flexible EO-specific architectures on most tasks, while offering post-training control over inference cost and competitive compute--performance trade-offs. The same formulation extends without architectural redesign to unordered 3D point clouds, showing that the atomic interface generalizes beyond regular raster inputs. These results suggest that pixels, patches, and grids do not need to define the interface of a sensing architecture.