Transformer Geometry Observatory TGO-III: Semantic Geometry Observatory
2026-08-03 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors studied how Vision Transformers (ViTs), a type of AI model for images, organize information as they learn. They created a tool called TGO-III to track how the model’s understanding of different image classes improves across layers during training. Their findings show that the model’s representations become easier to separate by class, with clearer distinctions and more structured patterns emerging over time. This supports the idea that the model’s internal representations expand and organize semantically as learning progresses.
Vision TransformerRepresentation LearningSemantic GeometryClass SeparabilityFisher RatioLinear Probe AccuracyManifold ExpansionPrincipal Component AnalysisCovariance Structure
Authors
Kaustubh Kapil, Kishor P. Upla
Abstract
With the widespread adoption of Vision Transformers in modern AI, the need to analyze their inherent representational behavior has become increasingly important. While most existing studies emphasize token geometries and training dynamics, the evolution of representational covariance structures and class-level geometric organization remains comparatively underexplored. In this work, we investigate semantic geometry and class separability as representations evolve across the layers of ViT-Small/16 through TGO-III: Semantic Geometry Observatory. It is a framework designed to analyze the emergence of semantic organization, feature evolution, and class-wise representation geometry throughout training. The framework employs multiple complementary observatories, including Linear Probe Accuracy, Fisher Ratio, Class Centroid Distances, Local Intrinsic Dimension, and Local PCA Rank, to quantify the progressive evolution of discriminative representations. Our analysis reveals that class representations become progressively more linearly separable, Fisher discriminability increases, class centroids move farther apart, and local representation manifolds exhibit structured class-dependent geometric complexity. These observations provide empirical evidence supporting the Semantic Expansion Hypothesis, suggesting that the manifold expansion observed in previous observatories is accompanied by the progressive organization of representations into increasingly discriminative semantic structures. Collectively, TGO-III extends the Transformer Geometry Observatory framework by establishing a direct connection between manifold geometry, covariance evolution, and semantic organization during Transformer training.