# World Alignment: recovering cross-modal correspondence from corpora that share no example ## The correspondence can be searched rather than bought Every multimodal system in use buys its cross-modal correspondence with paired data. CLIP pays in hundreds of millions of image-text pairs; a frozen-backbone connector pays less but still pays; unsupervised captioning pays covertly, through a supervised detector's label space. We ask what happens when nothing is paid: two corpora, one of images and one of text, collected separately, overlapping in the world they describe but sharing no example and carrying no link between them. The premise is that a correspondence need not be learned if it is already determined. Two sufficient descriptions of one world must satisfy the same relational structure, and that structure is a constraint on the unknown correspondence — enough of it, and the correspondence becomes identifiable. The method that follows treats the cross-modal bridge as a variable to be solved for under that constraint, never as weights to be fitted, which is what separates it from both contrastive pairing and any cross-modal predictive objective. ## States are sets, relations are fields, and the bridge is a coupling Four commitments define the method, each forced by a measured failure of its alternative. **A scene state is a set of part states, not a pooled vector.** Pooled embeddings of independently trained encoders share a cross-modal subspace of roughly sixteen population directions — projecting onto the top sixteen raises relational agreement from 0.306 to 0.367, and whitening the full spectrum collapses it to 0.03. Sixteen coarse dimensions cannot separate five hundred instances, so no energy over pooled states passes a local-ordering test, and none did across eighteen conditions. **Relations are computed within each modality and compared as fields.** Writing a scene as a set of part states, its relation to another scene is a moment kernel of the two sets; the field of all such relations is the object the two modalities must agree on. Fields are invariant to how each side labels its own features, so no shared vocabulary is presupposed — a fact we established by accident, when an oracle class alignment and a frequency-rank alignment produced bit-identical fields while agreeing on none of twenty-three classes. **The energy is a functional of one matrix.** With `M` the elementwise product of the permuted text field and the visual field, the pairwise term is a constant less `2·sum(M)` and the all-triple term a constant less `2·tr(M³)/6`, because traces of powers survive simultaneous row-column permutation. One batched matrix product scores hundreds of candidate couplings over every triple, and the same identity opens `tr(M^k)` for any order. **Solving is spectral first, then local.** A spectral solver reads the coarse correspondence out of the eigenstructure in polynomial time and never becomes trapped, but rounds badly; exact steepest descent on the closed-form energy rounds well but cannot find the basin from a random start. Composed, they recover; separately, they do not — 53.5% and 7.0% alone against 95.3% together. ## The gate decides before the search runs A candidate energy is admitted only if the true configuration is the deepest state a strong searcher reaches. This sounds procedural and is the single most productive rule in the project, because the intuitive substitutes are wrong in a specific way: **descent retention measures the search operator, not the energy.** Sampled-proposal descent kept the truth for every candidate energy tested at five hundred scenes; long tempering on the same energies reached 10% below it. Three times a limit looked intrinsic and proved to belong to the solver — retention, a third-order energy that appeared to pass, and a value-alignment ceiling of six correct entries in ten that combining two signals lifted to ten. Each was settled by computing the bound rather than inferring it. The gate also produced our longest-lived mistake, and correcting it is the paper's second result. Recovery appeared to track one measurable quantity — the correlation between the two relation fields at the true pairing — with the phase boundary where correlated-matching theory puts it, near 0.9 at five hundred scenes. We priced representations against that number for months. **It does not govern recovery.** Truncating the synthetic fields that recover to rank eight leaves the correlation at 0.906, past the supposed threshold, and recovery at 13%; the same fields at rank sixteen recover 96%. Correlation is nearly fixed across the ladder while recovery crosses its whole range. The theory said so in advance and we misread it. Matching thresholds are derived for exchangeable full-rank noise, where each of the N²/2 field entries constrains the pairing independently; a rank-*r* shared component supplies about *rN*, an eighth as many at 256 scenes with *r* near ten. What replaces the correlation is an open question, and we report two failed answers rather than one plausible one. The obvious candidate is the width of the shared spectrum, and it does not survive its own control: independent noise that lowers the correlation to 0.828 without narrowing the underlying signal reaches the same measured width as a field whose captions omit one factor — 15.7 against 15.0 — with recovery at 94.7% and 5.6% respectively. A second instrument built to fix that, counting canonical directions that generalise to held-out scenes, rates the three failing fields at 16.6 to 17.2 and ranks the worst of them highest. **What is established is that the correlation is not the governing statistic; what is not established is what is.** ## A closed world settles the mechanism A procedurally generated world makes representational sufficiency a construction rather than a hope: scenes are sets of object groups with spatial relations, captions state exactly the discrete state in a 43-word vocabulary, and renders resample layout so that what views share is precisely the world state. Vision-only and text-only splits are disjoint scenes. Three unimodal operations — appearance grouping, distance-transform watershed for touching objects, and one-dimensional clustering of radial extent for size classes — take the field correlation from 0.570 to 0.928, and blind recovery from 5.1% to **95.3% exact scene identification** at 256 scenes against 0.39% chance, replicated at 96.1% and 96.1% over two further hidden shuffles and 84.7% at 600 scenes. Recovery would be a curiosity if it could not be spent. Treated as pseudo-pair supervision for a ridge map and applied to 200 held-out scenes that took no part in matching, it reaches **93.0% exact retrieval** against 0.5% chance, with colour-set F1 of 0.989 and count multisets exact 97.5% of the time. The random-pair control lands exactly on chance and the shuffled-image control drops colour F1 to 0.295, so the output is conditioned on the image and extrapolates beyond the matched population. Matching is a bootstrap for the map, not the deployment mechanism, which is why population size binds far less than the recovery curve alone suggests. Cross-modal value correspondence is derived, not declared. Text factor families come from distributional structure, vision classes from clustering raw appearance, and the two are paired by marginal frequency rank and colour-by-shape joint structure — read from 6,000 text-only and 6,000 vision-only scenes with no shared instance. The derivation is exactly right on ten of ten colour terms, and survives a world where captions name only half the objects present and the pixel and text frequency rankings diverge in the tail. ## Natural data is not there yet, and the deficit is located Ported to Visual Genome under the same recipe — DINOv2 patch features, spectral segmentation, distributional word statistics — the field correlation reaches 0.731 against 0.928 for the synthetic world, and blind search returns 0.0000 against a chance rate of 0.0039. The trajectory that got there is more informative than the number. Discretisation cost more than half the signal: hard class codes give 0.166 where continuous states give 0.489, because fields need only scene similarity and quantising states first is attrition the synthetic world hid by being discrete already. Segmentation is not the wall on photographs — annotated boxes buy 0.061 over unsupervised spectral segmentation, the inverse of the closed world where segmentation was the entire gap, since flat untextured scenes are adversarial for object discovery in a way photographs are not. Content projection is the largest lever at 0.489 to 0.656. General-purpose frozen encoders underperform corpus-fitted statistics, 0.19 against 0.656, so in-context features beat isolated crops and corpus-specific vectors beat a language model's pooled states. What does localise the gap is an intervention rather than a statistic. In the closed world the caption is encoded in separate factor blocks, so one can be suppressed — the caption stops stating a property while vision still sees it, and nothing else changes. Deleting a single word class takes recovery from 96.4% to 5.6% at a correlation of 0.830, higher than anything we reach on photographs, and the full ladder is monotone in how much the caption states. **Corpus overlap controls recovery**, whatever the mediating quantity turns out to be. Encoder-side changes do not. Five interventions on the vision side — near-tripling the segments, substituting annotated boxes for unsupervised segmentation, third-order moments in the set kernel, and doubling the backbone to DINOv2-large — each raise the vision field's own effective rank and leave both the correlation and every width measure where they were. The gains that did arrive came from the projection, the text vectors and the phrase encoding, and total 0.075 in correlation. This makes corpus choice a first-class variable rather than a fallback: a region description does not say as much about a photograph as the closed world's captions say about their scenes, and no vision encoder repairs that. ## What is not shown The closed-world result uses object states built by hand; no self-supervised recipe we tested — orbit contrastive, masked pixel reconstruction, their hybrid, masked feature prediction, or slot attention at two resolutions — produced their equivalent, each failing by a distinct measured mechanism. Natural data recovers no better than chance. No frozen language model has been driven end to end on natural data by a recovered correspondence, though a paired control establishes that the interface works once the correspondence is right. The lever the closed world named as a protocol requirement — two-sided multi-view observation — has now been tried on photographs and returned nothing: four augmentation orbits leave the correlation at 0.6559 either way. Visual Genome offers one photograph per scene, so this tests augmentation rather than genuine re-observation, and corpora carrying native orbits remain untested. But the width result reframes what to look for in them. The question is no longer whether a lever closes a scalar gap; it is whether a pair of corpora describes enough of the same world for their shared spectrum to be wide, and that is a property to select corpora on rather than a deficiency to engineer around.