1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
|
# World Alignment: recovering cross-modal correspondence from corpora that share no example
## The correspondence can be searched rather than bought
Every multimodal system in use buys its cross-modal correspondence with
paired data. CLIP pays in hundreds of millions of image-text pairs; a
frozen-backbone connector pays less but still pays; unsupervised
captioning pays covertly, through a supervised detector's label space.
We ask what happens when nothing is paid: two corpora, one of images and
one of text, collected separately, overlapping in the world they describe
but sharing no example and carrying no link between them.
The premise is that a correspondence need not be learned if it is already
determined. Two sufficient descriptions of one world must satisfy the
same relational structure, and that structure is a constraint on the
unknown correspondence — enough of it, and the correspondence becomes
identifiable. The method that follows treats the cross-modal bridge as a
variable to be solved for under that constraint, never as weights to be
fitted, which is what separates it from both contrastive pairing and any
cross-modal predictive objective.
## States are sets, relations are fields, and the bridge is a coupling
Four commitments define the method, each forced by a measured failure of
its alternative.
**A scene state is a set of part states, not a pooled vector.** Pooled
embeddings of independently trained encoders share a cross-modal
subspace of roughly sixteen population directions — projecting onto the
top sixteen raises relational agreement from 0.306 to 0.367, and
whitening the full spectrum collapses it to 0.03. Sixteen coarse
dimensions cannot separate five hundred instances, so no energy over
pooled states passes a local-ordering test, and none did across eighteen
conditions.
**Relations are computed within each modality and compared as fields.**
Writing a scene as a set of part states, its relation to another scene is
a moment kernel of the two sets; the field of all such relations is the
object the two modalities must agree on. Fields are invariant to how each
side labels its own features, so no shared vocabulary is presupposed —
a fact we established by accident, when an oracle class alignment and a
frequency-rank alignment produced bit-identical fields while agreeing on
none of twenty-three classes.
**The energy is a functional of one matrix.** With `M` the elementwise
product of the permuted text field and the visual field, the pairwise
term is a constant less `2·sum(M)` and the all-triple term a constant
less `2·tr(M³)/6`, because traces of powers survive simultaneous
row-column permutation. One batched matrix product scores hundreds of
candidate couplings over every triple, and the same identity opens
`tr(M^k)` for any order.
**Solving is spectral first, then local.** A spectral solver reads the
coarse correspondence out of the eigenstructure in polynomial time and
never becomes trapped, but rounds badly; exact steepest descent on the
closed-form energy rounds well but cannot find the basin from a random
start. Composed, they recover; separately, they do not — 53.5% and 7.0%
alone against 95.3% together.
## The gate decides before the search runs
A candidate energy is admitted only if the true configuration is the
deepest state a strong searcher reaches. This sounds procedural and is
the single most productive rule in the project, because the intuitive
substitutes are wrong in a specific way: **descent retention measures the
search operator, not the energy.** Sampled-proposal descent kept the
truth for every candidate energy tested at five hundred scenes; long
tempering on the same energies reached 10% below it. Three times a limit
looked intrinsic and proved to belong to the solver — retention, a
third-order energy that appeared to pass, and a value-alignment ceiling
of six correct entries in ten that combining two signals lifted to ten.
Each was settled by computing the bound rather than inferring it.
The gate yields a cheap predictor. Recovery tracks a single measurable
quantity, the correlation between the two relation fields at the true
pairing, and the phase boundary sits where correlated-matching theory
puts it: at five hundred scenes the information-theoretic threshold is
near 0.29 and the polynomial-algorithm threshold near 0.9. Nothing we
built worked below 0.9, and recovery appeared as soon as fields crossed
it. Any new representation can therefore be priced in minutes, before a
single search is run.
## A closed world settles the mechanism
A procedurally generated world makes representational sufficiency a
construction rather than a hope: scenes are sets of object groups with
spatial relations, captions state exactly the discrete state in a
43-word vocabulary, and renders resample layout so that what views share
is precisely the world state. Vision-only and text-only splits are
disjoint scenes.
Three unimodal operations — appearance grouping, distance-transform
watershed for touching objects, and one-dimensional clustering of radial
extent for size classes — take the field correlation from 0.570 to
0.928, and blind recovery from 5.1% to **95.3% exact scene
identification** at 256 scenes against 0.39% chance, replicated at 96.1%
and 96.1% over two further hidden shuffles and 84.7% at 600 scenes.
Recovery would be a curiosity if it could not be spent. Treated as
pseudo-pair supervision for a ridge map and applied to 200 held-out
scenes that took no part in matching, it reaches **93.0% exact
retrieval** against 0.5% chance, with colour-set F1 of 0.989 and count
multisets exact 97.5% of the time. The random-pair control lands exactly
on chance and the shuffled-image control drops colour F1 to 0.295, so
the output is conditioned on the image and extrapolates beyond the
matched population. Matching is a bootstrap for the map, not the
deployment mechanism, which is why population size binds far less than
the recovery curve alone suggests.
Cross-modal value correspondence is derived, not declared. Text factor
families come from distributional structure, vision classes from
clustering raw appearance, and the two are paired by marginal frequency
rank and colour-by-shape joint structure — read from 6,000 text-only and
6,000 vision-only scenes with no shared instance. The derivation is
exactly right on ten of ten colour terms, and survives a world where
captions name only half the objects present and the pixel and text
frequency rankings diverge in the tail.
## Natural data is not there yet, and the deficit is located
Ported to Visual Genome under the same recipe — DINOv2 patch features,
spectral segmentation, distributional word statistics — the field
correlation reaches 0.656, against the 0.9 recovery needs. No recovery
run is warranted, and the trajectory that got there is more informative
than the number.
Discretisation cost more than half the signal: hard class codes give
0.166 where continuous states give 0.489, because fields need only
scene similarity and quantising states first is attrition the synthetic
world hid by being discrete already. Segmentation is not the wall on
photographs — annotated boxes buy 0.061 over unsupervised spectral
segmentation, the inverse of the closed world where segmentation was the
entire gap, since flat untextured scenes are adversarial for object
discovery in a way photographs are not. Content projection is the
largest lever at 0.489 to 0.656. General-purpose frozen encoders
underperform corpus-fitted statistics, 0.19 against 0.656, so in-context
features beat isolated crops and corpus-specific vectors beat a language
model's pooled states.
## What is not shown
The closed-world result uses object states built by hand; no
self-supervised recipe we tested — orbit contrastive, masked pixel
reconstruction, their hybrid, masked feature prediction, or slot
attention at two resolutions — produced their equivalent, each failing
by a distinct measured mechanism. Natural data sits 0.24 below the
recovery threshold. No frozen language model has been driven end to end
on natural data by a recovered correspondence, though a paired control
establishes that the interface works once the correspondence is right.
The untried lever is the one the closed world identified as a protocol
requirement: two-sided multi-view observation. Visual Genome offers one
photograph per scene, so orbits must come from augmentation or from
corpora that carry them natively. Whether that closes 0.656 to 0.9 is
the next measurement, and the gate will answer it before any search is
run.
|