# The gate was measuring the wrong thing *2026-08-01. Supersedes the correlation threshold used as the project's go/no-go statistic since the synthetic world results.* ## The correlation does not govern recovery The standing rule was that the field correlation at the true pairing has to reach about 0.9, that nothing works below it, and that any new representation can therefore be priced in minutes without running a search. Natural data sat at 0.656 and no recovery run was made on it, on the strength of that rule. **A controlled truncation refutes the rule.** Take the synthetic fields that recover at 95.3%, project both to rank *r*, and match blind from a spectral start with exact refinement. The correlation barely moves across the ladder; recovery moves across its entire range. (What replaces the correlation is *not* the shared-direction count in the next column — see the correction that follows this section. The column is reported because it is what motivated the hypothesis, not because it survived it.) | Field | ρ at truth | Shared directions | Recovery (256 scenes, chance 0.4%) | |---|---|---|---| | synthetic, rank 4 | 0.902 | 11 | 6.2% | | synthetic, rank 8 | 0.906 | 13 | 12.9% | | synthetic, rank 16 | 0.928 | 18 | **95.6%** | | synthetic, rank 32 | 0.928 | 27 | 93.8% | | synthetic, rank 64 | 0.928 | 27 | 95.8% | | synthetic, rank 128 | 0.929 | 27 | 96.1% | | synthetic, full | 0.929 | 27 | 95.8% | A rank-8 field correlating at 0.906 — comfortably past the supposed threshold — recovers 13%. The same field at rank 16 recovers 96%. Correlation is held almost fixed while recovery crosses from failure to success, so **something other than the strength of the agreement is doing the work.** Identifying that something is a separate problem, and the next section records two attempts at it that both failed. The theory says the same thing in advance and we misread it. Correlated-matching thresholds are derived for exchangeable full-rank noise, where each of the N²/2 entries carries an independent constraint on the pairing. A rank-*r* shared component supplies about *rN*. At 256 scenes with *r* near ten that is an eighth of what the threshold calculation assumes, so quoting the 0.9 figure against a low-rank field compares a number to a bound derived for a different object. The correction cuts against us as well as for us. With the old gate retired, the reason for never having run a search on natural data went with it, so the search was run: five blind trials on the best natural field, spectral start chosen by energy over five candidates, exact refinement to convergence. **Recovery is 0.0000 against a chance rate of 0.0039** — the spectral start lands exactly on chance and refinement moves away from the truth, which is what descent does when the truth is not the optimum. Natural data does not recover, and now that is a measurement rather than an inference from a statistic that does not govern it. ## Correction, same day: width is not the statistic either The section below proposed the width of the shared spectrum as the replacement gate. **A control run later the same session refutes it**, and the refutation is recorded here rather than folded away because it was my own claim and it lasted four hours. Adding independent noise to both synthetic fields lowers the correlation without narrowing the shared signal underneath, which is still full-rank. That control was run to guard against overcorrecting, and it did more than that: | Field | ρ | shared width | shared dimension (held out) | recovery | |---|---|---|---|---| | synthetic, noise 0.35 | 0.828 | 15.7 | 16.8 | **94.7%** | | synthetic, caption omits size | 0.830 | 15.0 | 16.6 | **5.6%** | | Visual Genome, best | 0.731 | 15.0 | 17.2 | **0.0%** | | synthetic, noise 0.50 | 0.743 | 13.3 | — | 70.7% | Three fields agree on the correlation to within 0.1, agree on both width measures to within 1.5, and differ in recovery across the entire range. The principal-angle count fails because noise rotates eigenvectors and lowers the measured overlap without narrowing anything. A second instrument built specifically to fix that — canonical correlations fitted on one third of the scenes and scored on another, so a direction counts only if it generalises — fails the same way, and rates Visual Genome *highest* of the three failures. So **two candidate statistics have now been proposed and refuted in one session**, and the honest position is narrower than the one I wrote four hours ago: - **Established.** The correlation does not govern recovery. Rank-8 truncation holds it at 0.906 and drops recovery to 12.9%; independent noise drops it to 0.743 and keeps recovery at 70.7%. Both directions are controlled, and together they retire the 0.9 rule for good. - **Established.** Genuine rank truncation below about 8 destroys recovery, and suppressing a caption factor destroys it while the correlation stays at 0.83. - **Not established.** That any single width or dimension statistic predicts recovery. Neither of the two tried does. What separates the noise case from the omission case is a live question with a specific candidate. Suppressing a factor makes scenes that differed only in that factor *exactly* interchangeable in the text field, so the energy acquires an exact symmetry and the truth stops being a unique minimum — it becomes one member of a degenerate orbit. Independent noise creates no such ties. That predicts a three-way split which the gate can read directly: the truth deeper than everything found (recoverable), tied with what is found (symmetry), or shallower than what is found (information deficit). The gate is running on both fields; the sections below should be read against this correction. ## What the shared spectrum measures Each field is eigendecomposed and the principal angles between the two leading 32-dimensional eigenspaces are computed; directions with cosine above 0.7 count as shared. A scene-shuffled null puts the chance count at 1.0 in every condition below, so the counts are not an artefact of subspace dimension. | Field | ρ | vision eff. rank | text eff. rank | shared (net of null) | |---|---|---|---|---| | synthetic watershed (recovers 95.3%) | 0.928 | 14.5 | 11.8 | **26** | | synthetic v0 (fails, 5.1%) | 0.508 | 31.3 | 27.9 | 18 | | Visual Genome, previous best | 0.656 | 18.7 | 13.9 | 10 | | Visual Genome, best today | 0.716 | 39.7 | 48.0 | 15 | The failing synthetic field settles the sufficiency question in the other direction: 18 shared directions with a correlation of 0.508 also fails. **Recovery needs both a wide shared spectrum and a strong correlation, and neither alone predicts it.** *(Superseded by the correction above: the pair does not predict it either. The counts below remain as measured; what they do not support is the inference that they govern recovery.)* ## Natural data is narrow in the intersection, not in either modality The diagnosis is sharper than "the features are not good enough". Each side is individually rich — vision effective rank 40, text 48 — while the two agree on 15 directions. **Each modality is rich about something, and they are rich about different things.** That intersection is close to invariant under everything tunable on the vision side, all measured at matched width: | Vision configuration | ρ | shared | |---|---|---| | 6 spectral segments | 0.692 | 15 | | 16 spectral segments | 0.679 | 14 | | annotated region boxes (oracle) | 0.703 | 15 | | DINOv2-base, 12 layers, 768-dim | 0.731 | 15 | | DINOv2-large, 24 layers, 1024-dim | 0.709 | 15 | The backbone row was a pre-registered prediction rather than a hope: the width had failed to respond to every other vision-side change, so a larger encoder was predicted to leave it alone. **Doubling the depth and widening the representation leaves the shared count at exactly 15** and moves the correlation by less than the segmentation noise floor, in the wrong direction. The prediction survives, and the row closes on a measurement. Handing the pipeline ground-truth boxes buys one direction over unsupervised segmentation and nothing over using fewer segments. This confirms from a new angle what the earlier oracle-box comparison found: **segmentation is not the constraint on photographs.** The correlation column in that table should be read against a noise floor. Moving the segmentation eigendecomposition from CPU to GPU changes nothing in the recipe, yet re-deriving segments through it takes the baseline correlation from 0.6559 to 0.6767 — the eigenvectors differ in sign and, where eigenvalues are near-degenerate, in rotation, so the clustering that follows lands differently. **Segmentation reseeding is worth about 0.02 in correlation**, so the 6-versus-16-segment and oracle-box differences in that table are inside the noise and only the shared counts distinguish them. The session's headline movements, 0.656 to 0.716 to 0.725, are three times the floor. The pipeline itself is unchanged: run against the original CPU-derived segments it reproduces 0.6559 exactly. The text side does move it, and saturates: | Text vector dimension | text eff. rank | ρ | shared | |---|---|---|---| | 24 | 19.0 | 0.661 | 10 | | 48 | 31.8 | 0.692 | 15 | | 128 | 48.0 | **0.716** | 16 | | 256 | 53.0 | 0.700 | 17 | Tripling the text representation's own rank from 19 to 53 buys seven shared directions and then stops. Enlarging the scene population does not help either — at matched width the shared count goes 23, 20, 19 for 256, 512 and 1024 scenes, so the ceiling is a property of what the two corpora are about rather than of how many scenes are sampled. ## Corpus overlap sets the width, by intervention The invariance results say the width does not come from the encoders. A controlled intervention says where it does come from. In the synthetic world the caption is encoded into separate factor blocks, so a block can be suppressed — the caption then never states that property while vision continues to see it. Nothing else changes: same images, same segmentation, same kernel. | Caption content | ρ | shared | recovery | |---|---|---|---| | states everything | 0.928 | 20 | **96.4%** | | never states size | 0.830 | 15 | **5.6%** | Deleting one word class costs five shared directions and the entire result. The correlation is still 0.830 — higher than anything we have achieved on photographs — and recovery is at chance. **The synthetic world degraded to Visual Genome's shared width fails exactly the way Visual Genome fails**, at a correlation Visual Genome never reaches. The full ladder is monotone in how much the caption states: | Caption content | ρ | shared | recovery | |---|---|---|---| | states everything | 0.928 | 20 | 96.4% | | never states size | 0.830 | 15 | 5.6% | | never states count | 0.678 | 13 | 1.3% | | states neither size nor count | 0.544 | 10 | 1.3% | | never states colour | 0.736 | 9 | 0.0% | What this establishes is that **corpus overlap controls recovery**, which is the claim that matters for what to do next. What it does not establish is the intermediate step — that it does so *through* the width, since the noise control reaches the same width with recovery intact. The mechanism is under test; the corpus-level conclusion does not depend on which way that test goes, and it reframes the natural-data deficit as a statement about Visual Genome rather than about our pipeline: a region description simply does not say as much about a photograph as the synthetic captions say about their scenes. (Widths in this table are computed on standardised fields and are not comparable entry-for-entry with the tables above, which use raw fields; within the table the convention is identical, which is what the comparison needs.) ## Two smaller results, one refuted hypothesis **The projection improved.** Scaling each discriminant direction by its own eigenvalue to the power 0.5 lets a wide basis be kept without the weak directions drowning the strong ones. It dominates the unweighted projection at every setting tested and, unlike the unweighted version, improves with width — which is the direction that raises rank. With text vectors at 128 dimensions this takes the field correlation from 0.656 to **0.716** and the shared count from 10 to 16, at no cost. **Hubness was the wrong suspect.** The hypothesis was that the correlation is inflated by a shared "this scene resembles everything" component carrying no matching information. It is not: the additive row-and-column model accounts for 0.9% of the visual field's variance and 4.9% of the text field's, and removing it *raises* the correlation slightly, from 0.656 to 0.671. The headline statistic was honest; it was simply not the statistic that governs recovery. ## Shrinking the problem is dead, and now by measurement Small blocks were the last standing non-correlation route, on the reasoning that the hard phase is joint in correlation and population size. The reasoning had the sign backwards. The information-theoretic threshold is ρ_IT ≈ √(4 log N / N), which **rises** as N falls: 0.29 at 256 scenes, 0.51 at 64, 0.83 at 16. Shrinking the population raises the bar. The project's own gate confirms it directly. At every size tested, the best state a strong searcher reaches — spectral start plus 200 random restarts, each run to a local optimum under exact steepest descent — is **deeper than the truth**, in three replicates out of three: | N | ρ | E(truth) | E(best found) | accuracy of best | truth deepest? | |---|---|---|---|---|---| | 16 | 0.705 | 2.95 | 2.37 | 0.271 | 0/3 | | 24 | 0.703 | 3.24 | 2.29 | 0.194 | 0/3 | | 32 | 0.679 | 3.69 | 2.38 | 0.073 | 0/3 | | 48 | 0.706 | 3.48 | 2.65 | 0.125 | 0/3 | | 64 | 0.700 | 3.24 | 2.15 | 0.151 | 0/3 | | 96 | 0.698 | 3.65 | 2.57 | 0.115 | 0/3 | When the truth is not the optimum, no searcher of any cost finds it, and the question of algorithm class does not arise. Accepting exponential cost buys nothing here — this is the same conclusion tempering reached by equilibrating below the truth, now established at sizes small enough that search cannot be blamed. Combined with the earlier rank-ladder result — at natural field quality every truncation from rank 4 to full returns chance — three of the five routes in the candidate register are closed: reshaping the landscape, searching harder, and shrinking the problem. Raising the shared spectrum and adding orthogonal signal are what remain. ## What this changes The target moves from a scalar to a pair. Raising the correlation from 0.656 to 0.716 was worth having, but the measurement that matters is that it came with the shared count going from 10 to 16, and 16 is inside the interval where the synthetic ladder crosses from 13% to 96%. The corpus, not the encoder, is where the width comes from. The synthetic world has 26 shared directions because its captions state exactly the world state; Visual Genome has 15 because a region description and a patch descriptor overlap on roughly that many aspects of a photograph, and no amount of segmentation quality, backbone capacity, or scene count moves it. That reframes the next step as a question about which corpora have naturally wide overlap — dense descriptions, product listings with photographs and full specifications, screenshots paired with their accessibility trees — rather than a question about better features for this one.