# Synthetic closed world v0: the objective becomes the bottleneck Run date: 2026-07-30 ## Question The natural-data experiments could never verify representational sufficiency: closure and over-completeness were properties of the data we happened to have. The synthetic world makes them constructions. Scenes are sets of object groups (count, size, color, shape) with spatial relation constraints; captions mention exactly the discrete state (vocabulary 43 words); renders realize it with resampled continuous nuisance, so visual orbits share exactly the world state and differ exactly by modality-private variation. 26,000 scenes, 56,000 renders, 3,578 ground-truth intervention variants; vision-only and text-only splits disjoint; every dial is a generator argument. Both towers are trained from scratch on their own modality: a compact ViT (5.3M parameters) with InfoNCE over natural orbit positives -- no flips, which would erase left/right, and no color jitter, which is world content -- and a word-level causal LM (5.6M parameters) to cross-entropy 0.42, near the floor of the templated language. Closure is verified visually and referentially (color-shape pairs unique per scene, exact multiplicity by ring placement). ## Result: perfect closure, collapsed relational sharing On 1,000 held-out pairs, evaluation-only: | Quantity | Synth v0 | Natural (Flickr, frozen) | |---|---:|---:| | Linear CKA | 0.278 | 0.262 | | Pairwise relational Spearman | 0.030 | 0.184 | | Shuffle z | 2.4 | 12.3 | The closed world shares global statistics as strongly as natural data and relational structure six times more weakly. The assignment gate agrees: the true configuration keeps 6.6% of nodes under descent with a counterfeit 26% below it, and content projection -- which on VG passed the basin audit -- rescues ranking power (z = 117, improving swaps 0.35%) but not the basin (retention 24.4%, counterfeit at -13%). ## The mechanism is the objective, not the data InfoNCE with a uniformity term optimizes scene discrimination: every scene is pushed apart from every other, and two scenes sharing "three red circles" gain nothing from sharing it. The vision tower converged to a near-orthogonal instance codebook whose between-scene geometry is arbitrary -- exactly the structure the relational energy needs and exactly what the objective erases. Data sufficiency was total by construction, so the deficit is attributable: **sufficiency of the data does not transfer to sufficiency of the representation; the training objective decides which world geometry survives.** The requirement R1 of `STRUCTURE_DESIGN.md` is a property of the data-objective pair. For the project's central question this is a constructive turn: alignment without pairs needs unimodal objectives whose representations retain compositional world geometry, and the synthetic world measures exactly that retention. The SSL objective joins vocabulary, ontology, scene count, and orbit multiplicity as a phase-diagram axis. ## Objective battery: the two canonical objectives fail in opposite ways Ground-truth factor probes (linear, fitted per modality on its own split) localize what each representation retains. The text tower is nearly lossless and is not the bottleneck. The vision towers: | Probe target | InfoNCE (CLS) | SimMIM (patch mean) | SimMIM (foreground-weighted) | Text tower | |---|---:|---:|---:|---:| | Color presence | 0.992 | 0.572 | 0.717 | 0.997 | | Shape presence | 0.540 | 0.535 | 0.645 | 0.992 | | Group count R2 | 0.61 | 0.12 | 0.11 | 0.94 | | Object total R2 | 0.44 | 0.33 | 0.35 | 0.90 | | Relation count R2 | 0.20 | 0.03 | 0.04 | 0.94 | InfoNCE performs shortcut feature selection: color multisets are nearly unique per scene, so discrimination is solvable from color alone and shape is discarded at chance level. SimMIM inpaints well (reconstruction MSE 0.0088) and pools to almost nothing: the scene is 80% background, the information lives distributed in patch tokens, and no simple pooling recovers it (CKA drops to 0.114, relational z to 0.8; the gate collapses to 0.2% retention). Discriminative-global objectives select shortcuts; reconstructive-local objectives never build a global summary. The naive hybrid (orbit-InfoNCE on the CLS plus masked pixel reconstruction on the patches, weight 25) is falsified: color stays at 0.983, shape stays at chance (0.532), relational Spearman turns negative (-0.063, z = -3.9), and the gate keeps 4.3%. The two losses settle separately -- the CLS keeps the color shortcut while shape remains distributed in patch tokens. Joint pressure on different tokens does not merge into one summary. A first prediction-transfer energy was built and gated on the weak towers: the vision substitution kernel from masked inpainting compared by foreground-patch Chamfer distance, the text kernel from relation-sentence continuation likelihood (relation sentences force referential dependence; whole-caption continuations only reward repetition, and the LM separates same-scene from different-scene continuations at two-to-one per-token likelihood but grades partial overlap poorly). The gated energy reaches z = 2.9 at 512 scenes with 31% improving swaps: falsified as built. Every energy constructed today bottlenecks at the same place -- the text side is near-lossless and the vision side is factor-blind, so the vision objective remains the single point of failure. ## The objective battery, complete: five recipes, five distinct failures | Vision objective | Color | Shape | Groups R2 | Relational rho (z) | Failure mode | |---|---:|---:|---:|---:|---| | InfoNCE (orbit positives) | 0.992 | 0.540 | 0.61 | 0.030 (2.4) | shortcut selection: color only | | SimMIM (pixel targets) | 0.572 | 0.535 | 0.12 | 0.024 (0.8) | local knowledge, no global summary | | Hybrid (InfoNCE + SimMIM) | 0.983 | 0.532 | -- | -0.063 (-3.9) | losses settle separately | | data2vec-lite (EMA feature targets) | 0.743 | 0.545 | 0.18 | 0.059 (3.7) | same, softer | | Slot attention (5-6 slots) | 0.549 | 0.618 | 0.48 | 0.128 (7.9) | slots tile space, not objects | No standard self-supervised recipe at this scale produces factor-complete global states. Slot attention comes closest structurally -- best shape decodability, four times the relational sharing of InfoNCE -- and fails by degenerating into near-uniform spatial tiles (alpha masses 0.13--0.20 across six slots where object binding would put 0.7+ on background), a known failure on flat untextured scenes. Its pooled readout also scrambles per-object factors, which is the set-versus-pool doctrine restated: color drops from 0.99 (InfoNCE pooled) to 0.55 the moment states are sets averaged without correspondence. ## The set-kernel machinery, taken to its ceiling With no learned tower delivering object states, model-free connected-component analysis provides them (the background is flat, so components are objects; ring merging at distance 30 restores multiplicity at 70% count accuracy). Text sets are per-group phrase bag-of-words vectors. Scene-to-scene relations are unweighted Hungarian matching values; the gate runs on these set-kernel fields. Iterations, each fixing one measured defect: | Iteration | z | Improving swaps | Kept | Counterfeit margin | |---|---:|---:|---:|---:| | Slot-vector sets | -0.1 | 45% | 3% | deep | | Decoded sprite sets | 2.7 | 32% | 6% | deep | | CC pixels x LM phrases | 1.8 | 32% | 5% | deep | | CC descriptors x BOW | 4.5 | 13% | 9% | deep | | + unweighted matching | 52.0 | 0.73% | 37% | -11% | | + multi-view field average | 59.9 | 0.13% | 69% | -6.8% | | + shape descriptors | 51.2 | 0.15% | 77% | -4.1% | Three lessons priced along the way: mass-weighting inside the matching destroys similarity grading (0.16 to 0.45 against soft truth); averaging per-view fields cancels segmentation errors (the orbit dividend again); and a missing factor on one side (shape, absent from the first visual descriptor while present in every text set) is exactly the kind of asymmetry counterfeits exploit. ## End to end: close, and honestly not there The best configuration holds 77% of the truth under exact descent, and blind tempering still finds assignments 4--6% below the true energy at chance accuracy. One further asymmetry was identified -- matching values correlate with set sizes differently per modality (0.29 text, 0.13 vision) -- and residualizing both fields against size features does not close the gap either: tempering ends 9.2% below the truth at 0.78% accuracy, because the correction removes shared signal along with the nuisance. Each counterfeit so far has pointed to a specific asymmetry between the two fields; the remaining ones (segmentation error, descriptor-versus-word semantics) resisted this window's budget. The comparison across datasets brackets the identifiability threshold. On VG, both-sides content projection reached cross-modal field validity near 0.61 and the basin audit passed with margins; here the upper-bound fields sit near 0.45 and every recovery run counterfeits. At N = 512, the pass boundary for this energy family lies somewhere between those two numbers -- the first coordinates on the phase diagram this world was built to draw. ## The Fock lift: matching-free set kernels pass at 64 and counterfeit at 512 Replacing the Hungarian matching value with a symmetric-tensor moment kernel -- phi(S) concatenates the degree-1 and degree-2 moments of the object set, scenes compare by normalized inner product -- removes the matching step and with it every measured matching bias, while making composition algebraic and multiplicity an observable. Paired with factor-mirrored one-hot descriptors (color, small/large with medium unmarked as in text, count, k-means shape cluster; size binned after residualizing the renderer's count-dependent radius shrink): - N = 64: the first full basin pass of the synthetic world -- no counterfeit in any quench, best random descent 11% above the truth, 92% retention. - N = 512: retention 75--76% and z = 38--45, but counterfeits return at 18--22% below the truth, and blind recovery stays at chance. The remaining asymmetries -- 30% segmentation count errors and shape-cluster impurity -- are enough at eight times the configuration space. The bracket sharpens: this energy family passes at N = 64 and fails at N = 512 on fields of validity near 0.45, while VG passed at N = 512 near 0.61. Identifiability is a joint (validity, N) boundary, and the world now measures both axes. ## Third-order invariants, and the correction they forced Triangle invariants (for a node triple, the product of the three edge values, compared across modalities) were added to the moment-kernel fields. Under a descent harness that proposes 64 sampled pairs per step, the result looked decisive: at N=512 the truth kept 99.6--100% of its assignment with no counterfeit below it and z = 48--50, replicated over three hidden shuffles. Two controls dismantled that reading. The pairwise-only energy passes identically in the same harness (retention 100%, z = 40), so the third order was not what changed; and long tempering on the triangle energy at N=128 reaches 10.3% below the truth at zero accuracy. What differed from the earlier failing runs was the searcher: exact all-pairs enumeration takes the globally best swap each step, sampled proposals do not, and tempering with a long budget beats both. **Descent retention measures the search operator, not the energy.** The gate statistic is corrected to compare the energy of the truth against the deepest state a strong searcher reaches, and every candidate energy is being re-measured on that basis. The ablation does carry one clean finding that survives: triangle-only has 11.3% improving transpositions against 0.63% for the sum, so the two orders are complementary -- third-order terms deepen the basin while second-order terms smooth the local landscape. ## Precision weighting, tried and instructive The channel-matching idea was tested and fails with a clean mechanism: weighting relation entries by orbit-derived precision (inverse variance across re-rendered views) makes N=64 worse than the unweighted kernel (counterfeit 46% below the truth; descent from the truth keeps 56%). Orbit variance measures reproducibility, not correctness -- the channel splits into variance and bias, view resampling exposes only the variance half, and systematically biased entries are the most reproducible ones, so precision weighting amplifies exactly the entries counterfeits live in. A full Nishimori match needs the cross-modal residual model, which requires pairs. The legal half-channel is now measured to be the wrong half. ## Closed form, and what it revealed The triangle energy has a closed form. Writing M for the elementwise product of the permuted text field with the visual field, the pairwise term is a constant minus 2 sum(M) and the all-triple term is a constant minus 2 trace(M^3)/6, because traces of powers are invariant under simultaneous row-column permutation. One batched matrix product scores hundreds of candidate permutations over every C(N,3) triple, replacing a 200,000-triple sample and a 5.7 ms per-proposal Python loop. Validated against brute force to 3e-7. The same identity opens the whole family trace(M^k): fourth and fifth order closed loops now cost one more matrix product each. Three diagnostics followed, each changing what we thought we knew. **Descent retention measures the searcher.** Sampled-proposal descent keeps the truth for every candidate energy at N=512; exact all-pairs enumeration is stronger; long tempering is stronger still and reaches 10% below the truth. The gate statistic is corrected to compare the truth against the deepest state a strong searcher reaches. **Spectral solvers place us in the hard phase.** Umeyama and GRAMPA run in polynomial time on N x N fields and need no common embedding dimension, so they sidestep the glassy landscape entirely -- and they return 3--13x chance, never recovery. Correlated-matching theory says so: at N=256 the information-theoretic threshold is near rho = 0.29 and the polynomial-algorithm threshold near rho = 0.9, and our fields sit at rho = 0.51--0.57. Identifiable, not efficiently recoverable, exactly as observed. (Our fields are structured rather than Wigner, so the theory guides rather than proves.) **Correlation and identifiability are different axes.** The colour channel alone reaches rho = 0.863, far above the mixed descriptor, but its scenes collide: only 126 distinct colour multisets among 256 scenes, capping any matcher at 49% and collapsing to 25% at N=1000. The eigenvalue degeneracy that breaks spectral methods is that collision structure in spectral form. Full descriptors are 100% distinct, so the information is present; mixing strong and weak channels into one moment kernel is what dilutes rho from 0.86 to 0.51. ## Tier 0: the cross-modal dictionary, derived rather than declared Anchors that map "red" to a hue band are not labels -- no pair is ever consulted -- but they are hand-supplied cross-modal prior. Tier 0 forbids them, and the derivation splits by factor type: ordered factors (count, size) match by their intrinsic order with the residual ambiguity enumerated; unordered factors (colour, shape) can only be matched by distributional statistics. That exposed a defect in world v0: factor values were sampled uniformly, so colour marginals and co-occurrences are symmetric and the correspondence is information-theoretically unrecoverable. Real corpora are the opposite, and this is precisely the structure unsupervised word translation runs on. World v1 samples colour and shape with a Zipf exponent of 1 (colour marginals 0.317 down to 0.035), which is both more realistic and sufficient. On v1 the derivation runs entirely on unimodal statistics of disjoint splits: 1. text families emerge from mutual exclusivity -- tokens of one factor never co-occur inside a group phrase, so a greedy pass recovers {a, two, three, four}, the ten colour words, and {small, large} with no declared list; 2. vision colour classes come from k-means on raw component RGB, with no hue rule; 3. the two sides pair by marginal frequency rank. **The derived dictionary is 10/10 semantically correct.** Tier 0 fields then reach rho = 0.570 -- above the v0 declared-dictionary Tier 1 value of 0.508 -- with a 97.3% identifiability ceiling, and GRAMPA rises to 13x chance at eta = 1. Zero declared cross-modal information. ## First recovery: 54% at Tier 0 The rho decomposition made the target unambiguous. With oracle segmentation the field correlation is exactly 1.0 and the derived dictionary costs nothing (oracle and derived give identical fields); with connected-component segmentation it is 0.5696. **The entire gap was segmentation.** Three unimodal fixes, each measured: | Fix | Measure | Before | After | |---|---|---:|---:| | Gestalt appearance grouping | exact group count | 71.9% | 90.2% | | Distance-transform watershed | exact object count | 74.2% | 100% | | One-dimensional k-means on radial extent | size class | 52.7% | 87.9% | The watershed matters because connected components merge touching ring members: the renderer leaves a gap of 0.37 radii between neighbours, under two pixels for small objects, so a quarter of all scenes lost objects before any grouping logic ran. The size fix matters because area confounds size with shape -- a cross of a given radius has half the area of a circle -- so extent replaces area, and natural breaks replace terciles because the size classes are separated by gaps rather than equally populated in extent. Field correlation rises 0.5696 to 0.6708 to **0.9285**, crossing the polynomial-recovery threshold, and GRAMPA follows it exactly: | Fields | rho | GRAMPA best | |---|---:|---:| | distance-merge | 0.570 | 5.1% | | Gestalt grouping | 0.671 | 9.4% | | watershed + extent sizes | **0.928** | **53.9%** | **Recovery is 53.9--54.7% exact scene identification across three hidden shuffles, against 0.39% chance -- 138x.** No pairs, no declared dictionary, no learned cross-modal map: the colour correspondence is derived from marginal frequency rank, the factor families from mutual exclusivity, the object states from pixels, and the assignment from a polynomial spectral solver on two independently built relation fields. This is the first end-to-end world matching result in the project, and it lands exactly where correlated-matching theory said it would: nothing worked below rho = 0.9, and recovery appeared as soon as the fields crossed it. The phase diagram now carries measured recovery rather than pass/fail: | rho | N | Recovery | Chance | |---:|---:|---:|---:| | 0.570 | 256 | 5.1% | 0.39% | | 0.671 | 256 | 9.4% | 0.39% | | 0.928 | 256 | **53.9%** | 0.39% | | 0.926 | 600 | **26.3%** | 0.17% | Recovery falls with population size at fixed correlation, which is the expected shape: the correlation a polynomial solver needs grows with N. Both large-N and small-N points sit far above chance (158x and 138x), so the boundary is a gradient rather than a cliff in this range. ## Composing the two solvers: 95% The spectral solver and the energy are complementary in exactly the way their failure modes suggest. Spectral reads the coarse correspondence out of the eigenstructure with no search, so it never gets trapped, but its rounding is noisy. Exact steepest descent on the closed-form energy fixes rounding but cannot find the basin from a random start. Composed at N=256: | Stage | Accuracy | |---|---:| | GRAMPA initialisation | 53.5% | | plus pairwise-energy refinement | 86.7% | | plus pairwise-and-triangle refinement | **95.3%** | Scaling is gentler than the spectral numbers alone suggest. At N=600 the spectral initialisation halves to 26.2% but the composed solver only falls to 84.7%, and there the refined energy sits *above* the truth (0.7056 against 0.6623), meaning the truth remains the better configuration and the solver simply stopped short -- a healthier state than the marginal undercut seen at N=256. | N | Spectral | plus pair | plus pair and triangle | Refined energy vs truth | |---:|---:|---:|---:|---| | 256 | 53.9% | 87.5% | **96.1%** | 0.04% below | | 600 | 26.2% | 80.0% | **84.7%** | 6.5% above | The final configuration at N=256 has energy 0.64738 against the truth's 0.64981 -- below it by 0.04% at 95.3% accuracy. That is a different creature from the counterfeits of the previous weeks: not a wrong configuration exploiting a blind spot, but residual field noise making a nearly correct assignment marginally preferable. The remaining 5% is measurement error in the fields, not a failure of identifiability. **Blind recovery of 95.3% of a hidden image-caption correspondence, with no pairs, no declared dictionary, and no learned cross-modal map.** ## The recovered correspondence transfers to unseen scenes Matching is transductive, so the result only matters if the recovered assignment can be spent. Treating it as pseudo-pair supervision for a ridge map from vision states to text states, and applying that map to 200 held-out scenes that took no part in matching: | Measure | Recovered pairs | Random-pair control | Shuffled-image control | |---|---:|---:|---:| | Held-out retrieval, exact | **93.0%** | 0.5% | -- | | Colour set F1 | **0.989** | 0.399 | 0.295 | | Count multiset exact | **97.5%** | 2.5% | 8.5% | | Group count exact | **97.5%** | 17.0% | 58.0% | Chance retrieval is 0.5%. The random-pair control lands exactly on chance, so nothing here comes from the map family or the state geometry; the shuffled-image control collapses colour F1 from 0.989 to 0.295, so the output is conditioned on the image rather than on a description prior. This closes rungs four and five of the evidence ladder inside the synthetic world: the output is image-conditioned, and the correspondence extrapolates to scenes the matcher never saw. Matching is a bootstrap, not the deployment mechanism -- once enough correct pairs exist, the map carries the rest forward. ## What the demonstration actually shows: matching verifies, statistics align The framing matters more than the numbers. Finding a permutation presupposes a bijection between the two corpora, and real unpaired data has none: image and text collections overlap only partially in the world states they cover, one image answers to many descriptions or none, and granularity differs. If the method needed that bijection it would not transfer. It does not. The step that establishes cross-modal correspondence -- the derived dictionary -- reads the text families from the 6,000 text-only scenes and the colour classes from the 6,000 vision-only scenes, and those two sets are disjoint. No instance appears on both sides. The pairing comes from marginal frequency, a statistic of each corpus alone. The 95.3% matching and the 93.0% transfer are therefore verification, not mechanism: they confirm at instance level that the factor coordinates recovered from disjoint corpora are the right ones. On real data, where no hidden truth exists to score, held-out generation against a shuffled-image control replaces matching accuracy as the measurement. The conditions this leaves are weaker than pairing but not vacuous. The two corpora must reflect a similar world distribution, since frequency ranks are only comparable when both sides sample the same world; reporting bias shifts the text marginal away from the visual one, which the bias worlds show degrades the dictionary gracefully (rank correlation 1.00/0.98/0.84/0.77 gives 8/8/6/6 correct entries) rather than breaking it; and co-occurrence structure, which real corpora have in abundance and this world lacks by construction, is the bias-robust upgrade over bare frequency. What remains open is unchanged and should not be blurred by this result: the object states here are hand-built upper-bound descriptors, and no self-supervised recipe tested has produced their equivalent; and the world is procedurally generated. The demonstration establishes that the matching machinery works end to end when the representations are good enough, not that learning produces such representations. ## Reporting bias, and the discipline of separating solver from information The frequency-rank dictionary depends on both corpora ranking values the same way, which real corpora do not: text mentions the salient, pixels count the common. Two worlds test it. World v3 adds colour-shape coupling (each colour prefers one shape) and strong reporting bias, so captions name only 50% of the groups present and the pixel and text colour rankings diverge in the tail (brown moves from ninth to fourth, pink from sixth to tenth). World v4 replaces the cyclic coupling with a distinct Dirichlet shape profile per colour. On v3 the derivation degrades and the recovery of the cause matters more than the number: | Evidence and solver | Correct | |---|---:| | Marginal frequency rank | 4/10 | | Joint shape structure, alternating solver | 6/10 | | Joint plus marginal, exact enumeration of shape permutations | 7/10 | | Oracle shape columns, measured clusters | 8/10 | | Oracle labels, all signals | 10/10 | The first reading was that 6/10 is an information-theoretic ceiling, since four of ten vision colours have near-identical shape profiles -- ten colours cycling over six shapes leaves 10 mod 6 = 4 duplicates, and the twin count came out at exactly four. That reading was wrong, and checking it rather than asserting it is the point. The twins differ sharply in frequency; combining both signals separates all ten (minimum pairwise evidence distance rises from 0.0008 to 0.2322), so the ceiling is 10/10. What remains is a one-entry solver gap and a two-entry estimation gap, the latter because reporting bias leaves common colours with a quarter of their mentions and thins the joint table. On v4, where the coupling gives every colour a distinct shape profile, exact enumeration recovers **10/10 under the same strong reporting bias and 50% mention rate**, and does so across a range of frequency weights including zero. Joint structure alone suffices when it is non-degenerate, which is the situation real corpora present: colour and object co-occurrence there is far richer than a cyclic preference. The pattern to carry forward is procedural. Three times in two days a result looked like a property of the problem and turned out to be a property of the solver: descent retention, the triangle-energy pass, and this ceiling. Each was settled by computing the bound rather than inferring it. ## Next 1. Push field validity: segmentation is the dominant residual noise (70% exact group counts); texture-world slot binding (running) or a supervised-free segmentation refinement would move validity directly. 2. Third-order triangle/holonomy invariants -- the remaining queued invention: counterfeits to date satisfy pairwise statistics only, and triangle tensors multiply the constraint count by N per node. 3. The objective-battery finding stands: no standard SSL recipe meets the ceiling spec; binding fails on flat scenes at both resolutions (sorted alpha masses 0.23--0.18 against uniform 0.20), and the texture world tests whether within-object structure fixes it. ## Main artifacts - `artifacts/synth_v0/manifest.json`, `scenes.private.json`, `captions.json`, `interventions.private.json` - `artifacts/synth_v0/{vision,text}_tower.pt` - `artifacts/synth_v0/{vision,text,text_orbits}.pt` - `artifacts/synth_v0/diagnostics.json` - `artifacts/synth_v0/gate_test512_orbit.json` - `artifacts/synth_v0/content_gate_k32.json`