summaryrefslogtreecommitdiff
path: root/LAB_NOTES.md
diff options
context:
space:
mode:
Diffstat (limited to 'LAB_NOTES.md')
-rw-r--r--LAB_NOTES.md575
1 files changed, 575 insertions, 0 deletions
diff --git a/LAB_NOTES.md b/LAB_NOTES.md
new file mode 100644
index 0000000..d8d235f
--- /dev/null
+++ b/LAB_NOTES.md
@@ -0,0 +1,575 @@
+# WorldAlign research log
+
+This file records research decisions, failed hypotheses, and the current
+experimental gate. Exact metrics live in the result reports and machine-readable
+artifacts. The collaborator-facing statement of the project is maintained in
+`PROJECT_CONCEPT.md`.
+
+## Research target
+
+The target is a real multimodal large-model system, not a geometry result or a
+toy correspondence benchmark.
+
+The working hypothesis is that independently trained, sufficiently informative
+unimodal representations may recover the same world structure up to a change
+of coordinates. Multimodal alignment could then be obtained by descending a
+high-level shared-world energy, without paired contrastive supervision.
+
+Important constraints:
+
+- Do not treat the word "graph" as a commitment to explicit semantic nodes.
+ It is only an analogy for richly constrained relations in learned states.
+- Node-wise or sample-wise sufficiency is not enough. The observations must be
+ overcomplete enough to constrain relations among world states.
+- Generic encyclopedia text is not a closed counterpart of the visible world:
+ it contains many latent factors that images cannot identify.
+- The target method is not an image-to-language MLP bridge. The bridge baseline
+ is useful only as an interface/control.
+- Paired examples may be used for held-out diagnosis, never by the unpaired
+ energy, optimizer, initialization, model selection, or stopping rule.
+
+## 2026-07-28: learned-bridge baseline
+
+Flickr30k, frozen DINOv2 and frozen Qwen were aligned with static
+GW/SWD/isometry objectives.
+
+- Cross-modal relational similarity is measurably non-random.
+- Static GW produces sharp but mostly incorrect correspondences.
+- A paired bridge verifies that the frozen-language-model interface works.
+- The unpaired learned bridge does not produce usable multimodal ability.
+
+See `RESULTS.md`.
+
+## 2026-07-29: overcomplete observations
+
+Five independent captions of the same scene were treated as an observation
+orbit. On 512 held-out examples, the correct-vs-shuffled relation-energy gap
+increased from 0.316 with one caption to 0.582 with five captions.
+
+This supports the narrower claim that relational overcompleteness makes the
+correct configuration easier to distinguish. It does not establish that the
+current energy identifies the correct absolute configuration.
+
+## 2026-07-29: energy descent without a cross-modal map
+
+Free Qwen-side latent particles were initialized from real text states and
+optimized using:
+
+- standardized relation-field matching;
+- multiscale conditional-neighborhood matching;
+- sliced text-population matching;
+- proximity to text-only prototypes.
+
+Weak manifold regularization briefly improved held-out R@10 from 2.34% to
+12.11%, but left the decodable language manifold. Strong regularization
+improved it only to 6.25% at the best intermediate step.
+
+Language-prior-only and shuffled-vision controls showed that caption-generation
+improvements were not reliably image-conditioned.
+
+## 2026-07-29: convergence audit and falsification
+
+The original logged total was not a fixed objective because relation and
+conditional weights were annealed. Therefore that curve could not establish
+convergence.
+
+Starting from the strong-manifold (`m30`) endpoint, the final weights were
+frozen and optimization was continued at one tenth of the learning rate:
+
+- fixed energy decreased from 2.767 to approximately 1.54;
+- R@10 decreased from 5.86% to approximately 3--4%;
+- median rank worsened from 103 to approximately 120--123;
+- random R@10 for 256 candidates is 3.91%;
+- gradients remained nonzero and Adam continued to oscillate.
+
+Most decisively, the true paired language configuration has fixed energy about
+4.49, while an incorrect synthetic configuration reaches about 1.54.
+
+Therefore the present energy is falsified as an alignment objective:
+
+\[
+E(\text{correct paired state}) >
+E(\text{incorrect synthetic state}).
+\]
+
+The early retrieval gain was a trajectory transient, not evidence that further
+energy descent would recover the correct gauge. More steps or optimizer tuning
+cannot repair this ordering.
+
+## 2026-07-29: on-manifold assignment gate
+
+Decision, confirmed with the project owner: restrict the configuration space
+to permutations of real frozen text states. On this space the language-only
+terms of the falsified energy are permutation-invariant constants, and the
+continuous counterfeit route is closed by construction. Per the standing
+gate, only ranking and local-ordering diagnostics were run; no new descent.
+
+Result: the falsification survives the restriction.
+
+- Global ordering always passes: the true assignment ranks below all 1,000
+ random permutations in every condition (z 2.7--18.7).
+- Local ordering always fails: 6.8--33% of transpositions of the true
+ assignment lower the energy (exact closed-form enumeration); at N=5,000
+ the count is 2.70M of 12.50M. Scale does not repair it.
+- Exact steepest 2-swap descent from the truth keeps 2.7--13.9% of nodes;
+ random-start quenches reach energies 0.9--1.9 below the true
+ configuration at 0.0--0.4% accuracy. On-manifold counterfeits exist in
+ every condition.
+
+Mechanism, from three audits: pooled states are extremely anisotropic
+(99.5% of Flickr caption pairs above cosine 0.8); the cross-modal share of
+standardized relational structure is Spearman 0.19--0.31; that share is
+spectrally concentrated (top-16 population directions: 0.367; full space:
+0.306; whitening collapses it to 0.03). Five-caption orbits and visible
+relations denoise within this coarse subspace; distribution-valued edge
+channels (std/q10/q90 of view-pair cosines) carry no shared signal and
+dilute the energy; model scale was already known to be inert. A roughly
+sixteen-dimensional topical code cannot separate 512 instances, so no
+energy over these node states can identify the assignment.
+
+See `MANIFOLD_RESULTS.md`.
+
+## 2026-07-29: view-level shared structure inside VG nodes
+
+The per-node view permutation hidden at preparation time is exactly
+replayable from the seed and the private image IDs; the replay reconstructs
+all 5,000 released phrase bundles verbatim and is used only for evaluation.
+
+- Within-node relation fields align across modalities: matched Spearman
+ 0.128 against -0.002 shuffled, z = 46.6 over 5,000 nodes.
+- With the true node-level alignment as a coarse frame, population-context
+ profiles match the 16 views at 14.5% accuracy against 6.25% chance
+ (top-3 29.5%); a larger frame helps (200 nodes: 12.9%). This is an upper
+ bound: a blind system must substitute its own recovered coarse frame.
+- Within-node relations alone stay near chance (7.0%), so the population
+ context carries the signal.
+
+The fine-grained shared structure missing at node level exists one level
+down, in the compositional structure that pooling discards.
+
+## 2026-07-29: Ricci-flow control and static view-level gate
+
+Two same-day continuations, both requested before the structural redesign.
+
+Ollivier-Ricci flow on each modality's kNN geometry (independent,
+identical hyperparameters, diffusion-only baseline for attribution) more
+than doubles the global ranking z (Flickr 21.4 to 47.8; VG 6.9 to 22.5)
+and changes no verdict: 9--14% improving transpositions, descent retention
+1--5%, counterfeits at about half the true energy in all four conditions.
+Geometric smoothing is a frame ingredient, not a resolver; it cannot be
+part of the alignment transform itself (information loss), matching the
+whitening result.
+
+The static view-level gate fails at scale: with the true coarse frame,
+none of 5,000 nodes has the true view assignment as a strict local
+minimum (32.0% improving for the within-node relational energy, 28.3% for
+the profile energy, 25.5% joint). Frame-free internal-field matching is
+null at 512 nodes over three seeds (R@10 at chance, zero margin-decile
+precision, true-pair z 0.1--1.8 after a capacity null; the negative z
+seen at 96-node smoke scale did not replicate). Consequence: the view
+level needs model-internal relations or true re-encoding interventions --
+the R3/R8 battery in `STRUCTURE_DESIGN.md` is the critical path.
+
+## 2026-07-29: representation batteries and the first passed basin audit
+
+The R6 and R3 batteries of `STRUCTURE_DESIGN.md` were run the same day.
+
+R6, within-scene noise whitening: directions maximizing between-scene
+over within-scene variance, fitted per modality on unpaired orbits
+(train-only for Flickr, holdout-only for VG). Cross-modal relational
+Spearman rises 0.298 to 0.358 (Flickr orbit), 0.189 to 0.326 (single
+caption, now above the raw orbit mean), 0.213 to 0.614 (VG both sides,
+k=8); improving-swap fractions collapse 7.8% to 0.5% and 23.5% to 0.25%.
+Population whitening had destroyed the same signal; the only difference
+is whitening the within-scene rather than the total covariance.
+
+R3, model-layer readout: joint-context encoding on all 5,000 VG nodes.
+Attention weights are modality-specialized (best attention-attention
+matched rho 0.076); in-context states carry rho 0.216 against the 0.128
+isolated baseline (z 69.6 vs 46.6). Relations do live in the model layer,
+and the shared readout is the mixed states, not the mixing weights.
+
+Full gate on the new states: nine of nine VG conditions pass the basin
+audit -- ten random-restart quenches per run all end 5--19% above the
+true energy, replicated over subset seeds 0--2, for both isolated
+projected states and the in-context composite (which roughly doubles the
+margin at matched k). First energies in the project whose descent finds
+nothing below the truth. Flickr still fails by under 1%: one visual view
+means no visual orbit to whiten, so two-sided multi-view observation is a
+data-protocol requirement. The true assignment is still not a strict
+local minimum (0.17--0.28% improving swaps; descent keeps 52--59%), so
+what is licensed is soft-coupling recovery, not hard matching.
+
+See `BATTERY_RESULTS.md`.
+
+## 2026-07-29: blind recovery enters the hard phase
+
+First Phase-2 experiment on the gate-passing states, hidden-shuffle
+protocol, three search arms. Parallel tempering (eight replicas, exact
+closed-form deltas, 60k rounds) equilibrates into the quench band of the
+gate report -- ending 9--20% above the true energy at 1--3% accuracy --
+and never enters the true funnel; entropic Sinkhorn and spectral-band
+homotopy stall at +62--100% with zero verified seed mass. Discrete
+annealing beats continuous relaxation everywhere and still loses to the
+landscape.
+
+Diagnosis: the classic hard phase of planted assignment -- the gate shows
+the configuration is information-theoretically identifiable while local
+dynamics equilibrate into an exponential shelf of wrong states. The gap
+between the statistical and algorithmic thresholds is now the central
+quantity.
+
+The N-scaling probe closes the cheapest hypothesis: at matched per-node
+move budget (N=2,048, 320k rounds), the reachable shelf returns to the
+same +9--10% band as N=512 and the small accuracy signal vanishes.
+Density alone does not approach an algorithmic threshold in the tested
+range.
+
+The first C5 anchor wave closes the second-cheapest. Pixel- and
+lexicon-derived unary anchors carry population signal (combined z = 17.2)
+with no per-pair reliability (top-margin precision 2%, mutual-NN 0.17%),
+and dense injection into the tempering energy degrades recovery
+monotonically in the weight (accuracy 3.1% to 1.0--2.0%, relational
+energy +9% to +30--115%). Scene-level anchors are coarse-redundant with
+the shared subspace; usable side information must be per-pair reliable or
+independent of the relational field. Fine-grained anchors (numerals,
+spatial predicates, box geometry) live at view level with sixteen
+candidates, so the second wave merges into the view recursion. Remaining
+search-side lever: cavity/message-passing dynamics.
+
+The second wave (2026-07-30) delivers the first seeds. Per-view color,
+light, size, and position anchors against crop pixels and box geometry:
+frame-free view matching 11.5% vs 6.25% chance, 35.6% precision on
+top-1%-margin views (354 seeds; every prior high-confidence channel was
+0--4%). Aggregated per-pair Hungarian values give the best frame-free
+node affinity to date (R@10 4.4--4.9x chance over three seeds), with
+node-level seed precision still at 2--6%. Pixel and lexicon statistics
+are exhausted; the next package is a factor graph over node and view
+assignments -- anchor unaries, relational pairwise, message passing --
+plus the untested numeral-to-repetition anchor.
+
+See `RECOVERY_RESULTS.md`.
+
+## 2026-07-30: synthetic closed world v0 -- the objective becomes the bottleneck
+
+A procedural world with exact closure (43-word vocabulary, unique
+referents, ring-placement multiplicity, relation-constrained rendering,
+ground-truth interventions) and from-scratch unimodal towers on disjoint
+splits. First measurement inverts the natural-data pattern: linear CKA
+matches (0.278 vs 0.262) while relational Spearman collapses six-fold
+(0.030 vs 0.184, z 2.4 vs 12.3). The gate fails accordingly; content
+projection restores ranking (z = 117) but not the basin (counterfeit at
+-13%).
+
+Attribution is clean because data sufficiency is total by construction:
+InfoNCE's uniformity pressure turned the vision tower into a
+near-orthogonal instance codebook, erasing between-scene geometry.
+Sufficiency of the data does not transfer to sufficiency of the
+representation; the unimodal objective decides which world geometry
+survives, and R1 is a property of the data-objective pair. The SSL
+objective is now a phase-diagram axis. Next: reconstruction-family
+towers (doubling as the masked predictors of the prediction-transfer
+energy), then the structure identification battery.
+
+See `SYNTH_RESULTS.md` and the 2026-07-30 additions to
+`STRUCTURE_DESIGN.md` (unified element-relation object; the three-rung
+energy ladder ending in lifted unimodal world models).
+
+## 2026-07-30/31: objective battery, set machinery, and the first phase coordinates
+
+The synthetic world's second day closes three questions and opens one.
+
+Objective battery, complete: five standard self-supervised recipes
+(orbit-InfoNCE, SimMIM, their hybrid, data2vec-lite, slot attention at
+two resolutions) all fail to produce factor-complete global states, each
+by a distinct mechanism -- shortcut selection, local-only knowledge,
+non-merging losses, softened shortcut, and spatial tiling instead of
+object binding. Slot attention carries the most relational sharing
+(0.128 against 0.030 for InfoNCE) and the set-versus-pool doctrine is
+confirmed quantitatively: pooling slot sets without correspondence drops
+color decodability from 0.99 to 0.55.
+
+Set-kernel machinery, taken to a hand-crafted ceiling
+(connected-component descriptors against phrase bag-of-words): seven
+measured iterations moved the gate from z = -0.1 and 45% improving swaps
+to z = 51, 0.15% improving swaps, and 77% descent retention. Priced en
+route: mass-weighting inside matching destroys grading; multi-view field
+averaging cancels segmentation noise; a factor present on one side only
+(shape) is a counterfeit lever; so is size-dependent matching bias.
+
+End to end, honestly: blind tempering on the best fields still finds
+assignments 4--9% below the truth at chance accuracy, and size
+residualization removes shared signal along with the nuisance. With VG's
+passing configuration at field validity 0.61 and this ceiling near 0.45,
+the basin-pass boundary at N = 512 is bracketed between them: the first
+coordinates of the identifiability phase diagram.
+
+See `SYNTH_RESULTS.md`.
+
+## 2026-07-31: the Fock lift and the (validity, N) boundary
+
+The out-of-the-box program's first invention went to flight the same
+morning. A symmetric-tensor moment kernel on object sets (the truncated
+second-quantized lift: composition as algebra, multiplicity as an
+observable, no matching step and hence none of the measured matching
+biases) with factor-mirrored one-hot descriptors gives the synthetic
+world its first full basin pass at N = 64 -- no counterfeit, 92%
+retention. At N = 512 counterfeits return 18--22% below the truth and
+recovery stays at chance; residual segmentation error (30% wrong counts)
+and shape-cluster impurity suffice at eight times the configuration
+space. Identifiability is now measured as a joint boundary in field
+validity and population size: pass at (0.45, 64), fail at (0.45, 512),
+pass at (0.61, 512). Queued inventions, in order: orbit-derived
+per-entry precision weighting (channel-matched energy), third-order
+triangle invariants (holonomy against pairwise counterfeits), moment
+ideals as gauge-free content coordinates; mixed-curvature products noted
+as a symmetry-breaking device and deprioritized.
+
+See `SYNTH_RESULTS.md` and the tool inventory in `STRUCTURE_DESIGN.md`.
+
+## 2026-07-31: the gate statistic was measuring the searcher
+
+A methodological correction that supersedes part of the same day's
+optimism. Third-order (triangle) invariants were added and, in a harness
+whose descent proposes 64 sampled pairs per step, the truth became
+essentially immovable at N=512 (retention 99.6%, no counterfeit, z=50),
+replicated over three hidden shuffles. Two controls dismantled the
+reading: the pairwise-only energy passes identically in that harness
+(retention 100%), and long tempering on the triangle energy at N=128
+reaches 10% below the truth at zero accuracy.
+
+Conclusion: descent retention measures the strength of the search
+operator, not the identifiability of the energy. Exact all-pairs
+enumeration (yesterday's harness) is a strictly stronger searcher than
+sampled proposals, and long tempering is stronger still. The gate
+statistic is therefore redefined:
+
+\[
+E(\text{truth}) \le E(\text{deepest state a strong searcher reaches}),
+\]
+
+measured with long parallel tempering from random starts plus a
+truth-initialized arm, uniformly across candidate energies. Retention and
+improving-swap fractions remain as diagnostics, never as verdicts. All
+candidate energies are being re-measured on this basis
+(`synth_deep_gate.py`); readings from the weaker harness are marked
+superseded in `SYNTH_RESULTS.md`.
+
+## 2026-07-31: closed form, spectral solvers, and a clean Tier 0
+
+Three results and one protocol decision, all same day.
+
+Closed form: the triangle energy equals a constant minus 2 trace(M^3)/6
+with M the elementwise product of permuted text and visual fields, so
+one batched matrix product scores hundreds of permutations over all
+C(N,3) triples. Validated to 3e-7 against brute force; opens trace(M^k)
+for any k. Iteration time per configuration falls from hours to minutes.
+
+Spectral solvers (Umeyama, GRAMPA) were added because they are
+polynomial and dimension-free -- relation fields are N x N whatever the
+embedding widths -- and therefore bypass the glassy landscape. They
+return 3--13x chance. Correlated-graph-matching theory explains it: at
+N=256 the information-theoretic threshold sits near rho = 0.29 and the
+polynomial threshold near rho = 0.9, and our fields are at 0.51--0.57.
+The project is squarely in the computationally hard phase, which now has
+a quantitative account rather than an empirical one.
+
+Identifiability is not correlation. The colour-only field reaches rho =
+0.863 yet caps at 49% because only 126 of 256 scenes have distinct
+colour multisets; its eigenvalue degeneracy is that collision structure.
+Full descriptors are 100% distinct. Channel mixing inside one moment
+kernel dilutes 0.86 to 0.51.
+
+Protocol, decided with the user: anchors that declare "red means this
+hue" are not labels but are hand-supplied cross-modal prior, so results
+must be reported by tier. Tier 0 forbids any declared correspondence.
+Achieved the same day on world v1 (Zipf-skewed factor marginals, since
+uniform v0 makes the correspondence information-theoretically
+unrecoverable): text factor families discovered by mutual exclusivity,
+vision colour classes by k-means on raw RGB, the two paired by marginal
+frequency rank -- 10/10 entries semantically correct, fields at rho =
+0.570 with a 97.3% identifiability ceiling.
+
+See `SYNTH_RESULTS.md`.
+
+## 2026-07-31: first recovery, and it was segmentation all along
+
+The rho decomposition settled the target: oracle segmentation gives field
+correlation exactly 1.0 and the Tier 0 derived dictionary is free (oracle
+and derived fields are identical), while connected-component segmentation
+gives 0.5696. Three unimodal fixes closed most of it -- Gestalt
+appearance grouping (exact group count 71.9 to 90.2%),
+distance-transform watershed for touching ring members (exact object
+count 74.2 to 100%), and one-dimensional k-means on radial extent for
+size classes (52.7 to 87.9%, because area confounds size with shape and
+the classes are gap-separated rather than equally populated).
+
+Field correlation went 0.570 to 0.671 to 0.928, and GRAMPA recovery
+followed it 5.1% to 9.4% to **53.9%** at N=256 (0.39% chance), replicated
+at 53.9/53.9/54.7% over three hidden shuffles, and 26.3% at N=600 (0.17%
+chance). First end-to-end world matching in the project: no pairs, no
+declared dictionary, no learned cross-modal map -- factor families from
+mutual exclusivity, colour correspondence from marginal frequency rank,
+object states from pixels, assignment from a polynomial spectral solver
+on two independently built relation fields.
+
+Composing the two solvers closes most of the remaining gap: GRAMPA
+initialisation 53.5%, plus pairwise-energy steepest descent 86.7%, plus
+pairwise-and-triangle refinement **95.3%**. Spectral never gets trapped
+but rounds badly; the energy rounds well but cannot find the basin; the
+composition does both. The final state sits 0.04% below the truth in
+energy at 95.3% accuracy -- residual field noise, not a counterfeit.
+
+The theoretical account held throughout: nothing worked below rho = 0.9
+and recovery appeared as soon as the fields crossed it. Reporting-bias
+worlds show the derivation degrades gracefully -- text-versus-pixel rank
+correlation 1.00/0.98/0.84/0.77 gives dictionaries of 8/8/6/6 out of ten
+-- with errors concentrated in the close-frequency tail, so
+margin-calibrated entries are the natural seed set.
+
+The recovered assignment then transfers: used as pseudo-pairs for a ridge
+map and applied to 200 held-out scenes that took no part in matching,
+retrieval is 93.0% exact against 0.5% chance, colour set F1 0.989, count
+multiset 97.5%. The random-pair control sits exactly at chance and the
+shuffled-image control drops colour F1 to 0.295, so the output is
+image-conditioned and extrapolates. Rungs four and five of the evidence
+ladder are closed inside the synthetic world; matching is a bootstrap for
+the map, not the deployment mechanism.
+
+Unchanged by this: object states are hand-built upper-bound descriptors
+that no tested self-supervised recipe reproduces, and the world is
+procedurally generated.
+
+See `SYNTH_RESULTS.md`.
+
+## 2026-08-01: reporting bias, and three solver-versus-information calls
+
+Worlds v3 and v4 test whether the derived dictionary survives reporting
+bias. v3 couples colour to shape cyclically and biases mentions toward
+rare colours, so captions name half the groups and the pixel and text
+colour rankings diverge in the tail; v4 replaces the cyclic coupling with
+a distinct Dirichlet shape profile per colour.
+
+v3 results: marginal rank 4/10, joint structure 6/10, joint plus marginal
+under exact enumeration of the 720 shape permutations 7/10, oracle shape
+columns with measured clusters 8/10, oracle labels 10/10. Vision-side
+extraction is not the limit (colour cluster purity 1.000, shape 0.998).
+The 6/10 initially looked like an information bound -- ten colours over
+six shapes leaves exactly four twin profiles, and exactly four appeared
+-- but the twins separate by frequency, and the combined evidence has
+minimum pairwise distance 0.2322 against 0.0008 for shape alone. The
+bound is 10/10; what is left is one entry of solver gap and two of
+estimation noise from the thinned joint table.
+
+v4 results: exact enumeration recovers 10/10 under the same bias and
+mention rate, across frequency weights including zero. Non-degenerate
+joint structure suffices on its own, which is the regime real corpora
+occupy.
+
+Three times now a limit has looked intrinsic and proved to be the
+solver's: descent retention, the triangle-energy pass, and this ceiling.
+The rule stands -- compute the bound, do not infer it.
+
+See `SYNTH_RESULTS.md`.
+
+## 2026-08-01: factor families on real language
+
+First step of the port to natural data, text side only, on 80,000 Visual
+Genome region descriptions with no lexicon and no labels.
+
+The synthetic pipeline found families by mutual exclusivity: template
+phrases carry exactly one colour word, so colour terms never co-occur.
+That does not transfer. In real descriptions colour words are
+statistically independent or better -- observed over expected
+co-occurrence is 1.29 for black with white, 0.93 for red with blue --
+while head nouns are the genuinely exclusive family (0.14 for car with
+truck). Run as written, the method recovers clean noun classes and no
+attribute families at all. Mutual exclusivity is a template artifact.
+
+Standard distributional word-class induction, of which exclusivity is a
+degenerate case, does transfer. Positive PMI context vectors, truncated
+SVD, k-means over the 300 most frequent words yields a colour family at
+0.67 purity (red, brown, gray, dark, grey, pink, tan, purple) plus
+coherent object categories: vehicles, animals, clothing, people, water
+scenes. Colour terms split into three clusters rather than one, and the
+split is structural rather than noisy -- white and blue cluster with sky
+and clouds, because what a colour word modifies shapes its distribution.
+That dependence is the signal cross-modal value alignment feeds on.
+
+Next: the vision side under unsupervised segmentation, then the field
+correlation as the go/no-go before any recovery run.
+
+## 2026-08-01: first natural-data coordinate, and a correction
+
+The Tier 0 recipe was ported to Visual Genome: DINOv2 patch features,
+spectral segmentation of the affinity graph, segment feature classes on
+one side; distributional word classes on the other. Field correlation at
+the hidden pairing is **0.166**, far under the 0.9 that recovery needs,
+so by our own go/no-go no recovery run is warranted yet.
+
+A correction came with it. Aligning vision classes to text classes by
+frequency rank and by an oracle give bit-identical fields -- 0.1662 both
+ways -- while agreeing on 0 of 23 classes. Moment fields are built from
+within-modality inner products, and permuting feature coordinates
+consistently leaves every inner product unchanged, so relation fields are
+gauge-invariant in the class labels and the dictionary never enters them.
+The same was true unnoticed in the synthetic world, where oracle and
+derived dictionaries produced identical fields. The dictionary is
+required for encoding a new image into text coordinates, not for
+matching; matching has always run on the fields alone.
+
+The 0.166 therefore reads as representation quality, and localises the
+deficit at vision-side granularity: six segments per image, background
+included, against sixteen phrases about specific objects. A sixteen-
+segment run is measuring whether granularity is the binding constraint.
+
+## 2026-08-01: natural data reaches 0.656, and segmentation was not the wall
+
+Full port to Visual Genome, measured by the go/no-go statistic. Three
+results overturned expectations, each checked rather than inferred.
+
+Discretisation was the dominant loss: hard class codes give 0.166,
+continuous states 0.489. Relation fields need scene similarity only, so
+quantising states first is pure attrition; the synthetic world hid this
+because its states are discrete by construction.
+
+Segmentation is not the bottleneck on photographs. Oracle region boxes
+buy 0.061 over unsupervised spectral segmentation of DINOv2 patches
+(0.550 against 0.489), the inverse of the synthetic world where
+segmentation was the whole gap. Flat untextured scenes are adversarial
+for object discovery in a way photographs are not.
+
+Content projection is again the largest lever: 0.489 to **0.656**,
+matching its earlier role at node level (0.21 to 0.61). Protocol-legal
+and oracle-box configurations land within 0.007, so the residual is not
+about object localisation.
+
+Control: frozen general-purpose encoders underperform corpus-fitted
+statistics (0.19 for DINOv2 crop CLS against Qwen phrase states), so
+in-context features and corpus-specific distributional vectors are the
+better substrate.
+
+Verdict: 0.656 against the 0.9 recovery needs; no recovery run. The
+untried lever is observation orbits, which VG lacks by construction --
+one photograph per scene. Multi-view natural data is the next protocol
+requirement.
+
+See `NATURAL_RESULTS.md`.
+
+## Current experimental gate
+
+The node-level ordering-and-basin gate is passed on VG by content-projected
+states (isolated and in-context composite). The standing rules extend
+rather than retire:
+
+1. Blind recovery next: unpaired soft-coupling optimization (Sinkhorn over
+ real states, preserving the decodable path) must place calibrated
+ coupling mass on the hidden truth from scratch. Its score is coupling
+ mass at truth and verified seed precision, never retrieval means alone.
+2. The view-level gate stands unchanged for the coarse-to-fine recursion:
+ static view fields failed it; in-context view states plus projection
+ are the next candidate, with replayed view truth evaluation-only.
+3. Any new energy or state family still passes ordering before descent,
+ including single-transposition local ordering.
+4. Flickr-style single-view data cannot host the current pass; two-sided
+ multi-view observation is a protocol requirement for now.