diff options
Diffstat (limited to 'LAB_NOTES.md')
| -rw-r--r-- | LAB_NOTES.md | 575 |
1 files changed, 575 insertions, 0 deletions
diff --git a/LAB_NOTES.md b/LAB_NOTES.md new file mode 100644 index 0000000..d8d235f --- /dev/null +++ b/LAB_NOTES.md @@ -0,0 +1,575 @@ +# WorldAlign research log + +This file records research decisions, failed hypotheses, and the current +experimental gate. Exact metrics live in the result reports and machine-readable +artifacts. The collaborator-facing statement of the project is maintained in +`PROJECT_CONCEPT.md`. + +## Research target + +The target is a real multimodal large-model system, not a geometry result or a +toy correspondence benchmark. + +The working hypothesis is that independently trained, sufficiently informative +unimodal representations may recover the same world structure up to a change +of coordinates. Multimodal alignment could then be obtained by descending a +high-level shared-world energy, without paired contrastive supervision. + +Important constraints: + +- Do not treat the word "graph" as a commitment to explicit semantic nodes. + It is only an analogy for richly constrained relations in learned states. +- Node-wise or sample-wise sufficiency is not enough. The observations must be + overcomplete enough to constrain relations among world states. +- Generic encyclopedia text is not a closed counterpart of the visible world: + it contains many latent factors that images cannot identify. +- The target method is not an image-to-language MLP bridge. The bridge baseline + is useful only as an interface/control. +- Paired examples may be used for held-out diagnosis, never by the unpaired + energy, optimizer, initialization, model selection, or stopping rule. + +## 2026-07-28: learned-bridge baseline + +Flickr30k, frozen DINOv2 and frozen Qwen were aligned with static +GW/SWD/isometry objectives. + +- Cross-modal relational similarity is measurably non-random. +- Static GW produces sharp but mostly incorrect correspondences. +- A paired bridge verifies that the frozen-language-model interface works. +- The unpaired learned bridge does not produce usable multimodal ability. + +See `RESULTS.md`. + +## 2026-07-29: overcomplete observations + +Five independent captions of the same scene were treated as an observation +orbit. On 512 held-out examples, the correct-vs-shuffled relation-energy gap +increased from 0.316 with one caption to 0.582 with five captions. + +This supports the narrower claim that relational overcompleteness makes the +correct configuration easier to distinguish. It does not establish that the +current energy identifies the correct absolute configuration. + +## 2026-07-29: energy descent without a cross-modal map + +Free Qwen-side latent particles were initialized from real text states and +optimized using: + +- standardized relation-field matching; +- multiscale conditional-neighborhood matching; +- sliced text-population matching; +- proximity to text-only prototypes. + +Weak manifold regularization briefly improved held-out R@10 from 2.34% to +12.11%, but left the decodable language manifold. Strong regularization +improved it only to 6.25% at the best intermediate step. + +Language-prior-only and shuffled-vision controls showed that caption-generation +improvements were not reliably image-conditioned. + +## 2026-07-29: convergence audit and falsification + +The original logged total was not a fixed objective because relation and +conditional weights were annealed. Therefore that curve could not establish +convergence. + +Starting from the strong-manifold (`m30`) endpoint, the final weights were +frozen and optimization was continued at one tenth of the learning rate: + +- fixed energy decreased from 2.767 to approximately 1.54; +- R@10 decreased from 5.86% to approximately 3--4%; +- median rank worsened from 103 to approximately 120--123; +- random R@10 for 256 candidates is 3.91%; +- gradients remained nonzero and Adam continued to oscillate. + +Most decisively, the true paired language configuration has fixed energy about +4.49, while an incorrect synthetic configuration reaches about 1.54. + +Therefore the present energy is falsified as an alignment objective: + +\[ +E(\text{correct paired state}) > +E(\text{incorrect synthetic state}). +\] + +The early retrieval gain was a trajectory transient, not evidence that further +energy descent would recover the correct gauge. More steps or optimizer tuning +cannot repair this ordering. + +## 2026-07-29: on-manifold assignment gate + +Decision, confirmed with the project owner: restrict the configuration space +to permutations of real frozen text states. On this space the language-only +terms of the falsified energy are permutation-invariant constants, and the +continuous counterfeit route is closed by construction. Per the standing +gate, only ranking and local-ordering diagnostics were run; no new descent. + +Result: the falsification survives the restriction. + +- Global ordering always passes: the true assignment ranks below all 1,000 + random permutations in every condition (z 2.7--18.7). +- Local ordering always fails: 6.8--33% of transpositions of the true + assignment lower the energy (exact closed-form enumeration); at N=5,000 + the count is 2.70M of 12.50M. Scale does not repair it. +- Exact steepest 2-swap descent from the truth keeps 2.7--13.9% of nodes; + random-start quenches reach energies 0.9--1.9 below the true + configuration at 0.0--0.4% accuracy. On-manifold counterfeits exist in + every condition. + +Mechanism, from three audits: pooled states are extremely anisotropic +(99.5% of Flickr caption pairs above cosine 0.8); the cross-modal share of +standardized relational structure is Spearman 0.19--0.31; that share is +spectrally concentrated (top-16 population directions: 0.367; full space: +0.306; whitening collapses it to 0.03). Five-caption orbits and visible +relations denoise within this coarse subspace; distribution-valued edge +channels (std/q10/q90 of view-pair cosines) carry no shared signal and +dilute the energy; model scale was already known to be inert. A roughly +sixteen-dimensional topical code cannot separate 512 instances, so no +energy over these node states can identify the assignment. + +See `MANIFOLD_RESULTS.md`. + +## 2026-07-29: view-level shared structure inside VG nodes + +The per-node view permutation hidden at preparation time is exactly +replayable from the seed and the private image IDs; the replay reconstructs +all 5,000 released phrase bundles verbatim and is used only for evaluation. + +- Within-node relation fields align across modalities: matched Spearman + 0.128 against -0.002 shuffled, z = 46.6 over 5,000 nodes. +- With the true node-level alignment as a coarse frame, population-context + profiles match the 16 views at 14.5% accuracy against 6.25% chance + (top-3 29.5%); a larger frame helps (200 nodes: 12.9%). This is an upper + bound: a blind system must substitute its own recovered coarse frame. +- Within-node relations alone stay near chance (7.0%), so the population + context carries the signal. + +The fine-grained shared structure missing at node level exists one level +down, in the compositional structure that pooling discards. + +## 2026-07-29: Ricci-flow control and static view-level gate + +Two same-day continuations, both requested before the structural redesign. + +Ollivier-Ricci flow on each modality's kNN geometry (independent, +identical hyperparameters, diffusion-only baseline for attribution) more +than doubles the global ranking z (Flickr 21.4 to 47.8; VG 6.9 to 22.5) +and changes no verdict: 9--14% improving transpositions, descent retention +1--5%, counterfeits at about half the true energy in all four conditions. +Geometric smoothing is a frame ingredient, not a resolver; it cannot be +part of the alignment transform itself (information loss), matching the +whitening result. + +The static view-level gate fails at scale: with the true coarse frame, +none of 5,000 nodes has the true view assignment as a strict local +minimum (32.0% improving for the within-node relational energy, 28.3% for +the profile energy, 25.5% joint). Frame-free internal-field matching is +null at 512 nodes over three seeds (R@10 at chance, zero margin-decile +precision, true-pair z 0.1--1.8 after a capacity null; the negative z +seen at 96-node smoke scale did not replicate). Consequence: the view +level needs model-internal relations or true re-encoding interventions -- +the R3/R8 battery in `STRUCTURE_DESIGN.md` is the critical path. + +## 2026-07-29: representation batteries and the first passed basin audit + +The R6 and R3 batteries of `STRUCTURE_DESIGN.md` were run the same day. + +R6, within-scene noise whitening: directions maximizing between-scene +over within-scene variance, fitted per modality on unpaired orbits +(train-only for Flickr, holdout-only for VG). Cross-modal relational +Spearman rises 0.298 to 0.358 (Flickr orbit), 0.189 to 0.326 (single +caption, now above the raw orbit mean), 0.213 to 0.614 (VG both sides, +k=8); improving-swap fractions collapse 7.8% to 0.5% and 23.5% to 0.25%. +Population whitening had destroyed the same signal; the only difference +is whitening the within-scene rather than the total covariance. + +R3, model-layer readout: joint-context encoding on all 5,000 VG nodes. +Attention weights are modality-specialized (best attention-attention +matched rho 0.076); in-context states carry rho 0.216 against the 0.128 +isolated baseline (z 69.6 vs 46.6). Relations do live in the model layer, +and the shared readout is the mixed states, not the mixing weights. + +Full gate on the new states: nine of nine VG conditions pass the basin +audit -- ten random-restart quenches per run all end 5--19% above the +true energy, replicated over subset seeds 0--2, for both isolated +projected states and the in-context composite (which roughly doubles the +margin at matched k). First energies in the project whose descent finds +nothing below the truth. Flickr still fails by under 1%: one visual view +means no visual orbit to whiten, so two-sided multi-view observation is a +data-protocol requirement. The true assignment is still not a strict +local minimum (0.17--0.28% improving swaps; descent keeps 52--59%), so +what is licensed is soft-coupling recovery, not hard matching. + +See `BATTERY_RESULTS.md`. + +## 2026-07-29: blind recovery enters the hard phase + +First Phase-2 experiment on the gate-passing states, hidden-shuffle +protocol, three search arms. Parallel tempering (eight replicas, exact +closed-form deltas, 60k rounds) equilibrates into the quench band of the +gate report -- ending 9--20% above the true energy at 1--3% accuracy -- +and never enters the true funnel; entropic Sinkhorn and spectral-band +homotopy stall at +62--100% with zero verified seed mass. Discrete +annealing beats continuous relaxation everywhere and still loses to the +landscape. + +Diagnosis: the classic hard phase of planted assignment -- the gate shows +the configuration is information-theoretically identifiable while local +dynamics equilibrate into an exponential shelf of wrong states. The gap +between the statistical and algorithmic thresholds is now the central +quantity. + +The N-scaling probe closes the cheapest hypothesis: at matched per-node +move budget (N=2,048, 320k rounds), the reachable shelf returns to the +same +9--10% band as N=512 and the small accuracy signal vanishes. +Density alone does not approach an algorithmic threshold in the tested +range. + +The first C5 anchor wave closes the second-cheapest. Pixel- and +lexicon-derived unary anchors carry population signal (combined z = 17.2) +with no per-pair reliability (top-margin precision 2%, mutual-NN 0.17%), +and dense injection into the tempering energy degrades recovery +monotonically in the weight (accuracy 3.1% to 1.0--2.0%, relational +energy +9% to +30--115%). Scene-level anchors are coarse-redundant with +the shared subspace; usable side information must be per-pair reliable or +independent of the relational field. Fine-grained anchors (numerals, +spatial predicates, box geometry) live at view level with sixteen +candidates, so the second wave merges into the view recursion. Remaining +search-side lever: cavity/message-passing dynamics. + +The second wave (2026-07-30) delivers the first seeds. Per-view color, +light, size, and position anchors against crop pixels and box geometry: +frame-free view matching 11.5% vs 6.25% chance, 35.6% precision on +top-1%-margin views (354 seeds; every prior high-confidence channel was +0--4%). Aggregated per-pair Hungarian values give the best frame-free +node affinity to date (R@10 4.4--4.9x chance over three seeds), with +node-level seed precision still at 2--6%. Pixel and lexicon statistics +are exhausted; the next package is a factor graph over node and view +assignments -- anchor unaries, relational pairwise, message passing -- +plus the untested numeral-to-repetition anchor. + +See `RECOVERY_RESULTS.md`. + +## 2026-07-30: synthetic closed world v0 -- the objective becomes the bottleneck + +A procedural world with exact closure (43-word vocabulary, unique +referents, ring-placement multiplicity, relation-constrained rendering, +ground-truth interventions) and from-scratch unimodal towers on disjoint +splits. First measurement inverts the natural-data pattern: linear CKA +matches (0.278 vs 0.262) while relational Spearman collapses six-fold +(0.030 vs 0.184, z 2.4 vs 12.3). The gate fails accordingly; content +projection restores ranking (z = 117) but not the basin (counterfeit at +-13%). + +Attribution is clean because data sufficiency is total by construction: +InfoNCE's uniformity pressure turned the vision tower into a +near-orthogonal instance codebook, erasing between-scene geometry. +Sufficiency of the data does not transfer to sufficiency of the +representation; the unimodal objective decides which world geometry +survives, and R1 is a property of the data-objective pair. The SSL +objective is now a phase-diagram axis. Next: reconstruction-family +towers (doubling as the masked predictors of the prediction-transfer +energy), then the structure identification battery. + +See `SYNTH_RESULTS.md` and the 2026-07-30 additions to +`STRUCTURE_DESIGN.md` (unified element-relation object; the three-rung +energy ladder ending in lifted unimodal world models). + +## 2026-07-30/31: objective battery, set machinery, and the first phase coordinates + +The synthetic world's second day closes three questions and opens one. + +Objective battery, complete: five standard self-supervised recipes +(orbit-InfoNCE, SimMIM, their hybrid, data2vec-lite, slot attention at +two resolutions) all fail to produce factor-complete global states, each +by a distinct mechanism -- shortcut selection, local-only knowledge, +non-merging losses, softened shortcut, and spatial tiling instead of +object binding. Slot attention carries the most relational sharing +(0.128 against 0.030 for InfoNCE) and the set-versus-pool doctrine is +confirmed quantitatively: pooling slot sets without correspondence drops +color decodability from 0.99 to 0.55. + +Set-kernel machinery, taken to a hand-crafted ceiling +(connected-component descriptors against phrase bag-of-words): seven +measured iterations moved the gate from z = -0.1 and 45% improving swaps +to z = 51, 0.15% improving swaps, and 77% descent retention. Priced en +route: mass-weighting inside matching destroys grading; multi-view field +averaging cancels segmentation noise; a factor present on one side only +(shape) is a counterfeit lever; so is size-dependent matching bias. + +End to end, honestly: blind tempering on the best fields still finds +assignments 4--9% below the truth at chance accuracy, and size +residualization removes shared signal along with the nuisance. With VG's +passing configuration at field validity 0.61 and this ceiling near 0.45, +the basin-pass boundary at N = 512 is bracketed between them: the first +coordinates of the identifiability phase diagram. + +See `SYNTH_RESULTS.md`. + +## 2026-07-31: the Fock lift and the (validity, N) boundary + +The out-of-the-box program's first invention went to flight the same +morning. A symmetric-tensor moment kernel on object sets (the truncated +second-quantized lift: composition as algebra, multiplicity as an +observable, no matching step and hence none of the measured matching +biases) with factor-mirrored one-hot descriptors gives the synthetic +world its first full basin pass at N = 64 -- no counterfeit, 92% +retention. At N = 512 counterfeits return 18--22% below the truth and +recovery stays at chance; residual segmentation error (30% wrong counts) +and shape-cluster impurity suffice at eight times the configuration +space. Identifiability is now measured as a joint boundary in field +validity and population size: pass at (0.45, 64), fail at (0.45, 512), +pass at (0.61, 512). Queued inventions, in order: orbit-derived +per-entry precision weighting (channel-matched energy), third-order +triangle invariants (holonomy against pairwise counterfeits), moment +ideals as gauge-free content coordinates; mixed-curvature products noted +as a symmetry-breaking device and deprioritized. + +See `SYNTH_RESULTS.md` and the tool inventory in `STRUCTURE_DESIGN.md`. + +## 2026-07-31: the gate statistic was measuring the searcher + +A methodological correction that supersedes part of the same day's +optimism. Third-order (triangle) invariants were added and, in a harness +whose descent proposes 64 sampled pairs per step, the truth became +essentially immovable at N=512 (retention 99.6%, no counterfeit, z=50), +replicated over three hidden shuffles. Two controls dismantled the +reading: the pairwise-only energy passes identically in that harness +(retention 100%), and long tempering on the triangle energy at N=128 +reaches 10% below the truth at zero accuracy. + +Conclusion: descent retention measures the strength of the search +operator, not the identifiability of the energy. Exact all-pairs +enumeration (yesterday's harness) is a strictly stronger searcher than +sampled proposals, and long tempering is stronger still. The gate +statistic is therefore redefined: + +\[ +E(\text{truth}) \le E(\text{deepest state a strong searcher reaches}), +\] + +measured with long parallel tempering from random starts plus a +truth-initialized arm, uniformly across candidate energies. Retention and +improving-swap fractions remain as diagnostics, never as verdicts. All +candidate energies are being re-measured on this basis +(`synth_deep_gate.py`); readings from the weaker harness are marked +superseded in `SYNTH_RESULTS.md`. + +## 2026-07-31: closed form, spectral solvers, and a clean Tier 0 + +Three results and one protocol decision, all same day. + +Closed form: the triangle energy equals a constant minus 2 trace(M^3)/6 +with M the elementwise product of permuted text and visual fields, so +one batched matrix product scores hundreds of permutations over all +C(N,3) triples. Validated to 3e-7 against brute force; opens trace(M^k) +for any k. Iteration time per configuration falls from hours to minutes. + +Spectral solvers (Umeyama, GRAMPA) were added because they are +polynomial and dimension-free -- relation fields are N x N whatever the +embedding widths -- and therefore bypass the glassy landscape. They +return 3--13x chance. Correlated-graph-matching theory explains it: at +N=256 the information-theoretic threshold sits near rho = 0.29 and the +polynomial threshold near rho = 0.9, and our fields are at 0.51--0.57. +The project is squarely in the computationally hard phase, which now has +a quantitative account rather than an empirical one. + +Identifiability is not correlation. The colour-only field reaches rho = +0.863 yet caps at 49% because only 126 of 256 scenes have distinct +colour multisets; its eigenvalue degeneracy is that collision structure. +Full descriptors are 100% distinct. Channel mixing inside one moment +kernel dilutes 0.86 to 0.51. + +Protocol, decided with the user: anchors that declare "red means this +hue" are not labels but are hand-supplied cross-modal prior, so results +must be reported by tier. Tier 0 forbids any declared correspondence. +Achieved the same day on world v1 (Zipf-skewed factor marginals, since +uniform v0 makes the correspondence information-theoretically +unrecoverable): text factor families discovered by mutual exclusivity, +vision colour classes by k-means on raw RGB, the two paired by marginal +frequency rank -- 10/10 entries semantically correct, fields at rho = +0.570 with a 97.3% identifiability ceiling. + +See `SYNTH_RESULTS.md`. + +## 2026-07-31: first recovery, and it was segmentation all along + +The rho decomposition settled the target: oracle segmentation gives field +correlation exactly 1.0 and the Tier 0 derived dictionary is free (oracle +and derived fields are identical), while connected-component segmentation +gives 0.5696. Three unimodal fixes closed most of it -- Gestalt +appearance grouping (exact group count 71.9 to 90.2%), +distance-transform watershed for touching ring members (exact object +count 74.2 to 100%), and one-dimensional k-means on radial extent for +size classes (52.7 to 87.9%, because area confounds size with shape and +the classes are gap-separated rather than equally populated). + +Field correlation went 0.570 to 0.671 to 0.928, and GRAMPA recovery +followed it 5.1% to 9.4% to **53.9%** at N=256 (0.39% chance), replicated +at 53.9/53.9/54.7% over three hidden shuffles, and 26.3% at N=600 (0.17% +chance). First end-to-end world matching in the project: no pairs, no +declared dictionary, no learned cross-modal map -- factor families from +mutual exclusivity, colour correspondence from marginal frequency rank, +object states from pixels, assignment from a polynomial spectral solver +on two independently built relation fields. + +Composing the two solvers closes most of the remaining gap: GRAMPA +initialisation 53.5%, plus pairwise-energy steepest descent 86.7%, plus +pairwise-and-triangle refinement **95.3%**. Spectral never gets trapped +but rounds badly; the energy rounds well but cannot find the basin; the +composition does both. The final state sits 0.04% below the truth in +energy at 95.3% accuracy -- residual field noise, not a counterfeit. + +The theoretical account held throughout: nothing worked below rho = 0.9 +and recovery appeared as soon as the fields crossed it. Reporting-bias +worlds show the derivation degrades gracefully -- text-versus-pixel rank +correlation 1.00/0.98/0.84/0.77 gives dictionaries of 8/8/6/6 out of ten +-- with errors concentrated in the close-frequency tail, so +margin-calibrated entries are the natural seed set. + +The recovered assignment then transfers: used as pseudo-pairs for a ridge +map and applied to 200 held-out scenes that took no part in matching, +retrieval is 93.0% exact against 0.5% chance, colour set F1 0.989, count +multiset 97.5%. The random-pair control sits exactly at chance and the +shuffled-image control drops colour F1 to 0.295, so the output is +image-conditioned and extrapolates. Rungs four and five of the evidence +ladder are closed inside the synthetic world; matching is a bootstrap for +the map, not the deployment mechanism. + +Unchanged by this: object states are hand-built upper-bound descriptors +that no tested self-supervised recipe reproduces, and the world is +procedurally generated. + +See `SYNTH_RESULTS.md`. + +## 2026-08-01: reporting bias, and three solver-versus-information calls + +Worlds v3 and v4 test whether the derived dictionary survives reporting +bias. v3 couples colour to shape cyclically and biases mentions toward +rare colours, so captions name half the groups and the pixel and text +colour rankings diverge in the tail; v4 replaces the cyclic coupling with +a distinct Dirichlet shape profile per colour. + +v3 results: marginal rank 4/10, joint structure 6/10, joint plus marginal +under exact enumeration of the 720 shape permutations 7/10, oracle shape +columns with measured clusters 8/10, oracle labels 10/10. Vision-side +extraction is not the limit (colour cluster purity 1.000, shape 0.998). +The 6/10 initially looked like an information bound -- ten colours over +six shapes leaves exactly four twin profiles, and exactly four appeared +-- but the twins separate by frequency, and the combined evidence has +minimum pairwise distance 0.2322 against 0.0008 for shape alone. The +bound is 10/10; what is left is one entry of solver gap and two of +estimation noise from the thinned joint table. + +v4 results: exact enumeration recovers 10/10 under the same bias and +mention rate, across frequency weights including zero. Non-degenerate +joint structure suffices on its own, which is the regime real corpora +occupy. + +Three times now a limit has looked intrinsic and proved to be the +solver's: descent retention, the triangle-energy pass, and this ceiling. +The rule stands -- compute the bound, do not infer it. + +See `SYNTH_RESULTS.md`. + +## 2026-08-01: factor families on real language + +First step of the port to natural data, text side only, on 80,000 Visual +Genome region descriptions with no lexicon and no labels. + +The synthetic pipeline found families by mutual exclusivity: template +phrases carry exactly one colour word, so colour terms never co-occur. +That does not transfer. In real descriptions colour words are +statistically independent or better -- observed over expected +co-occurrence is 1.29 for black with white, 0.93 for red with blue -- +while head nouns are the genuinely exclusive family (0.14 for car with +truck). Run as written, the method recovers clean noun classes and no +attribute families at all. Mutual exclusivity is a template artifact. + +Standard distributional word-class induction, of which exclusivity is a +degenerate case, does transfer. Positive PMI context vectors, truncated +SVD, k-means over the 300 most frequent words yields a colour family at +0.67 purity (red, brown, gray, dark, grey, pink, tan, purple) plus +coherent object categories: vehicles, animals, clothing, people, water +scenes. Colour terms split into three clusters rather than one, and the +split is structural rather than noisy -- white and blue cluster with sky +and clouds, because what a colour word modifies shapes its distribution. +That dependence is the signal cross-modal value alignment feeds on. + +Next: the vision side under unsupervised segmentation, then the field +correlation as the go/no-go before any recovery run. + +## 2026-08-01: first natural-data coordinate, and a correction + +The Tier 0 recipe was ported to Visual Genome: DINOv2 patch features, +spectral segmentation of the affinity graph, segment feature classes on +one side; distributional word classes on the other. Field correlation at +the hidden pairing is **0.166**, far under the 0.9 that recovery needs, +so by our own go/no-go no recovery run is warranted yet. + +A correction came with it. Aligning vision classes to text classes by +frequency rank and by an oracle give bit-identical fields -- 0.1662 both +ways -- while agreeing on 0 of 23 classes. Moment fields are built from +within-modality inner products, and permuting feature coordinates +consistently leaves every inner product unchanged, so relation fields are +gauge-invariant in the class labels and the dictionary never enters them. +The same was true unnoticed in the synthetic world, where oracle and +derived dictionaries produced identical fields. The dictionary is +required for encoding a new image into text coordinates, not for +matching; matching has always run on the fields alone. + +The 0.166 therefore reads as representation quality, and localises the +deficit at vision-side granularity: six segments per image, background +included, against sixteen phrases about specific objects. A sixteen- +segment run is measuring whether granularity is the binding constraint. + +## 2026-08-01: natural data reaches 0.656, and segmentation was not the wall + +Full port to Visual Genome, measured by the go/no-go statistic. Three +results overturned expectations, each checked rather than inferred. + +Discretisation was the dominant loss: hard class codes give 0.166, +continuous states 0.489. Relation fields need scene similarity only, so +quantising states first is pure attrition; the synthetic world hid this +because its states are discrete by construction. + +Segmentation is not the bottleneck on photographs. Oracle region boxes +buy 0.061 over unsupervised spectral segmentation of DINOv2 patches +(0.550 against 0.489), the inverse of the synthetic world where +segmentation was the whole gap. Flat untextured scenes are adversarial +for object discovery in a way photographs are not. + +Content projection is again the largest lever: 0.489 to **0.656**, +matching its earlier role at node level (0.21 to 0.61). Protocol-legal +and oracle-box configurations land within 0.007, so the residual is not +about object localisation. + +Control: frozen general-purpose encoders underperform corpus-fitted +statistics (0.19 for DINOv2 crop CLS against Qwen phrase states), so +in-context features and corpus-specific distributional vectors are the +better substrate. + +Verdict: 0.656 against the 0.9 recovery needs; no recovery run. The +untried lever is observation orbits, which VG lacks by construction -- +one photograph per scene. Multi-view natural data is the next protocol +requirement. + +See `NATURAL_RESULTS.md`. + +## Current experimental gate + +The node-level ordering-and-basin gate is passed on VG by content-projected +states (isolated and in-context composite). The standing rules extend +rather than retire: + +1. Blind recovery next: unpaired soft-coupling optimization (Sinkhorn over + real states, preserving the decodable path) must place calibrated + coupling mass on the hidden truth from scratch. Its score is coupling + mass at truth and verified seed precision, never retrieval means alone. +2. The view-level gate stands unchanged for the coarse-to-fine recursion: + static view fields failed it; in-context view states plus projection + are the next candidate, with replayed view truth evaluation-only. +3. Any new energy or state family still passes ordering before descent, + including single-transposition local ordering. +4. Flickr-style single-view data cannot host the current pass; two-sided + multi-view observation is a protocol requirement for now. |
