summaryrefslogtreecommitdiff
path: root/BATTERY_RESULTS.md
diff options
context:
space:
mode:
Diffstat (limited to 'BATTERY_RESULTS.md')
-rw-r--r--BATTERY_RESULTS.md117
1 files changed, 117 insertions, 0 deletions
diff --git a/BATTERY_RESULTS.md b/BATTERY_RESULTS.md
new file mode 100644
index 0000000..a92d257
--- /dev/null
+++ b/BATTERY_RESULTS.md
@@ -0,0 +1,117 @@
+# Representation batteries: content projection and model-layer readout
+
+Run date: 2026-07-29
+
+## Question
+
+Every energy tested before today failed the basin audit: descent from any
+start found configurations below the true one. The manifold gate located
+the cause in the node states, and `STRUCTURE_DESIGN.md` distilled it into
+requirement batteries. This report covers the first two batteries -- R6
+(content-subspace projection) and R3 (model-layer relational readout) --
+and the full assignment gate on the resulting states. Both batteries
+change the states only; the energy, the gate machinery, and the hidden
+evaluation protocol are unchanged.
+
+## R6: within-scene noise whitening lifts the ceiling
+
+Views of the same scene share content and differ in style. Directions
+maximizing between-scene over within-scene variance are the solution of a
+generalized eigenproblem fitted per modality on unpaired orbit structure:
+Flickr directions on the 12,000 text-only training orbits, VG directions
+only on nodes outside the evaluated subset. The contrast with population
+whitening carries the entire lesson: population whitening equalizes the
+total covariance and collapsed the cross-modal signal to 0.03; whitening
+the within-scene covariance instead amplifies exactly the directions where
+independent redescriptions agree.
+
+| Condition | Baseline rho / improving swaps | Projected best |
+|---|---|---|
+| Flickr orbit mean | 0.298 / 7.8% | 0.358 / 0.5% (k=32) |
+| Flickr single caption | 0.189 / 15.6% | 0.326 (k=16) |
+| VG region_closed, both sides | 0.213 / 23.5% | 0.614 (k=8); 0.25% swaps (k=128) |
+
+A single projected caption carries more shared relational signal than the
+five-caption orbit mean of the raw states. On VG the cross-modal Spearman
+nearly triples.
+
+## R3: the model layer shares its states, not its attention
+
+For each VG node the sixteen phrases were encoded jointly in one Qwen
+context (four phrase orders, cross-phrase attention pooled per layer
+group, in-context phrase states from the final layer), and each image was
+encoded once with DINO patch attention pooled over region-box pairs. At
+the replayed true view correspondence, on all 5,000 nodes:
+
+| Relation source | Matched rho | Gap z |
+|---|---:|---:|
+| Isolated-encoding cosine (baseline) | 0.128 | 46.6 |
+| Attention x attention, best layer pair | 0.076 | 6.8 |
+| In-context cosine x in-context cosine | **0.216** | **69.6** |
+
+Attention weights are modality-specialized -- DINO's spatial attention and
+Qwen's causal attention share almost nothing directly. What the model
+layer shares is the result of context mixing: **in-context states carry
+1.7 times the cross-modal relational signal of isolated encodings.** The
+"relations live in the model layer" hypothesis holds with that
+refinement: the readout is the mixed states, not the mixing weights.
+
+## The assignment gate passes on Visual Genome
+
+Full gate on content-projected states (`content_gate.py`), with ten
+random-restart quenches per run and subset seeds 0--2. "Margin" is how far
+the best quench ends above the true energy; positive margin with near-zero
+accuracy quenches means descent found nothing below the truth.
+
+| Condition | z | Improving swaps | Kept | Margin | Passes |
+|---|---:|---:|---:|---:|---|
+| VG isolated states, k=128, seeds 0/1/2 | 207--219 | 0.25--0.28% | 52--54% | +5.1--10.8% | yes, 3/3 |
+| VG isolated states, k=256, seeds 0/1/2 | 199--212 | 0.19--0.23% | 54--59% | +7.8--15.8% | yes, 3/3 |
+| VG in-context + projection, k=64/128/256 | 209--213 | 0.17--0.23% | 52--59% | +11.6--18.5% | yes, 3/3 |
+| Flickr text-side projection, k=128/256 | 141--147 | 0.23--0.25% | 56--57% | -0.9---0.6% | no |
+
+Nine of nine VG conditions pass: for the first time in the project, no
+descent trajectory from any start reaches an energy below the true
+configuration. The in-context composite widens the margin over isolated
+states by roughly a factor of two at matched k. Flickr fails by under one
+percent: with a single visual view there is no within-scene visual orbit
+to whiten, so the visual side's private fine structure stays in the
+energy. The pass is therefore attributable to two-sided multi-view
+structure, which is a data-protocol requirement, not a dataset accident.
+
+## What did not pass
+
+- The true assignment is still not a strict local minimum anywhere:
+ 0.17--0.28% of transpositions improve it (roughly 200--350 swaps of
+ 131k), and exact descent from the truth keeps 52--59% of nodes. The
+ passed gate licenses soft-coupling recovery, not exact hard matching.
+- Flickr at node level, per the single-view limitation above.
+- Blind recovery is untested: everything here evaluates orderings around
+ the hidden truth. Nothing yet demonstrates that unpaired optimization
+ finds this basin from scratch.
+
+## Conclusion
+
+The information ceiling diagnosed by the manifold gate was real and was
+liftable by two purely unimodal operations: whiten the within-scene view
+noise, and read states after the model's own context mixing. On
+two-sided multi-view data the resulting energies order the true
+configuration below everything descent can reach, with margins of 5--19%
+replicated over three subset seeds and ten restarts.
+
+The next phase is the alignment process under the P-requirements of
+`STRUCTURE_DESIGN.md`: soft-coupling recovery (Sinkhorn over real states,
+which preserves the decodable path back to the frozen LLM), scored by
+coupling mass on the hidden truth and by calibrated seed precision, then
+the coarse-to-fine recursion onto view level. The passed node-level gate
+is the license for that phase, not a claim of usable multimodal ability.
+
+## Main artifacts
+
+- `artifacts/manifold_gate/content_projection.json`
+- `artifacts/manifold_gate/content_gate_{flickr,vg}_k{16,32,128}.json`
+- `artifacts/manifold_gate/confirm_vg_k{128,256}_seed{0,1,2}.json`
+- `artifacts/manifold_gate/confirm_vg_ctx_k{64,128,256}.json`
+- `artifacts/manifold_gate/confirm_flickr_k256.json`
+- `artifacts/manifold_gate/attn_{text,vision}_full.pt`
+- `artifacts/manifold_gate/attention_probe_full.json`