From a62cf4d2a99b4a7985c61b2a7feb92a82a8218b7 Mon Sep 17 00:00:00 2001 From: Yuren Hao Date: Sat, 1 Aug 2026 14:10:03 -0500 Subject: World Alignment: unpaired cross-modal correspondence by relational identifiability Method: scene states are sets of part states; relation fields are built within each modality and are invariant to how each side labels its own features; the cross-modal bridge is a coupling searched under an energy that is a closed-form functional of one matrix; solving is spectral initialisation followed by exact local refinement. Evidence: in a procedurally generated closed world, blind recovery of a hidden image-caption correspondence reaches 95.3% at 256 scenes against 0.39% chance, and the recovered pairs transfer to 200 held-out scenes at 93.0% exact retrieval with random-pair and shuffled-image controls at or near chance. Cross-modal value correspondence is derived from disjoint corpora rather than declared. On Visual Genome the field correlation reaches 0.656 against the 0.9 that polynomial recovery needs, with the deficit attributed away from segmentation and discretisation. Protocol: no image-text pair enters any objective, optimiser, initialisation, or model selection; hidden pairs score orderings only. Co-Authored-By: Claude --- VG_RESULTS.md | 127 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 127 insertions(+) create mode 100644 VG_RESULTS.md (limited to 'VG_RESULTS.md') diff --git a/VG_RESULTS.md b/VG_RESULTS.md new file mode 100644 index 0000000..e390756 --- /dev/null +++ b/VG_RESULTS.md @@ -0,0 +1,127 @@ +# Visual Genome hidden-permutation result + +Run date: 2026-07-28 + +## Question + +Does relational overcompleteness make independently pretrained visual and +language representations identifiable without exposing any image-text pairs +to the matcher? + +This experiment is deliberately more favorable than the Flickr30k experiment. +Both modalities observe the same set of 5,000 real scenes, but all node IDs and +orders are independently randomized. The permutation is stored in a private +file and loaded only for evaluation. + +## Protocol + +- 5,000 Visual Genome scenes, sampled from 107,019 scenes with at least 32 + annotated regions. +- 16 views per scene and modality. +- Vision: full-image and region-crop features from frozen + `facebook/dinov2-small`; no language-supervised vision model. +- Language: frozen final hidden states from `Qwen/Qwen2.5-1.5B`. +- No pair, source image ID, URL, coordinate, label, or region ID is available + to the matching code. +- `region_closed`: 16 human region descriptions. +- `visible_relations`: a fixed 16-view mixture of region descriptions and + explicitly visible attribute/SPO statements. +- `qa_expanded`: a fixed 16-view mixture that also includes image QA. + +Two independent invariants are evaluated: + +1. a within-scene unlabeled bundle signature (edge-distribution quantiles and + Gram spectra); +2. a between-scene 32-nearest-neighbor graph signature (local density and + multi-scale diagonal heat kernels). + +Both signatures are computed independently in each modality. They require no +learned cross-modal map. + +## Closure ablation + +Random chance at 5,000 candidates is 0.02% R@1 and 0.20% R@10. + +| Text bundle | True edge Spearman | Bundle R@10 | Graph R@1 | Graph R@10 | Graph median rank | +|---|---:|---:|---:|---:|---:| +| Region descriptions | 0.1286 | 0.26% | 0.06% (3x) | 0.62% (3.1x) | 2,069 | +| Visible attributes/relations | **0.1646** | 0.16% | **0.12% (6x)** | **0.70% (3.5x)** | **1,928** | +| QA-expanded | 0.1579 | 0.26% | 0.02% (1x) | 0.50% (2.5x) | 2,025 | + +The important result is not the absolute accuracy, which remains far from +usable. It is the controlled ordering: + +- explicit visible relations make the two scene graphs more similar and + improve blind graph matching; +- adding QA increases linguistic content but reduces identifiability relative + to the visible-relation condition; +- concatenating a weak or modality-specific bundle signature can reduce the + stronger graph-only result. + +Thus, more information in the text representation is not automatically +helpful. The useful quantity is information about the common observable +relational structure. + +## Graph-size scaling + +For the `region_closed` condition: + +| Nodes | Graph R@10 | Chance | Lift | +|---:|---:|---:|---:| +| 500 | 2.60% | 2.00% | 1.3x | +| 1,000 | 1.30% | 1.00% | 1.3x | +| 2,000 | 1.00% | 0.50% | 2.0x | +| 5,000 | 0.62% | 0.20% | 3.1x | + +Absolute recall declines because the candidate pool grows, but lift over +chance increases. Under an independence-only binomial reference, the 20 +top-ten hits at 2,000 nodes have `p=0.0034`, and the 31 hits at 5,000 nodes +have `p=7.6e-8`. These p-values are descriptive rather than calibrated because +graph-derived query ranks are dependent. + +## Negative ablation + +Three rounds of parameter-free neighborhood message passing did not improve +the 5,000-node result. Graph R@10 fell from 0.62% to 0.36%, although R@1 rose +from 0.06% to 0.08%. More graph depth is therefore not sufficient by itself. + +## Conclusion + +This experiment supports a narrow version of the hypothesis: + +> Independently pretrained vision and language representations contain shared +> relational structure, and relational coverage can accumulate enough signal +> to improve blind correspondence as the graph grows. + +It does not yet support the stronger claim that the hidden semantic gauge is +recoverable well enough to train a multimodal bridge. The best setting finds +only 6 exact top-one matches among 5,000 nodes. Feeding those assignments into +a frozen-LLM bridge would mostly train on incorrect pseudo-pairs. + +Confidence calibration makes this boundary sharper. In the +`visible_relations` setting, none of the six correct top-one matches appears +among the 1,000 queries with the largest top-one/top-two score margin. +Likewise, the 1,030 mutual-nearest assignments contain zero correct pairs. +The signal is therefore a population-level rank shift, not a usable +high-precision seed signal. + +The next gating experiment should increase common relational information +without increasing language-private information: + +1. preserve object identity across many images and text contexts; +2. use explicit visible transformations and relations rather than generic + encyclopedia prose; +3. measure confidence-calibrated seed precision, not only aggregate recall; +4. train the bridge only after a high-precision seed set can be recovered + without using the private permutation. + +## Main artifacts + +- `artifacts/vg_5k/diagnostics_region_closed.json` +- `artifacts/vg_5k/graph_diagnostics_region_closed.json` +- `artifacts/vg_5k/graph_diagnostics_region_closed_{500,1000,2000}.json` +- `artifacts/vg_5k/graph_diagnostics_region_closed_mp3.json` +- `artifacts/vg_5k_tiers/diagnostics_visible_relations.json` +- `artifacts/vg_5k_tiers/graph_diagnostics_visible_relations.json` +- `artifacts/vg_5k_tiers/diagnostics_qa_expanded.json` +- `artifacts/vg_5k_tiers/graph_diagnostics_qa_expanded.json` -- cgit v1.2.3