diff options
Diffstat (limited to 'VG_RESULTS.md')
| -rw-r--r-- | VG_RESULTS.md | 127 |
1 files changed, 127 insertions, 0 deletions
diff --git a/VG_RESULTS.md b/VG_RESULTS.md new file mode 100644 index 0000000..e390756 --- /dev/null +++ b/VG_RESULTS.md @@ -0,0 +1,127 @@ +# Visual Genome hidden-permutation result + +Run date: 2026-07-28 + +## Question + +Does relational overcompleteness make independently pretrained visual and +language representations identifiable without exposing any image-text pairs +to the matcher? + +This experiment is deliberately more favorable than the Flickr30k experiment. +Both modalities observe the same set of 5,000 real scenes, but all node IDs and +orders are independently randomized. The permutation is stored in a private +file and loaded only for evaluation. + +## Protocol + +- 5,000 Visual Genome scenes, sampled from 107,019 scenes with at least 32 + annotated regions. +- 16 views per scene and modality. +- Vision: full-image and region-crop features from frozen + `facebook/dinov2-small`; no language-supervised vision model. +- Language: frozen final hidden states from `Qwen/Qwen2.5-1.5B`. +- No pair, source image ID, URL, coordinate, label, or region ID is available + to the matching code. +- `region_closed`: 16 human region descriptions. +- `visible_relations`: a fixed 16-view mixture of region descriptions and + explicitly visible attribute/SPO statements. +- `qa_expanded`: a fixed 16-view mixture that also includes image QA. + +Two independent invariants are evaluated: + +1. a within-scene unlabeled bundle signature (edge-distribution quantiles and + Gram spectra); +2. a between-scene 32-nearest-neighbor graph signature (local density and + multi-scale diagonal heat kernels). + +Both signatures are computed independently in each modality. They require no +learned cross-modal map. + +## Closure ablation + +Random chance at 5,000 candidates is 0.02% R@1 and 0.20% R@10. + +| Text bundle | True edge Spearman | Bundle R@10 | Graph R@1 | Graph R@10 | Graph median rank | +|---|---:|---:|---:|---:|---:| +| Region descriptions | 0.1286 | 0.26% | 0.06% (3x) | 0.62% (3.1x) | 2,069 | +| Visible attributes/relations | **0.1646** | 0.16% | **0.12% (6x)** | **0.70% (3.5x)** | **1,928** | +| QA-expanded | 0.1579 | 0.26% | 0.02% (1x) | 0.50% (2.5x) | 2,025 | + +The important result is not the absolute accuracy, which remains far from +usable. It is the controlled ordering: + +- explicit visible relations make the two scene graphs more similar and + improve blind graph matching; +- adding QA increases linguistic content but reduces identifiability relative + to the visible-relation condition; +- concatenating a weak or modality-specific bundle signature can reduce the + stronger graph-only result. + +Thus, more information in the text representation is not automatically +helpful. The useful quantity is information about the common observable +relational structure. + +## Graph-size scaling + +For the `region_closed` condition: + +| Nodes | Graph R@10 | Chance | Lift | +|---:|---:|---:|---:| +| 500 | 2.60% | 2.00% | 1.3x | +| 1,000 | 1.30% | 1.00% | 1.3x | +| 2,000 | 1.00% | 0.50% | 2.0x | +| 5,000 | 0.62% | 0.20% | 3.1x | + +Absolute recall declines because the candidate pool grows, but lift over +chance increases. Under an independence-only binomial reference, the 20 +top-ten hits at 2,000 nodes have `p=0.0034`, and the 31 hits at 5,000 nodes +have `p=7.6e-8`. These p-values are descriptive rather than calibrated because +graph-derived query ranks are dependent. + +## Negative ablation + +Three rounds of parameter-free neighborhood message passing did not improve +the 5,000-node result. Graph R@10 fell from 0.62% to 0.36%, although R@1 rose +from 0.06% to 0.08%. More graph depth is therefore not sufficient by itself. + +## Conclusion + +This experiment supports a narrow version of the hypothesis: + +> Independently pretrained vision and language representations contain shared +> relational structure, and relational coverage can accumulate enough signal +> to improve blind correspondence as the graph grows. + +It does not yet support the stronger claim that the hidden semantic gauge is +recoverable well enough to train a multimodal bridge. The best setting finds +only 6 exact top-one matches among 5,000 nodes. Feeding those assignments into +a frozen-LLM bridge would mostly train on incorrect pseudo-pairs. + +Confidence calibration makes this boundary sharper. In the +`visible_relations` setting, none of the six correct top-one matches appears +among the 1,000 queries with the largest top-one/top-two score margin. +Likewise, the 1,030 mutual-nearest assignments contain zero correct pairs. +The signal is therefore a population-level rank shift, not a usable +high-precision seed signal. + +The next gating experiment should increase common relational information +without increasing language-private information: + +1. preserve object identity across many images and text contexts; +2. use explicit visible transformations and relations rather than generic + encyclopedia prose; +3. measure confidence-calibrated seed precision, not only aggregate recall; +4. train the bridge only after a high-precision seed set can be recovered + without using the private permutation. + +## Main artifacts + +- `artifacts/vg_5k/diagnostics_region_closed.json` +- `artifacts/vg_5k/graph_diagnostics_region_closed.json` +- `artifacts/vg_5k/graph_diagnostics_region_closed_{500,1000,2000}.json` +- `artifacts/vg_5k/graph_diagnostics_region_closed_mp3.json` +- `artifacts/vg_5k_tiers/diagnostics_visible_relations.json` +- `artifacts/vg_5k_tiers/graph_diagnostics_visible_relations.json` +- `artifacts/vg_5k_tiers/diagnostics_qa_expanded.json` +- `artifacts/vg_5k_tiers/graph_diagnostics_qa_expanded.json` |
