summaryrefslogtreecommitdiff
path: root/VG_RESULTS.md
diff options
context:
space:
mode:
Diffstat (limited to 'VG_RESULTS.md')
-rw-r--r--VG_RESULTS.md127
1 files changed, 127 insertions, 0 deletions
diff --git a/VG_RESULTS.md b/VG_RESULTS.md
new file mode 100644
index 0000000..e390756
--- /dev/null
+++ b/VG_RESULTS.md
@@ -0,0 +1,127 @@
+# Visual Genome hidden-permutation result
+
+Run date: 2026-07-28
+
+## Question
+
+Does relational overcompleteness make independently pretrained visual and
+language representations identifiable without exposing any image-text pairs
+to the matcher?
+
+This experiment is deliberately more favorable than the Flickr30k experiment.
+Both modalities observe the same set of 5,000 real scenes, but all node IDs and
+orders are independently randomized. The permutation is stored in a private
+file and loaded only for evaluation.
+
+## Protocol
+
+- 5,000 Visual Genome scenes, sampled from 107,019 scenes with at least 32
+ annotated regions.
+- 16 views per scene and modality.
+- Vision: full-image and region-crop features from frozen
+ `facebook/dinov2-small`; no language-supervised vision model.
+- Language: frozen final hidden states from `Qwen/Qwen2.5-1.5B`.
+- No pair, source image ID, URL, coordinate, label, or region ID is available
+ to the matching code.
+- `region_closed`: 16 human region descriptions.
+- `visible_relations`: a fixed 16-view mixture of region descriptions and
+ explicitly visible attribute/SPO statements.
+- `qa_expanded`: a fixed 16-view mixture that also includes image QA.
+
+Two independent invariants are evaluated:
+
+1. a within-scene unlabeled bundle signature (edge-distribution quantiles and
+ Gram spectra);
+2. a between-scene 32-nearest-neighbor graph signature (local density and
+ multi-scale diagonal heat kernels).
+
+Both signatures are computed independently in each modality. They require no
+learned cross-modal map.
+
+## Closure ablation
+
+Random chance at 5,000 candidates is 0.02% R@1 and 0.20% R@10.
+
+| Text bundle | True edge Spearman | Bundle R@10 | Graph R@1 | Graph R@10 | Graph median rank |
+|---|---:|---:|---:|---:|---:|
+| Region descriptions | 0.1286 | 0.26% | 0.06% (3x) | 0.62% (3.1x) | 2,069 |
+| Visible attributes/relations | **0.1646** | 0.16% | **0.12% (6x)** | **0.70% (3.5x)** | **1,928** |
+| QA-expanded | 0.1579 | 0.26% | 0.02% (1x) | 0.50% (2.5x) | 2,025 |
+
+The important result is not the absolute accuracy, which remains far from
+usable. It is the controlled ordering:
+
+- explicit visible relations make the two scene graphs more similar and
+ improve blind graph matching;
+- adding QA increases linguistic content but reduces identifiability relative
+ to the visible-relation condition;
+- concatenating a weak or modality-specific bundle signature can reduce the
+ stronger graph-only result.
+
+Thus, more information in the text representation is not automatically
+helpful. The useful quantity is information about the common observable
+relational structure.
+
+## Graph-size scaling
+
+For the `region_closed` condition:
+
+| Nodes | Graph R@10 | Chance | Lift |
+|---:|---:|---:|---:|
+| 500 | 2.60% | 2.00% | 1.3x |
+| 1,000 | 1.30% | 1.00% | 1.3x |
+| 2,000 | 1.00% | 0.50% | 2.0x |
+| 5,000 | 0.62% | 0.20% | 3.1x |
+
+Absolute recall declines because the candidate pool grows, but lift over
+chance increases. Under an independence-only binomial reference, the 20
+top-ten hits at 2,000 nodes have `p=0.0034`, and the 31 hits at 5,000 nodes
+have `p=7.6e-8`. These p-values are descriptive rather than calibrated because
+graph-derived query ranks are dependent.
+
+## Negative ablation
+
+Three rounds of parameter-free neighborhood message passing did not improve
+the 5,000-node result. Graph R@10 fell from 0.62% to 0.36%, although R@1 rose
+from 0.06% to 0.08%. More graph depth is therefore not sufficient by itself.
+
+## Conclusion
+
+This experiment supports a narrow version of the hypothesis:
+
+> Independently pretrained vision and language representations contain shared
+> relational structure, and relational coverage can accumulate enough signal
+> to improve blind correspondence as the graph grows.
+
+It does not yet support the stronger claim that the hidden semantic gauge is
+recoverable well enough to train a multimodal bridge. The best setting finds
+only 6 exact top-one matches among 5,000 nodes. Feeding those assignments into
+a frozen-LLM bridge would mostly train on incorrect pseudo-pairs.
+
+Confidence calibration makes this boundary sharper. In the
+`visible_relations` setting, none of the six correct top-one matches appears
+among the 1,000 queries with the largest top-one/top-two score margin.
+Likewise, the 1,030 mutual-nearest assignments contain zero correct pairs.
+The signal is therefore a population-level rank shift, not a usable
+high-precision seed signal.
+
+The next gating experiment should increase common relational information
+without increasing language-private information:
+
+1. preserve object identity across many images and text contexts;
+2. use explicit visible transformations and relations rather than generic
+ encyclopedia prose;
+3. measure confidence-calibrated seed precision, not only aggregate recall;
+4. train the bridge only after a high-precision seed set can be recovered
+ without using the private permutation.
+
+## Main artifacts
+
+- `artifacts/vg_5k/diagnostics_region_closed.json`
+- `artifacts/vg_5k/graph_diagnostics_region_closed.json`
+- `artifacts/vg_5k/graph_diagnostics_region_closed_{500,1000,2000}.json`
+- `artifacts/vg_5k/graph_diagnostics_region_closed_mp3.json`
+- `artifacts/vg_5k_tiers/diagnostics_visible_relations.json`
+- `artifacts/vg_5k_tiers/graph_diagnostics_visible_relations.json`
+- `artifacts/vg_5k_tiers/diagnostics_qa_expanded.json`
+- `artifacts/vg_5k_tiers/graph_diagnostics_qa_expanded.json`