From 0ba1fe093c8393c473fd13ec44d07cacb83ec33f Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Sat, 1 Aug 2026 16:29:37 -0500 Subject: Corpus overlap sets the shared width, by intervention Suppressing one factor block from the synthetic captions -- vision still sees the property, nothing else changes -- costs five shared directions and the entire result: 96.4% recovery to 5.6%, at a correlation of 0.830 that is higher than anything achieved on photographs. The synthetic world degraded to Visual Genome's shared width fails exactly the way Visual Genome fails, at a correlation Visual Genome never reaches. That closes the chain: how much of the same world the two corpora describe sets the width, the width sets recovery, and the correlation reports on neither reliably. Also records the Delta compute-node offline trap in DELTA_HPC.md, and marks the superseded verdict in NATURAL_RESULTS. Co-Authored-By: Claude --- RANK_RESULTS.md | 31 +++++++++++++++++++++++++++++++ logs/obj_large_gpu.log | 2 +- logs/omit_rest.log | 0 3 files changed, 32 insertions(+), 1 deletion(-) create mode 100644 logs/omit_rest.log diff --git a/RANK_RESULTS.md b/RANK_RESULTS.md index 0e481b9..2998e47 100644 --- a/RANK_RESULTS.md +++ b/RANK_RESULTS.md @@ -118,6 +118,37 @@ directions and then stops. Enlarging the scene population does not help either scenes, so the ceiling is a property of what the two corpora are about rather than of how many scenes are sampled. +## Corpus overlap sets the width, by intervention + +The invariance results say the width does not come from the encoders. A +controlled intervention says where it does come from. In the synthetic world +the caption is encoded into separate factor blocks, so a block can be +suppressed — the caption then never states that property while vision +continues to see it. Nothing else changes: same images, same segmentation, +same kernel. + +| Caption content | ρ | shared | recovery | +|---|---|---|---| +| states everything | 0.928 | 20 | **96.4%** | +| never states size | 0.830 | 15 | **5.6%** | + +Deleting one word class costs five shared directions and the entire result. +The correlation is still 0.830 — higher than anything we have achieved on +photographs — and recovery is at chance. **The synthetic world degraded to +Visual Genome's shared width fails exactly the way Visual Genome fails**, at a +correlation Visual Genome never reaches. + +This closes the causal chain. How much of the same world the two corpora +describe sets the width of the shared spectrum; the width sets recovery; and +the correlation reports on neither reliably. It also reframes the natural-data +deficit as a statement about Visual Genome rather than about our pipeline: a +region description simply does not say as much about a photograph as the +synthetic captions say about their scenes. + +(Widths in this table are computed on standardised fields and are not +comparable entry-for-entry with the tables above, which use raw fields; within +the table the convention is identical, which is what the comparison needs.) + ## Two smaller results, one refuted hypothesis **The projection improved.** Scaling each discriminant direction by its own diff --git a/logs/obj_large_gpu.log b/logs/obj_large_gpu.log index 748f6bd..8dfb4fe 100644 --- a/logs/obj_large_gpu.log +++ b/logs/obj_large_gpu.log @@ -1,3 +1,3 @@ Using a slow image processor as `use_fast` is unset and a slow processor was saved with this model. `use_fast=True` will be the default behavior in v4.52, even if the model was saved with a slow processor. This will result in minor differences in outputs. You'll still be able to use a slow processor with `use_fast=False`. `torch_dtype` is deprecated! Use `dtype` instead! - segment: 0%| | 0/250 [00:00