summaryrefslogtreecommitdiff
path: root/RESULTS.md
blob: a4a88f17e32df462a0790b79b1f78c344d64b68c (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
# WorldAlign MVP: first real-data result

Run date: 2026-07-28

## Question

Can independently pretrained, frozen vision and language models be aligned
without image-text pairs or a contrastive objective, and can the resulting
bridge drive a frozen causal language model?

## Protocol

- Data: Flickr30k.
- Unpaired training: 12,000 image-only examples and 12,000 caption-only
  examples with disjoint Flickr image IDs.
- Held out: 1,000 validation and 1,000 test image-caption groups.
- Vision: frozen `facebook/dinov2-small`, 384-dimensional CLS feature.
- Language: frozen `Qwen/Qwen2.5-0.5B` or
  `Qwen/Qwen2.5-1.5B`, mean-pooled final hidden state.
- Bridge: approximately 2–3M trainable parameters.
- Unpaired energy: prototype Gromov-Wasserstein correspondence, sliced
  Wasserstein distribution matching, and within-modality isometry.
- Generation interface: a text-only-trained latent-to-prefix adapter followed
  by a frozen Qwen causal LM.
- Paired bridge: same features and bridge family, trained on 12,000 pairs only
  as an interface upper bound.

No validation/test pair is used by the unpaired objective or checkpoint
selection.

## Results

### Shared relational structure exists

On 1,000 held-out paired examples used only for diagnosis:

| Text backbone | Linear CKA | Pairwise cosine Spearman | Shuffled mean | Shuffle z |
|---|---:|---:|---:|---:|
| Qwen2.5-0.5B | 0.2622 | 0.1835 | -0.0003 | 12.34 |
| Qwen2.5-1.5B | 0.2611 | 0.1837 | 0.0011 | 12.87 |

The two independently pretrained representations therefore share a strong,
non-random relational signal. Increasing Qwen from 0.5B to 1.5B does not
materially change these global geometry metrics.

### Static geometry does not uniquely recover the semantic gauge

The 128-prototype GW coupling is sharp but mostly incorrect:

| Text backbone | GW distance | Row entropy | Held-out cluster match | Chance | Oracle permutation |
|---|---:|---:|---:|---:|---:|
| Qwen2.5-0.5B | 0.1642 | 0.3875 | 1.5% | 0.78% | 20.5% |
| Qwen2.5-1.5B | 0.1218 | 0.3814 | 1.7% | 0.78% | 20.0% |

For reference, uniform 128-way entropy is about 4.85. The low entropy says GW
finds a confident structural matching; the low true match says confidence is
not identifiability.

### Frozen-backbone retrieval

All values below are from the 1,000-example held-out test split.

| Bridge | i2t R@1 | i2t R@5 | i2t R@10 | i2t median rank | t2i R@10 |
|---|---:|---:|---:|---:|---:|
| Random chance | 0.1% | 0.5% | 1.0% | about 500 | 1.0% |
| Unpaired SWD, Qwen-0.5B | 0.1% | 0.5% | 1.3% | 473 | 0.9% |
| Unpaired GW, Qwen-0.5B | 0.0% | 0.5% | 0.8% | 465 | 1.1% |
| Unpaired GW, Qwen-1.5B | 0.2% | 1.3% | 2.2% | 379 | 1.2% |
| Paired upper bound, Qwen-0.5B | 13.7% | 35.4% | 46.4% | 13 | 15.6% |

The Qwen-1.5B unpaired result is weak but reproducible. Three additional bridge
initializations obtained i2t R@10 of 2.1%, 2.3%, and 2.2%, with median ranks
384, 371, and 379. Under an independent-query binomial chance approximation,
the corresponding one-sided p-values are between 0.00027 and 0.00150. This is
evidence for partial alignment, not usable multimodal ability.

SWD-only and a linear GW bridge both remain at chance, so the Qwen-0.5B
failure is not explained just by an overly flexible MLP or by conflict between
the GW and distribution terms.

### Frozen-LLM generation

The semantic-latent-to-prefix adapter is trained only on captions and frozen
Qwen states. It never receives image features or image-text pairs.

On the same first 100 held-out test examples:

| Input to text-only prefix adapter | Best-reference unigram F1 |
|---|---:|
| True Qwen text latent | 0.5370 |
| Paired image bridge | 0.2933 |
| Unpaired image bridge | 0.2202 |
| Paired image bridge with image order shuffled | 0.2009 |

The text-only interface can make frozen Qwen decode its semantic state, and a
correctly trained vision bridge can drive that interface. The current
unpaired bridge is only slightly above the shuffled lexical baseline and its
examples are generally fluent descriptions of the wrong scene.

## Conclusion

This MVP separates three claims that are easy to conflate:

1. Independently pretrained vision and language models share non-random
   relational structure: supported.
2. Static distribution and pairwise-geometry matching are sufficient to
   identify the correct cross-modal coordinate system: not supported.
3. Once the coordinate system is correct, a small bridge can invoke a frozen
   language model's generative capability: supported by the paired control.

The next experiment should preserve this real-data, frozen-LLM endpoint but
add constraints that break semantic automorphisms: matched transition
operators, interventions, or another functional signal. Merely scaling the
same static GW/SWD objective is unlikely to close the roughly 44-point i2t
R@10 gap to the paired interface.

## Main artifacts

- `artifacts/diagnostics.json`
- `artifacts/diagnostics_qwen1p5b.json`
- `artifacts/gw.pt`
- `artifacts/gw_qwen1p5b.pt`
- `artifacts/eval_paired.json`
- `artifacts/eval_unpaired.json`
- `artifacts/eval_paired_shuffled.json`
- `artifacts/eval_unpaired_qwen1p5b.json`
- `artifacts/eval_unpaired_qwen1p5b_seed{11,22,33}.json`
- `artifacts/prefix_evaluation.json`