summaryrefslogtreecommitdiff
path: root/README.md
diff options
context:
space:
mode:
authorYuren Hao <blackhao0426@gmail.com>2026-08-01 14:10:03 -0500
committerYuren Hao <blackhao0426@gmail.com>2026-08-01 14:10:03 -0500
commita62cf4d2a99b4a7985c61b2a7feb92a82a8218b7 (patch)
treeee2248078db7edf3812a07f195afa3d9bd6f10c6 /README.md
World Alignment: unpaired cross-modal correspondence by relational identifiability
Method: scene states are sets of part states; relation fields are built within each modality and are invariant to how each side labels its own features; the cross-modal bridge is a coupling searched under an energy that is a closed-form functional of one matrix; solving is spectral initialisation followed by exact local refinement. Evidence: in a procedurally generated closed world, blind recovery of a hidden image-caption correspondence reaches 95.3% at 256 scenes against 0.39% chance, and the recovered pairs transfer to 200 held-out scenes at 93.0% exact retrieval with random-pair and shuffled-image controls at or near chance. Cross-modal value correspondence is derived from disjoint corpora rather than declared. On Visual Genome the field correlation reaches 0.656 against the 0.9 that polynomial recovery needs, with the deficit attributed away from segmentation and discretisation. Protocol: no image-text pair enters any objective, optimiser, initialisation, or model selection; hidden pairs score orderings only. Co-Authored-By: Claude <noreply@anthropic.com>
Diffstat (limited to 'README.md')
-rw-r--r--README.md221
1 files changed, 221 insertions, 0 deletions
diff --git a/README.md b/README.md
new file mode 100644
index 0000000..3fbb763
--- /dev/null
+++ b/README.md
@@ -0,0 +1,221 @@
+# WorldAlign
+
+An end-to-end test of whether independently pretrained vision and language
+representations can be aligned without image-text pairs, then used as the input
+interface to a frozen causal language model.
+
+For the project-level hypothesis, intended final system, non-goals, evidence
+standards, and current experimental gate, read
+[PROJECT_CONCEPT.md](PROJECT_CONCEPT.md). The chronological research decisions
+and falsified directions are maintained in [LAB_NOTES.md](LAB_NOTES.md).
+
+The first completed real-data run finds strong shared relational geometry and
+a small reproducible unpaired retrieval signal with Qwen2.5-1.5B, but the
+static GW/SWD energy does not yet recover a usable semantic bridge. The paired
+control confirms that the frozen-backbone interface itself works. See
+[RESULTS.md](RESULTS.md) for exact numbers and interpretation.
+
+A second hidden-permutation experiment uses 5,000 Visual Genome scenes with
+16 visual/text views per node. It finds that explicitly visible relations
+improve blind graph matching to 6x chance at R@1, while adding QA weakens it.
+This isolates the relevant requirement as common relational information, not
+total information in either representation. See
+[VG_RESULTS.md](VG_RESULTS.md) and [VG_PROTOCOL.md](VG_PROTOCOL.md).
+
+The third experiment removes the bridge entirely. It performs iterative
+energy descent on free Qwen-side latent particles, conditioned by frozen DINO
+population relations. Five-caption observation orbits strengthen the
+correct-vs-shuffled energy gap. However, a convergence audit shows that further
+reduction of the fixed energy returns retrieval to chance, and the true paired
+configuration has higher energy than incorrect optimized states. The present
+energy is therefore falsified as an identifiable alignment objective; its
+early rank gain is only a trajectory transient. See
+[ENERGY_RESULTS.md](ENERGY_RESULTS.md) and [LAB_NOTES.md](LAB_NOTES.md).
+
+The current experiment closes the remaining escape route by restricting
+configurations to permutations of real frozen text states, then testing the
+ordering of the energy directly. The true assignment beats every random
+permutation globally but is never a local minimum: 7--33% of single swaps
+lower the energy, exact 2-swap descent walks away from the truth, and
+random-start descent finds on-manifold counterfeits below the true energy in
+every condition. Spectral audits locate the cause in the states rather than
+the objective: the cross-modal shared signal occupies roughly sixteen
+population directions, and the fine structure needed for instance identity is
+modality-private. A view-permutation replay then shows that fine-grained
+shared structure does exist one level down, between the sixteen region views
+inside each Visual Genome node (z = 46.6, coarse-frame view matching at 2.3x
+chance). The next candidate energy must therefore be defined over structured
+view-level states, not pooled node vectors. Two controls sharpen this: an
+Ollivier-Ricci flow more than doubles the global ranking separation while
+leaving every local verdict unchanged (geometry is a frame ingredient, not a
+resolver), and static view-level relation fields fail the same ordering gate
+(no strict local minimum among 5,000 nodes; frame-free matching at chance).
+See [MANIFOLD_RESULTS.md](MANIFOLD_RESULTS.md). The requirements
+decomposition for the next representation structure and alignment process is
+maintained in [STRUCTURE_DESIGN.md](STRUCTURE_DESIGN.md).
+
+The first two requirement batteries lift the ceiling. Whitening the
+within-scene view noise (fitted without pairs) triples the VG cross-modal
+relational correlation to 0.61 and collapses improving swaps to 0.2%;
+joint-context encoding shows the model layer's shared signal lives in its
+mixed states (rho 0.216 vs 0.128 isolated), not its attention weights. On
+the resulting states the basin audit passes for the first time: across nine
+VG conditions (two state families, three subset seeds, ten restarts each)
+every descent trajectory ends 5--19% above the true configuration. Flickr
+still fails by under 1% -- one visual view leaves nothing to whiten -- so
+two-sided multi-view observation is a protocol requirement. The passed gate
+licenses unpaired soft-coupling recovery, which is the next phase. See
+[BATTERY_RESULTS.md](BATTERY_RESULTS.md).
+
+A synthetic closed world with exact closure and from-scratch unimodal
+towers turns data sufficiency into a construction rather than an
+assumption. Its first measurement relocates the bottleneck: with perfect
+closure, global similarity matches natural data while relational sharing
+collapses six-fold, because InfoNCE's uniformity erases between-scene
+geometry. Sufficiency of the data does not transfer to sufficiency of the
+representation; the unimodal training objective is a phase-diagram axis.
+See [SYNTH_RESULTS.md](SYNTH_RESULTS.md).
+
+The synthetic world then produced the project's first end-to-end result.
+Fixing segmentation with unimodal perceptual operations (Gestalt
+appearance grouping, distance-transform watershed, extent-based size
+classes) raised the cross-modal field correlation from 0.57 to 0.93, past
+the threshold where polynomial matching becomes possible, and a spectral
+solver composed with exact energy refinement recovers **95.3%** of a
+hidden image-caption correspondence at N=256 -- no pairs, no declared
+dictionary, no learned cross-modal map, with the colour correspondence
+itself derived from unimodal marginal frequency ranks. The recovered
+correspondence then transfers: spent as pseudo-pairs for a ridge map, it
+reaches 93.0% exact retrieval on 200 held-out scenes that took no part in
+matching (0.5% chance), with the random-pair control at chance and the
+shuffled-image control collapsing colour F1 from 0.989 to 0.295. The
+machinery works end to end when representations are good enough; whether
+self-supervised learning produces such representations remains open. See
+[SYNTH_RESULTS.md](SYNTH_RESULTS.md).
+
+The first recovery attempt maps the search phase: parallel tempering,
+entropic Sinkhorn, and a spectral-band homotopy all fail to enter the true
+funnel, with tempering equilibrating exactly into the quench band 9--20%
+above the truth. The problem sits in the hard phase of planted assignment:
+statistically identifiable, algorithmically out of reach for local
+dynamics. The mathematical tool inventory for closing that gap -- symmetric
+group Fourier band-limiting, cavity methods, operator-algebraic
+intertwiners, sheaf-cohomological gluing, rigidity thresholds -- is mapped
+to measured obstacles in [STRUCTURE_DESIGN.md](STRUCTURE_DESIGN.md). See
+[RECOVERY_RESULTS.md](RECOVERY_RESULTS.md).
+
+The first experiment uses:
+
+- Flickr30k real images and captions;
+- a frozen image-only DINOv2 encoder;
+- a frozen text-only Qwen2.5 causal LM;
+- disjoint image-only and text-only training subsets;
+- Gromov-Wasserstein structure matching plus distributional and isometry
+ energies;
+- a text-only-trained semantic-latent-to-prefix adapter;
+- a paired bridge trained on the same backbones as an interface upper bound.
+
+The paired validation/test data are never used for unsupervised training or
+checkpoint selection.
+
+## Protocol
+
+```text
+unpaired images -> frozen DINO -> visual features --\
+ > structural bridge
+unpaired captions -> frozen Qwen -> text features ---/
+
+captions -> frozen Qwen text feature -> text-only prefix adapter
+
+image -> DINO -> bridge -> prefix adapter -> frozen Qwen -> caption
+```
+
+The unpaired split is stronger than shuffling captions: the image-only and
+text-only subsets contain disjoint Flickr image IDs.
+
+## Commands
+
+```bash
+python -m worldalign.prepare --output artifacts/manifest.json
+
+python -m worldalign.extract_vision \
+ --manifest artifacts/manifest.json \
+ --output artifacts/vision.pt \
+ --device cuda:1
+
+python -m worldalign.extract_text \
+ --manifest artifacts/manifest.json \
+ --output artifacts/text.pt \
+ --device cuda:3
+
+python -m worldalign.diagnose \
+ --manifest artifacts/manifest.json \
+ --vision artifacts/vision.pt \
+ --text artifacts/text.pt
+
+python -m worldalign.precompute_gw \
+ --manifest artifacts/manifest.json \
+ --vision artifacts/vision.pt \
+ --text artifacts/text.pt \
+ --output artifacts/gw.pt
+
+python -m worldalign.train_bridge \
+ --mode unpaired_gw \
+ --manifest artifacts/manifest.json \
+ --vision artifacts/vision.pt \
+ --text artifacts/text.pt \
+ --gw-cache artifacts/gw.pt \
+ --output artifacts/bridge_unpaired.pt \
+ --device cuda:1
+
+python -m worldalign.train_bridge \
+ --mode paired \
+ --manifest artifacts/manifest.json \
+ --vision artifacts/vision.pt \
+ --text artifacts/text.pt \
+ --output artifacts/bridge_paired.pt \
+ --device cuda:3
+
+python -m worldalign.train_prefix \
+ --manifest artifacts/manifest.json \
+ --text artifacts/text.pt \
+ --output artifacts/prefix.pt \
+ --device cuda:1
+
+python -m worldalign.evaluate_prefix \
+ --manifest artifacts/manifest.json \
+ --vision artifacts/vision.pt \
+ --text artifacts/text.pt \
+ --prefix artifacts/prefix.pt \
+ --device cuda:1
+
+python -m worldalign.evaluate \
+ --manifest artifacts/manifest.json \
+ --vision artifacts/vision.pt \
+ --text artifacts/text.pt \
+ --bridge artifacts/bridge_unpaired.pt \
+ --prefix artifacts/prefix.pt \
+ --device cuda:1
+```
+
+The on-manifold assignment gate and the view-level probe:
+
+```bash
+python -m worldalign.manifold_gate \
+ --dataset flickr --split test --samples 512 --text-mode orbit_mean \
+ --output artifacts/manifold_gate/flickr_test512_orbit.json
+
+python -m worldalign.manifold_gate \
+ --dataset vg --vg-text artifacts/vg_5k/text_features.pt \
+ --samples 512 --subset-seed 0 \
+ --output artifacts/manifold_gate/vg512_region_closed_seed0_scalar.json
+
+python -m worldalign.vg_view_probe \
+ --context-nodes 1000 --skip-floor \
+ --output artifacts/manifold_gate/vg_view_probe_full.json
+```
+
+Use `--help` on each command for smoke-test size controls. The default first
+run uses Qwen2.5-0.5B and DINOv2-small to validate the experimental signal.
+Once the protocol works, model names can be replaced without changing the
+method.