| Age | Commit message (Collapse) | Author |
|
Seven text encoders against a fixed DINOv2 vision side, priced by the
anchor bound rather than the retired field correlation:
PPMI 0.289 BGE-large 335M 0.308
MiniLM 22M 0.328 Qwen 0.5B 0.325
mpnet 110M 0.341 Qwen 1.5B 0.311
bert 110M 0.341 part-oracle ceiling 0.359
No correlation with scale and the largest model is not the best; a
retrieval-tuned 335M encoder is worse than bert-base. All seven sit in a
0.29-0.34 band against a part-oracle ceiling of 0.359, so swapping in a
stronger off-the-shelf encoder is closed as a route.
That the band is narrow across unrelated architectures and training sets
is itself the finding: what limits them is common to all of them. They
are generic text encoders whose similarity geometry is organised around
distinctions that do not line up with DINOv2's. Changing which one is
used cannot fix that; changing how both towers are trained might, and
that variable has never been touched.
Also records the VG part-level oracle: annotator-supplied box-to-phrase
correspondence buys 0.023 (0.336 -> 0.359), so part correspondence was
never the problem either.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
Each VG region box is index-aligned with its own description, so the crop
can be encoded and paired with the phrase describing it -- part
correspondence by annotation, which is cross-modal supervision a deployed
system would not have. The cross-modal anchor bound goes 0.336 -> 0.359.
So the two encoders do not covary at the part level either, and no
segmentation, aggregation, kernel or backbone repairs that. The 0.805
within-text ceiling was never the cross-modal ceiling; for this encoder
pair on this corpus the cross-modal ceiling is about 0.36 against the 0.9
recovery needs.
Six representational interventions this session were chosen without a
ceiling in view -- the same error being made on the matching side at the
same time. What replaces them is a three-number screening protocol that
runs before any pipeline is built.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
Adversarial review of the artifacts found three of my numbers to be
artifacts of my own code. All three reproduced here before acceptance:
- anchor_bound presented probe rows in the same index order on both
sides, so exact twins had their tie broken onto the diagonal. 0.997 ->
0.920 on omit-size. Fixed by scrambling the T-side presentation.
- The truth is not a strict local minimum: 51 transpositions have
exactly zero energy delta. fast_pair_descent only looked stationary
because its break test treats zero as no-improvement.
- scipy's FAQ takes no n_init, so it was swallowed into unknown_options
and 'FAQ x30 restarts' computed bit-identically to plain FAQ. Replaced
with a real restart loop over P0='randomized'.
The blind ceiling for omit-size is 0.836, not 1.0: the text field has 51
exact transposition automorphisms, so T[s,s] is bitwise identical to T
and no objective f(V, P T P^T) can separate an orbit at any order. Every
synthetic accuracy was being divided by the wrong denominator.
The fifteen failed solvers share one property -- all purely quadratic or
purely spectral, none with a node-level term. Eight moments of each
node's own field row, blended with the quadratic term through
Frank-Wolfe, reach 0.837 with an energy gap of exactly zero. The term
must stay in the loop: as a seed for pure-quadratic descent it scores
0.21, pinned through the iterations it scores 0.84 -- which is also why
amplification plateaued, being itself pure-quadratic.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
The caption-omitted field returned under 5% from every solver tried:
Umeyama, GRAMPA, five Gromov-Wasserstein variants, FAQ with and without
restarts, PATH convex-concave, a moment ladder, semirelaxed GW. The
measurements said the answer was not a sixteenth solver.
Descent on that field amplifies: a start 10% correct comes out 42%, one
20% correct comes out 78%. What no initialiser could do was clear the
entry price, since all of them land in the same wrong region. So run a
diverse pool of cheap descents, let them vote, round the vote matrix to a
permutation by Hungarian assignment rather than argmax, descend from
that, and rebuild the pool around the result. Each round feeds the
amplifier a better start.
0.044 -> 0.72 on the best run, 0.54 mean over two. Competitive elsewhere:
0.977 on the easy field against 0.961 for the best GW variant.
Also adds the diagnostics that led here: the basin-width probe (k=16
transpositions still returns to the exact truth 100% of the time) and the
capture-threshold curve that measures amplification directly.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
a benchmark
The gate settles which failure mode each field is in, and refutes the
symmetry hypothesis I proposed. On the caption-omitted field the truth is
a STRICT local minimum -- descent started at the truth does not move at
all -- the anchor bound says the information is 99.7% intact, and our
solver stops 0.44 above it at 4.9% accuracy. That is a pure optimiser
failure. Natural data is the opposite: descent from the truth falls a
further 0.167, so the truth is not even locally optimal, which is the
information-deficit signature the 0.291 bound predicted.
Steepest descent was brute-forcing all 32,640 candidate permutations
through the full energy every step, including a batched cube trace with
the triangle term active -- 203 seconds per descent, which is why the
gates were hopeless. The pairwise term needs one matrix product for the
whole table: swapping p,q changes the alignment sum by
2(C_pq + C_qp - C_pp - C_qq + 2 A_pq B_pq) with C = A @ B. Verified
against brute force to 1e-9 before use, and the fast descent reaches the
same optimum. 203s -> 0.79s.
Adds a matching benchmark with known-reachable answers and the solver
families never tried on these fields: Gromov-Wasserstein, entropic GW
with an annealed regulariser, BAPG.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
Three statistics failed the same way -- fields agreeing on the number
and disagreeing on recovery -- because each was invented by staring at
the fields rather than by asking what matching needs. The fourth asks
directly: declare half the scenes anchors, hand over their
correspondence, describe the rest by their field rows against the
anchors, and match one-to-one by Hungarian assignment. Seconds to
compute, and it upper-bounds blind recovery because blind recovery must
also discover the anchor correspondence.
Never violated across six fields spanning the full range of outcomes,
and it separates every case the refuted statistics collapsed:
synth full bound 0.989 blind 0.958 gap +0.03
synth noise 0.35 bound 0.984 blind 0.947 gap +0.04
synth omit size bound 0.997 blind 0.056 gap +0.94
synth rank 8 bound 0.930 blind 0.129 gap +0.80
Visual Genome bound 0.291 blind 0.000 gap +0.29
This corrects two claims from earlier today. Caption suppression does
not destroy information -- its bound is 0.997 -- it destroys blind
searchability, by making scenes interchangeable under permutation in a
way given anchors break. And Visual Genome's problem was never spectral
width: with the correspondence handed over, seven scenes in ten still
cannot be identified.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
The replacement gate proposed this morning is refuted by a control run
this afternoon. Independent noise lowers the correlation to 0.828
without narrowing the underlying signal and reaches the same measured
width as a field whose captions omit one factor -- 15.7 against 15.0 --
with recovery at 94.7% and 5.6%. A second instrument built specifically
to fix that, counting canonical directions that generalise to held-out
scenes, fails the same way and rates Visual Genome highest of the three
failing fields.
Established: the correlation does not govern recovery, in both
directions. Corpus overlap does control it -- the caption-suppression
ladder is monotone from 96.4% to 0.0%. Not established: any statistic
that predicts recovery cheaply. Documents corrected accordingly rather
than quietly rephrased; the refuted claim stood for four hours and is
recorded as such.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
The user-facing statement still carried the retired correlation
threshold. Adds the correction, the joint condition that replaces it,
the width diagnosis, and the closure of the shrink-N route.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
Four tests around today's additions. Two failed on first run and both
were worth having.
The degree decomposition left an O(1/n) residual on a field that is
purely additive: excluding the diagonal makes the two-way design
unbalanced, so one pass of row and column means does not remove a pure
degree effect. Swept to convergence instead. At N=256 the correction
moves the reported variance shares by under 0.001, so the refutation of
the hubness hypothesis stands unchanged -- but the instrument that
produced it now does what it claims.
The other failure was the test's own scale: two random 16-dimensional
subspaces of R^64 overlap above 0.7 by chance, which is why the real
measurements are made at N=256 where the null sits at 1.0.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
A controlled truncation refutes the project's central go/no-go rule.
Projecting the recovering synthetic fields to rank r holds the field
correlation at 0.902-0.929 while recovery moves 6.2% -> 12.9% -> 95.6%
across ranks 4, 8, 16. A field past the supposed 0.9 threshold recovers
13%, so correlation neither predicts nor forbids recovery and the width
of the shared spectrum is what moves it.
The gate becomes a joint condition on correlation and shared width,
measured by principal angles against a scene-shuffled null. Neither
suffices alone: 18 shared directions at 0.508 fails, 11 at 0.902 fails.
With the old gate retired, natural data was finally searched: 0.0000
against 0.0039 chance. The old verdict was right, its reasoning was not.
Also closes route D by measurement. rho_IT ~ sqrt(4 log N / N) rises as N
falls, and at N = 16 through 96 the deepest state a strong searcher
reaches is deeper than the truth in 3/3 replicates at every size.
Free gains: eigenvalue-weighted projection over a wide basis with
128-dim text vectors takes the correlation 0.656 -> 0.716 and shared
width 10 -> 16. Hubness refuted as an inflation hypothesis.
Moving the per-image segmentation eigendecomposition onto the GPU cut
batch time from 130s to 1.9s.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|
identifiability
Method: scene states are sets of part states; relation fields are built
within each modality and are invariant to how each side labels its own
features; the cross-modal bridge is a coupling searched under an energy
that is a closed-form functional of one matrix; solving is spectral
initialisation followed by exact local refinement.
Evidence: in a procedurally generated closed world, blind recovery of a
hidden image-caption correspondence reaches 95.3% at 256 scenes against
0.39% chance, and the recovered pairs transfer to 200 held-out scenes at
93.0% exact retrieval with random-pair and shuffled-image controls at or
near chance. Cross-modal value correspondence is derived from disjoint
corpora rather than declared. On Visual Genome the field correlation
reaches 0.656 against the 0.9 that polynomial recovery needs, with the
deficit attributed away from segmentation and discretisation.
Protocol: no image-text pair enters any objective, optimiser,
initialisation, or model selection; hidden pairs score orderings only.
Co-Authored-By: Claude <noreply@anthropic.com>
|