From ad8819a63e632e0ba7364c56127e1f31b832530f Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Sat, 1 Aug 2026 16:37:01 -0500 Subject: Close the backbone row: the prediction held DINOv2-large, twice the depth of base and wider, leaves the shared count at exactly 15 and moves the correlation by less than the segmentation noise floor. Four vision-side interventions now raise the vision field's own rank and leave the intersection alone. The cheap tier of the register is exhausted. What remains is a corpus chosen for naturally wide overlap, promoted to the central bet, and the unbalanced formulation, promoted to the critical path since such a corpus is unlikely to arrive in bijection. Co-Authored-By: Claude --- CANDIDATE_REGISTER.md | 49 +++++++++++++++++++++++++++++-------------------- 1 file changed, 29 insertions(+), 20 deletions(-) (limited to 'CANDIDATE_REGISTER.md') diff --git a/CANDIDATE_REGISTER.md b/CANDIDATE_REGISTER.md index eff860e..e4aa235 100644 --- a/CANDIDATE_REGISTER.md +++ b/CANDIDATE_REGISTER.md @@ -14,7 +14,7 @@ Status is one of: **done** (measured, number recorded), **running**, **queued** (specified, not started), **open** (idea, not specified), **closed** (measured and eliminated). -Current position: correlation 0.716, shared directions 15, against a +Current position: correlation 0.731, shared directions 15, against a recovering synthetic reference at 0.928 and 26. ## A. Widen the shared spectrum @@ -32,21 +32,24 @@ correlation. | Annotated boxes instead of segmentation (oracle) | low | closed: shared 15, buys one direction | | More scenes (256 to 1024) | low | closed: shared 23 to 19 at matched width | | Augmentation orbits, four views | medium | closed: no gain, ρ 0.6559 either way | -| Larger vision backbone (DINOv2-large) | medium | running: predicted useless, see below | -| Third-order moments in the set kernel | low | queued | -| Corpora with naturally wide overlap | high | queued | +| Larger vision backbone (DINOv2-large) | medium | closed: shared 15 either way, prediction held | +| Third-order moments in the set kernel | low | closed: vision rank +7, shared +1 | +| Structured phrase encoding: head and modifiers apart | low | done: ρ 0.716 to 0.725, shared unchanged | +| Suppressing a caption factor (synthetic, controlled) | low | done: shared 20 to 15, recovery 96.4% to 5.6% | +| Corpora with naturally wide overlap | high | queued -- now the central bet | | Intervention responses: occlude, re-encode, use the change | high | queued | | Corpus-trained text encoder instead of PPMI vectors | medium | queued | -| Structured phrase encoding: head and modifiers apart | low | queued | | Nonlinear content projection (kernel discriminant) | low | closed: 0.62 against 0.69 linear | -The backbone row is a live prediction test rather than a hope. The shared count -does not move when segmentation is replaced by ground-truth boxes, when -segments are nearly tripled, or when the vision field's own effective rank -doubles, so a larger backbone is predicted to leave it unchanged. The run is -finishing anyway because the project's rule is to measure the bound rather than -infer it, and a cheap refutation of my own prediction is worth more than the -GPU time. +The backbone row was a pre-registered prediction and it held: DINOv2-large, +twice the depth of base, leaves the shared count at exactly 15 and moves the +correlation less than the segmentation noise floor. Together with segments, +oracle boxes and third-order moments, that is four vision-side interventions +that raise the vision field's own rank and leave the intersection alone. + +Current position after today: correlation 0.731, shared 15. The gains came +from the projection, the text vectors and the phrase encoding; nothing on the +vision side moved either number. **Corpora with naturally wide overlap is the row today's result argues for and the one never attempted.** The synthetic world reaches 26 shared directions @@ -142,11 +145,17 @@ the only home for selection-based ideas. ## Order of attempt -Third-order moments first: it is the cheapest untried thing that raises field -rank by construction, since a set kernel carrying third moments spans more -directions than one carrying two. Then structured phrase encoding and the -corpus-trained text encoder, because the text side is the one that moved the -shared count at all. Then the two rows worth pre-committing to regardless of -cost — intervention responses, and a corpus chosen for naturally wide overlap — -the second of which is now the register's central bet rather than one candidate -among many. +The cheap tier is exhausted. Third-order moments, structured phrase encoding +and the larger backbone were the last three untried low-cost rows and all +three are now measured: the first two raise the correlation by 0.01 each and +the third by nothing, and none of them moves the shared count. Five separate +representational interventions have now left the intersection at 15 ± 1. + +What remains is not cheaper work but different work. **A corpus chosen for +naturally wide overlap is the register's central bet**, promoted from one +high-cost candidate among several, because the caption-suppression experiment +showed corpus overlap to be what sets the width and no encoder-side change to +touch it. Intervention responses stay pre-committed as the one signal that does +not pass through the relation field at all. The unbalanced formulation in F is +now on the critical path rather than a future tidy-up, since a corpus with wide +overlap is unlikely to arrive in bijection. -- cgit v1.2.3