From 3b89612174f16da526188e640f6e67a0b4b22974 Mon Sep 17 00:00:00 2001 From: YurenHao0426 Date: Mon, 10 Aug 2026 04:26:07 -0500 Subject: results: aggregate four full state-bias EP seeds --- RAIN_EP_RELEASED_PROFILE.md | 18 ++++++++++++++++++ 1 file changed, 18 insertions(+) (limited to 'RAIN_EP_RELEASED_PROFILE.md') diff --git a/RAIN_EP_RELEASED_PROFILE.md b/RAIN_EP_RELEASED_PROFILE.md index 6849a9d..1d0b4d6 100644 --- a/RAIN_EP_RELEASED_PROFILE.md +++ b/RAIN_EP_RELEASED_PROFILE.md @@ -145,3 +145,21 @@ predeclared endpoint is final accuracy, so its temporarily higher intermediate or best accuracy is not substituted for this result. Two further paired seeds are required to narrow the comparison, but R3 itself is mixed/negative for an SDIL-over-intercept claim. + +## R4 four-seed aggregate + +R4 added seeds 1990 and 1991 under the identical full-data protocol. Their +final clean/raw/intercept/SDIL accuracies were 86.93/10.00/83.56/84.37% and +83.17/10.00/83.37/85.72%, respectively. Raw became nonfinite after three and +one epochs. SDIL beat the intercept tracker by 0.81 and 2.35 points on these +seeds. + +Across all four full runs, final clean/raw/intercept/SDIL accuracy averages +86.89/10.00/84.37/85.11%. The paired SDIL-minus-intercept differences are +-1.39, +1.20, +0.81, and +2.35 points, for a +0.74-point mean with a +1.56-point sample standard deviation. This is directionally positive but not +decisive: the gain changes sign and is small relative to seed variation. It +also compares against a costly tracker that receives a fresh exact local probe +every ten updates. The next experiment freezes equal small upfront probe +budgets and then removes further probes, testing whether state prediction +improves accuracy when repeated recalibration is unavailable. -- cgit v1.2.3