Blind Spot Intelligence — 134 runs

2026-03-02 → 2026-09-21 · 119 scored runs (110 unique dates) carry a full Alignment Tracker · 527 distinct candidate headings · rebuilt 2026-09-21

AGI Proximity
92%
59% → 92%  +33.0
Alignment Maturity
21.6%
30% → 21.6%  -8.4
Capability Gap
+70.4
+29 → +70.4  +41.4
Risk: CRITICAL
113
CRITICAL in 113 of 116 scored runs

1 · The scissors

The log's central claim, drawn once: capability proximity and alignment maturity have moved in opposite directions for six months.

AGI Proximity rose monotonically. Alignment Maturity fell from 30% to a floor near 13%, and has only turned up in the last three weeks. The shaded band is the gap.

2 · The gap itself

CRITICAL was first logged 2026-03-28 and has held for 107 of 110 scored runs. The gap has widened in every month except September.

3 · Five dimensions — only one went up

Monthly means. Interpretability is the sole dimension above its March value.

Interpretability Coverage +1.9 (2.0 → 3.9). Every dimension that measures control over behaviour fell: Oversight Robustness −2.7 to near zero, Institutional Coordination −1.4, Goal Specification −1.3, Course-Correction −0.9. We can see into these systems better than in March and steer them worse.

4 · September is the first month the trend reversed

Change in each dimension's monthly mean, August → September. The recovery is entirely institutional — coordination and course-correction — while oversight robustness stayed pinned at the floor.

5 · Candidate churn: 54% of ideas appear once and never again

514 distinct candidate headings across 128 runs. 279 appeared in exactly one run. Only 21 survived more than 30 days; only 2 survived more than 90 — Latent Recurrent Depth (185d) and Active Inference / Free Energy Principle (147d).

6 · Score inflation — the log grading itself

Mean blind-spot score per month, and the share of scores at a perfect 10/10.

Through July the mean score sat at 7.8–8.5 with 4–13% at 10/10. In August the mean jumped to 9.38 and 59% of all candidate scores were 10/10; September, 46%. Simultaneously the log tracked far fewer candidates (742 scores in April → 50 in September). Either the world became dramatically more blind-spot-dense, or the scale collapsed. The second is more likely, and it is a measurement problem in the log's own instrument — the same critique run 128 levelled at the field.

7 · The two survivors

Score trajectories for the eight longest-tracked candidates.

8 · Prediction Timeline — how the nine estimates have moved

Left: each milestone's estimated arrival date over time (dotted diagonal = "today"). Right: every move, in order. Hover either side to link them. CAP capability-driven · ALN alignment-driven · earlier · later · confidence up · pressure logged, no move

    Reconstructed from the log's movement narratives and verified against the master table; every chain connects to the current state. Two things to see. First, the two capability-benchmark milestones (M3, M4) have both drifted later since May — in the same six months AGI Proximity went 57% → 92%. The log's capability score and its dated forecasts disagree about direction. Second, every move that came earlier in September was alignment-driven: a first-party lab statement or a disclosed incident, not a benchmark. The Hugging Face intrusion (contained 07-16, disclosed 07-21) appears twice as pressure — the log misclassified it for four runs, then it became M2's second E — and is the direct antecedent of the 09-13 confidence move.

    #MilestoneEstimateConf.MovesLast movedSequence
    1Fully Automated AI Research2026-H2High2 + conf2026-09-01▲▼◆
    2First Spontaneous Agentic Misalignment2026-Q3Med2 + conf2026-09-13▲▲◆
    3ARC-AGI-3 >15% RHAE (non-LLM)2028Med52026-09-05▲▼▼▼▼
    4General Superhuman Cognitive Labor2027-Q1Med62026-07-27▲▲▲▼▼▼
    5Meaningful RSI Threshold2027-Q4Low12026-09-13
    6World Models Dominate2029Low12026-09-03
    7Broadly Accepted AGI2029-2030Low0seed
    8Capability Explosion / Singularity2032-2035Low0seed
    9ASI2035+Low0seed

    Data-integrity note. The original seed values for M1–M4 (set 2026-03-28) are not recoverable from the memory file: the master table's history chains are truncated mid-sentence and the early run sections never recorded timeline paragraphs. The chart starts each milestone at its first fully-stated value (M3 2026-03-29, M1 2026-04-07, M2/M4 2026-04-08). Six early moves (runs 18–29) are known to have happened but their from/to values are lost.