The label on every gap
What kind of work stands in the way of each gap, and whether AI reaches it. What the label is, why it might be worth having, and where it breaks.
Two further attributes were built and are not proposed. A one-sentence outcome on every gap, and a progress indicator on eight of them. Both are in the CSV and the JSON for anyone who wants them. They are left out here because the critical paths do the same jobs better: a chain has to state the axis it runs on, which is what the outcome was for, and it carries a sourced quantity on every step, which is what the indicator was for.
1. What kind of work is in the way
Eight values, naming kinds of work and not kinds of model. This column answers one question only: what stands between here and the gap closing. It is not a claim that AI does that work — that is the separate question in the last column.
Six of the eight were originally named after the AI that would do the work (“LLM reasoning and synthesis”) and two after the work itself (“Physical build and manipulation”). Read together, that made the institutional category look out of place when it was the one naming the thing consistently. All eight are now named for the work. The stored labels and every recorded judgment are unchanged; only the words a reader sees moved.
| Kind of work | What it means | Gaps | Where the AI for it stands |
|---|---|---|---|
| Reading and synthesis | Reading, summarizing, connecting, proposing. Work whose product is text or an argument. | 10 | Language models. Working today. |
| Prediction and modeling | Learning a fast approximation of something slow to compute or measure, then using it in place of the slow thing. | 22 | Learned surrogates. Working today in several fields. |
| Design search | Searching a large space of candidate designs against a stated objective. | 16 | Generative design and search. Working today for proteins and materials. |
| Measurement and sensing | Getting a usable measurement out of a noisy or indirect one. | 19 | Learned reconstruction and denoising. Working today. |
| Running experiments | Choosing the next experiment and running it with nobody in the loop. | 9 | Self-driving labs. Early, and real. |
| Real-time control | Closed-loop sense, decide and actuate on hardware that already exists, at machine timescales. Added after a blind audit found it had no home in the original seven. | 4 | Learned control. Working today for plasma and adaptive optics. |
| Physical build | Fabricating, assembling, or handling matter. | 8 | Robotics, and it is moving fast. One of these gaps has a capability that works today; the rest are two-to-five years or speculative. |
| Coordination and institutions | Approval, funding, agreement, incentives, and who counts what. | 15 | Hardest of the eight, and not empty: matching, scheduling, drafting and forecasting all apply. |
Each gap gets exactly one primary and any number of secondaries. Alongside it sits a maturity: working now, two-to-five years, or speculative, describing the relevant AI capability rather than the gap.
Where it breaks, and this is the attribute I would most like torn apart.
A full independent relabel of all 103 gaps put type disagreement at 25%. It also confirmed the eighth category was worth adding: four gaps took real-time control as their primary. Two problems the audit found are still open. Gaps where AI is the object rather than the instrument now carry a separate frame flag instead of a type. Composite gaps, which bundle sub-problems needing different values, are still recorded under one label.
Maturity is the weakest thing measured here. The two passes agreed on the kind of work for 77 of 103 gaps and on maturity for only 63, and the disagreements moved overwhelmingly in one direction. The working-now gradient the first pass produced did not replicate and has been withdrawn.
Fusing “what kind of blocker” with “how mature is the AI for it” into one axis is probably the underlying mistake. Two fields would be cleaner than one, and would make the robotics trajectory legible instead of hiding it inside a maturity label.
2. The measurability tier
Whether the gap has something you could actually watch. Your own roadmapping criterion asks whether success is unambiguously measurable, and applying it to all 103 gaps turns out to sort them sharply.
| Tier | What it means | Gaps |
|---|---|---|
| Directly measurable | An observable quantity exists and everyone agrees which direction is an improvement. Elapsed years, cost per trial, cubic millimeters reconstructed. | 72 |
| Proxy only | You can measure inputs or side effects but not the thing itself. This tier failed its own audit at 78% disagreement and I would drop it. | 19 |
| Verification contested | A candidate observable exists and there is no agreement that moving it settles anything. Quantum gravity is the clean case. | 11 |
| Counterfactual required | The quantity of interest is something that did not happen. No observation of the world you are in contains it. | 1 |
Where it breaks, and I would ship three tiers rather than four. “Proxy only” ran 78% disagreement in the blind audit against 0% for directly measurable. Every auditor independently reported it was the nearest alternative and almost never the winner. A category that two careful readers apply differently four times in five is not a category.
Confidence flags
Every label carries confident or guess. A run that produced no guesses would not be a careful run.
One caveat on reading them. Because every proxy-only tier was downgraded as a class, 19 of the 24 flagged tiers come from that one rule and only 5 are independent judgments. The tier flag is closer to a synonym for proxy-only than to a measure of my uncertainty. The full audit