Method, audit, and what is wrong with this
Everything that would make you trust the labels less is here, and the whole thing is on GitHub — every label, every rationale, and the scripts that rebuild the database from Convergent’s own export.
- No human reviewed any label. All 103 gaps were labelled, audited and written up in 181 minutes of agent time.
- A second pass relabelled all 103 blind and disagreed often enough to be worth publishing. The rates are below.
- Three cold reviews then read the artifact as one of you would, and found real errors. Those are below too.
- One finding was withdrawn after it failed to replicate.
The result that did not replicate, in full
This was on the front page and is no longer, because it is a fact about this analysis rather than about the map. It is here in full because withdrawing a finding quietly is worse than never publishing it.
The first pass produced a clean gradient: the share of gaps each kind of work blocks where the AI for it already works ran from 60% for reading and synthesis down to 0% for physical build. A second pass relabelled all 103 gaps blind, against a revised taxonomy, by labelers who never saw the first set, with predictions registered in a commit beforehand. It did not reproduce that result.
| Kind of work | First pass | Independent relabel | Moved |
|---|---|---|---|
| Reading and synthesis | 60% 6/10 | 50% 5/10 | -10 points |
| Coordination and institutions | 7% 1/15 | 0% 0/15 | -7 points |
| Measurement and sensing | 20% 3/15 | 47% 9/19 | +27 points |
| Prediction and modeling | 30% 6/20 | 14% 3/22 | -16 points |
| Design search | 19% 4/21 | 13% 2/16 | -6 points |
| Running experiments | 40% 2/5 | 56% 5/9 | +16 points |
| Physical build | 0% 0/17 | 13% 1/8 | +13 points |
| Real-time control | — | 25% 1/4 | new category |
The ordering inverts at the top. Coordination and institutions go from last place to first. Physical build is no longer zero. The gradient is withdrawn.
The cause was a definitional hole I left open. Does “working now” mean the capability exists, or that applying it would move this gap? For technical categories those coincide. For institutional ones they come apart completely: convening a standards body is available this afternoon, and getting universal DNA-synthesis screening adopted is not. I labelled institutional gaps on efficacy and the relabelers read availability.
A third pass repaired maturity against the sharper definition — applying it would move this gap — which is why coordination and institutional now has nothing in the working-now column. That is a definition being fixed, not a result being found, and it is the reason the argument page makes no claim about a gradient.
What I would not rely on
- Maturity is the least reliable label here. The two independent passes agreed on the kind of work for 77 of 103 gaps and on maturity for only 63. A third pass then repaired it against a sharper definition. Treat a single gap’s maturity as a judgment, and the distribution as the thing worth reading.
- Every label is an AI judgment, not expert consensus, and no human has reviewed any of them.
- The argument leans on the category I was least sure of. Gaps whose primary blocker is coordination and institutional carry a confidence flag on some dimension 11 times out of 15.
- The kinds of work fuse two questions. What kind of work is in the way, and how mature the AI for it is, are separate facts sharing one axis. Every kind has an AI analogue, robotics included, so the axis is really about maturity and should probably be split in two.
- The tier confidence flag is close to a synonym for “proxy only”. Of the 24 flagged tiers, 19 come from one blanket rule and 5 are independent judgments.
- Outcomes are a text field on a gap. There are more outcomes than gaps, and one capability unlocks outcomes across several fields. Modelling them properly is a schema change.
- The indicators are a sample of eight, chosen across tiers. Nothing about them supports a claim about the other 95.
- Capability edges are untyped upstream, so the chains reconstruct link semantics by hand. That is why there are two of them.
The blind audit: how it was run and what it found
A second labeller, which never saw the first set, relabelled a stratified sample of 36 of the 103 gaps. The sample oversamples the rare tiers on purpose, because that is where the taxonomy is hardest, so the raw rate is biased upward and the population-weighted figure is the one that means anything.
| Stratum (original tier) | In population | Sampled | Agreed | Disagreement |
|---|---|---|---|---|
| Directly measurable | 72 | 15 | 15 | 0% |
| Proxy only | 19 | 9 | 2 | 78% |
| Verification contested | 11 | 11 | 10 | 9% |
| Counterfactual required | 1 | 1 | 1 | 0% |
| Raw sample disagreement (28 of 36 agreed) | 22% | |||
| Weighted to the population (each stratum’s rate by its share of the 103) | 15% | |||
One of my four tiers does not work. “Proxy only” ran 78% disagreement against 0% for directly measurable, and every auditor independently reported it was the nearest alternative and almost never the winner. If you adopt a measurability attribute, three tiers would work better than four. On the kind of work, raw sample disagreement was 33%, and the audit found three gap types the seven-value list handles badly: closed-loop control of a physical system, gaps where AI is the object rather than the instrument, and composite gaps that would need different values for different sub-problems. The first became an eighth category, real-time control of physical systems, which four gaps took as primary in the relabel.
Every disagreement, both dimensions
| Gap | Dimension | First labeller | Auditor |
|---|---|---|---|
| Uncertainty and Noise in the Science of Room-Temperature Superconductivity | tier | Verification contested | Directly measurable |
| Current “Model Systems” for Brain Function are Not Representative of the Real Human Brain | tier | Proxy only | Verification contested |
| AI Could Be Misused | tier | Proxy only | Verification contested |
| Major Planetary Science and Astrobiology Missions Are Not Realized by Existing Government Space Agencies | tier | Proxy only | Directly measurable |
| Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation | tier | Proxy only | Verification contested |
| Underdevelopment of Modern Tools in the Social Sciences | tier | Proxy only | Directly measurable |
| Lack of a Dedicated Field for Planetary Terraforming | tier | Proxy only | Counterfactual required |
| We Can Learn More from Nature’s Biological Designs | tier | Proxy only | Verification contested |
| Uncertainty and Noise in the Science of Room-Temperature Superconductivity | kind of work | Coordination and institutions | Prediction and modeling |
| Our Platforms for Civic Engagement and Democratic Decision-Making Don’t Take Advantage of 21st Century Scalable Technology | kind of work | Coordination and institutions | Reading and synthesis |
| Robust and Compact Plasma Confinement for Fusion is Still Not Solved | kind of work | Prediction and modeling | Design search |
| We Have a Limited Ability to Acquire, Concentrate and Substitute Chemical Elements in Processes | kind of work | Design search | Running experiments |
| Biological Life is Our Only Working Example of Complex Evolved Computation | kind of work | Prediction and modeling | Design search |
| Intervening in Earth Systems at Scale is Largely Untested | kind of work | Physical build | Prediction and modeling |
| AI Could Be Misused | kind of work | Coordination and institutions | Reading and synthesis |
| Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation | kind of work | Reading and synthesis | Coordination and institutions |
| Quantum Gravity is Experimentally Hard to Constrain | kind of work | Physical build | Design search |
| Difficulty Delivering Physical Probes for Imaging into Living Cells | kind of work | Design search | Measurement and sensing |
| Synthetic Biology Platforms Are Over-Reliant on Evolved Cells That We Don’t Fully Understand or Control | kind of work | Design search | Running experiments |
| Poor Scalability of Bioreactors Limits Biomanufacturing | kind of work | Physical build | Prediction and modeling |
Progress indicators, built and not proposed
Eight gaps across all four tiers. Every value was read off a page that was actually fetched. Four verify against Crossref or arXiv; two are reachable but not scholarly-verified and are recorded that way.
| Gap | Tier | Quantity | Current | Source |
|---|---|---|---|---|
| Silicon-Based Electronics Face Fundamental Limits in Dimensional Scaling | Directly measurable | Contacted gate pitch, most recent publicly disclosed volume-production node | 45 nm | IEDM 2022 – TSMC 3nm |
| Searching Through the Vast, Underexplored Space of Materials is Slow and Expensive | Directly measurable | Novel inorganic compounds experimentally realised per day of autonomous laboratory operation | 2.4 compounds per day | An autonomous laboratory for the accelerated synthesis of novel materials |
| Frontier Telescopes Are Expensive and Take Decades to Build | Directly measurable | Elapsed time from first concept study to launch, flagship space observatory | 32 years | Mission Timeline — James Webb Space Telescope |
| Quantum Gravity is Experimentally Hard to Constrain | Verification contested | Highest benchmarked quantum-gravity figure-of-merit eta across mechanical quantum-control platforms | 1.1e-7 dimensionless (eta) | Agafonova, Rossello, Mekonnen & Hosten, 'One-milligram torsional pendulum toward experiments at the quantum-gravity interface', arXiv:2408.09445v4 (20 December 2025), published as Communications Physics 9, 80 (30 January 2026). Table S1 of the supplement benchmarks eta across thirteen platforms. |
| Fraud in the Scientific Literature | Proxy only | Retractions per 10,000 published articles, averaged over 2000-2024 | 11.82 retractions per 10,000 articles | Zhou, Lou, Shen & Li, 'Mapping Academic Integrity: Global Retraction Trends Explored through a Topic Lens', arXiv:2511.21176v2 (6 July 2026). Preprint, not peer reviewed. |
| Clinical Trials Are Poorly Optimized for Evidence Gathering | Directly measurable | Median estimated cost of a pivotal clinical trial supporting a new FDA approval | 19.0 million USD | Estimated Costs of Pivotal Trials for Novel Therapeutic Agents Approved by the US Food and Drug Administration, 2015-2016 |
| A Limited Set of Rigid Organizational Structures for Organizing and Funding Research Constrains the Forms of R&D That Get Done | Counterfactual required | Progress toward less rigid organizational and funding structures for research | none found | — |
| Most Brain Circuitry is Still Invisible | Directly measurable | Volume of brain tissue reconstructed at synaptic resolution in a single dataset | 1.0 mm³ | Functional connectomics spanning multiple areas of mouse visual cortex |
A Limited Set of Rigid Organizational Structures for Organizing and Funding Research Constrains the Forms of R&D That Get Done. NULL CORROBORATED by Gate A, and this is now the strongest row in the sample. A second searcher, working from the gap description alone and without sight of the queries below, ran eight independent searches of its own wording and also found no maintained series measuring which forms of R&D get done. Two searchers failing with different wording is a materially different claim from one searcher failing, and it is the reason to trust this row. The original six searches (recorded in research-log/searches/phase-3.json, phase 3, this gap — the search_log database table is empty, so that file is the provenance): researcher time spent on administration; metrics for diversity of research funding mechanisms; counterfactual measurement of funding structure against discoveries; outcome metrics for new organizational forms including FROs; indicators of institutional innovation in science funding; and randomised experiments in grant allocation. Near-misses, all failing for stateable reasons. (1) Administrative burden: US federally funded principal investigators report spending roughly 42% of their research time on administration. The gate objected that the gap's own sentence opens 'scientists are spending a lot of time not doing science', so this measures the first clause and was rejected against the second. The objection is fair and the rejection still stands, but on narrower grounds than before: the 42% is a real measure of one symptom, and cutting it to 20% would be a genuine gain that left the gap's actual claim — which forms of R&D get done — untouched. It is the wrong quantity, not a trivial one, and the FDP survey behind it last ran in 2018. (2) NIH's median age at first R01-equivalent award, an official series updated annually, has held near 42 across FY21-25 (mean 43-44). It is maintained, current and genuinely about rigidity, which makes it a stronger near-miss than either of the two originally listed — and it still measures who gets funded rather than which forms of research become possible. (3) The randomised evidence on funding allocation, including grant lotteries, is the right instrument but is a set of individual studies, not a series, and its own conclusion is that counterfactual assessment of funding structures is usually impossible because researchers have alternative sources. This is what Counterfactual required means: the quantity of interest is the research a different structure would have produced and this one did not, and no observation of the world we are in contains it. The null is not a search failure. It is the tier being correct.
Distributions across all 103 gaps
Measurability tier
Kind of work in the way
Tier against kind of work
| Kind of work | Directly measurable | Proxy only | Verification contested | Counterfactual required | Total |
|---|---|---|---|---|---|
| Coordination and institutions | 4 | 8 | 2 | 1 | 15 |
| Design search | 15 | · | 1 | · | 16 |
| Measurement and sensing | 16 | 1 | 2 | · | 19 |
| Physical build | 8 | · | · | · | 8 |
| Prediction and modeling | 14 | 5 | 3 | · | 22 |
| Reading and synthesis | 4 | 4 | 2 | · | 10 |
| Real-time control | 4 | · | · | · | 4 |
| Running experiments | 7 | 1 | 1 | · | 9 |
Tier against field
| Field | Directly measurable | Proxy only | Verification contested | Counterfactual required | Total |
|---|---|---|---|---|---|
| Astrophysics | 3 | 1 | · | · | 4 |
| Biophysics | 5 | · | 2 | · | 7 |
| Biosecurity | 3 | 2 | · | · | 5 |
| Cellular and Molecular Biology | 3 | · | · | · | 3 |
| Chemistry | 7 | · | · | · | 7 |
| Computation | 4 | 1 | 3 | · | 8 |
| Ecology | 1 | 2 | 1 | · | 4 |
| Geophysics and Climate | 3 | 2 | 1 | · | 6 |
| Global Health | 3 | · | · | · | 3 |
| Immunology | 2 | · | · | · | 2 |
| Materials Science | 4 | · | · | · | 4 |
| Mechanical Engineering | 7 | · | · | · | 7 |
| Metascience | 2 | 2 | · | 1 | 5 |
| Nanoscale Fabrication | 3 | · | · | · | 3 |
| Neuroscience | 3 | 1 | · | · | 4 |
| Physics | 5 | · | 3 | · | 8 |
| Physiology and Medicine | 5 | 1 | · | · | 6 |
| Social Science | 2 | 6 | 1 | · | 9 |
| Space Engineering | 2 | 1 | · | · | 3 |
| Synthetic Biology | 5 | · | · | · | 5 |
What it cost to build
| Phase | Kind | Units | Elapsed | What happened |
|---|---|---|---|---|
| phase-0 | agent | — | 46 min | Repo scaffold, schema, importer, additive guardrail, integrity report, search client, baseline snapshot. |
| phase-1 | agent | 103 | 11 min | Full-coverage labeling of all 103 gaps across 20 field batches: outcome, AI type and maturity, measurability tier. |
| phase-2 | agent | 36 | 19 min | Stratified blind audit across 3 independent auditors, adjudication, findings report. |
| phase-2b | agent | 103 | 55 min | Taxonomy revision (8th type, frame dimension), full independent blind relabel of all 103 gaps by 5 labelers, mechanical adjudication, withdrawal of the maturity gradient. |
| phase-3 | agent | 8 | 9 min | progress indicators: 8 rows across 4 tiers, 2 honest nulls, 29 logged searches. Local-session setup before this point (rebuild, network verification, ingest plumbing) is not counted. |
| phase-4 | agent | 4 | 6 min | new gaps: 4 proposed, 1 candidate dropped; dedup over the full export plus funding checks. |
| phase-5 | agent | 2 | 7 min | two critical paths, 15 steps, plus the cross-field intersection. Expectations committed before the analysis in a separate commit. |
| phase-6 | agent | 1 | 27 min | artifact (Next.js static export), CSV keyed on their id and slug, findings summary, cover note draft; found and fixed the adjudication rebuild defect. |
| revision-1 | agent | — | 73 min | rebuilt around the argument in Convergent's voice, after a blind cold review |
| revision-2 | agent | — | 50 min | second cold review: replotted the chart, fixed three dead links, stated three holes |
| revision-3 | agent | — | 41 min | cut page one to a fifth, split into five pages, rebuilt the map on their components |
| revision-4 | agent | — | 22 min | new opening, all 103 outcome sentences rewritten, indicators page |
| revision-5 | agent | — | 33 min | third cold review: two reasoning errors corrected, outcomes deticked, attributes page |
| revision-6 | agent | — | open | robotics as the AI analogue for physical build, time ledger, drafting cost corrected |
| all | human-review | — | 244 min | David's own time reviewing, directing and rewriting, over the same window as the revision cycle. His estimate rather than an instrumented figure, and roughly equal to the agent's. The build phases above had no human review at all. |
Agent time and human time are tracked separately, because a single blended number would be the first thing worth objecting to. Phase 0’s start was never instrumented, so it counts as zero and the 181 minutes for the build is a lower bound.
Calls that could have gone the other way
31 of them, each with the runner-up and what would reverse it. The runner-up was usually the more flattering option.