david@sendorai.comLinkedInAn independent contribution to Convergent Research’s Fundamental Development Gap Map v1.0. Not affiliated with, endorsed by, or published by Convergent Research.

Method, audit, and what is wrong with this

AI writtenThis page was written by Claude, and so was everything it describes.

Everything that would make you trust the labels less is here, and the whole thing is on GitHub — every label, every rationale, and the scripts that rebuild the database from Convergent’s own export.

  • No human reviewed any label. All 103 gaps were labelled, audited and written up in 181 minutes of agent time.
  • A second pass relabelled all 103 blind and disagreed often enough to be worth publishing. The rates are below.
  • Three cold reviews then read the artifact as one of you would, and found real errors. Those are below too.
  • One finding was withdrawn after it failed to replicate.
The result that did not replicate, in full

This was on the front page and is no longer, because it is a fact about this analysis rather than about the map. It is here in full because withdrawing a finding quietly is worse than never publishing it.

The first pass produced a clean gradient: the share of gaps each kind of work blocks where the AI for it already works ran from 60% for reading and synthesis down to 0% for physical build. A second pass relabelled all 103 gaps blind, against a revised taxonomy, by labelers who never saw the first set, with predictions registered in a commit beforehand. It did not reproduce that result.

Kind of workFirst passIndependent relabelMoved
Reading and synthesis60% 6/1050% 5/10-10 points
Coordination and institutions7% 1/150% 0/15-7 points
Measurement and sensing20% 3/1547% 9/19+27 points
Prediction and modeling30% 6/2014% 3/22-16 points
Design search19% 4/2113% 2/16-6 points
Running experiments40% 2/556% 5/9+16 points
Physical build0% 0/1713% 1/8+13 points
Real-time control25% 1/4new category
Share of each kind of work whose gaps have an AI capability that works today, in both passes. The two agreed on the kind of work for 77 of 103 gaps and on maturity for only 63.

The ordering inverts at the top. Coordination and institutions go from last place to first. Physical build is no longer zero. The gradient is withdrawn.

The cause was a definitional hole I left open. Does “working now” mean the capability exists, or that applying it would move this gap? For technical categories those coincide. For institutional ones they come apart completely: convening a standards body is available this afternoon, and getting universal DNA-synthesis screening adopted is not. I labelled institutional gaps on efficacy and the relabelers read availability.

A third pass repaired maturity against the sharper definition — applying it would move this gap — which is why coordination and institutional now has nothing in the working-now column. That is a definition being fixed, not a result being found, and it is the reason the argument page makes no claim about a gradient.

What I would not rely on

  • Maturity is the least reliable label here. The two independent passes agreed on the kind of work for 77 of 103 gaps and on maturity for only 63. A third pass then repaired it against a sharper definition. Treat a single gap’s maturity as a judgment, and the distribution as the thing worth reading.
  • Every label is an AI judgment, not expert consensus, and no human has reviewed any of them.
  • The argument leans on the category I was least sure of. Gaps whose primary blocker is coordination and institutional carry a confidence flag on some dimension 11 times out of 15.
  • The kinds of work fuse two questions. What kind of work is in the way, and how mature the AI for it is, are separate facts sharing one axis. Every kind has an AI analogue, robotics included, so the axis is really about maturity and should probably be split in two.
  • The tier confidence flag is close to a synonym for “proxy only”. Of the 24 flagged tiers, 19 come from one blanket rule and 5 are independent judgments.
  • Outcomes are a text field on a gap. There are more outcomes than gaps, and one capability unlocks outcomes across several fields. Modelling them properly is a schema change.
  • The indicators are a sample of eight, chosen across tiers. Nothing about them supports a claim about the other 95.
  • Capability edges are untyped upstream, so the chains reconstruct link semantics by hand. That is why there are two of them.
The blind audit: how it was run and what it found

A second labeller, which never saw the first set, relabelled a stratified sample of 36 of the 103 gaps. The sample oversamples the rare tiers on purpose, because that is where the taxonomy is hardest, so the raw rate is biased upward and the population-weighted figure is the one that means anything.

Stratum (original tier)In populationSampledAgreedDisagreement
Directly measurable7215150%
Proxy only199278%
Verification contested1111109%
Counterfactual required1110%
Raw sample disagreement (28 of 36 agreed)22%
Weighted to the population (each stratum’s rate by its share of the 103)15%

One of my four tiers does not work. “Proxy only” ran 78% disagreement against 0% for directly measurable, and every auditor independently reported it was the nearest alternative and almost never the winner. If you adopt a measurability attribute, three tiers would work better than four. On the kind of work, raw sample disagreement was 33%, and the audit found three gap types the seven-value list handles badly: closed-loop control of a physical system, gaps where AI is the object rather than the instrument, and composite gaps that would need different values for different sub-problems. The first became an eighth category, real-time control of physical systems, which four gaps took as primary in the relabel.

Every disagreement, both dimensions
GapDimensionFirst labellerAuditor
Uncertainty and Noise in the Science of Room-Temperature SuperconductivitytierVerification contestedDirectly measurable
Current “Model Systems” for Brain Function are Not Representative of the Real Human BraintierProxy onlyVerification contested
AI Could Be MisusedtierProxy onlyVerification contested
Major Planetary Science and Astrobiology Missions Are Not Realized by Existing Government Space AgenciestierProxy onlyDirectly measurable
Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation tierProxy onlyVerification contested
Underdevelopment of Modern Tools in the Social SciencestierProxy onlyDirectly measurable
Lack of a Dedicated Field for Planetary TerraformingtierProxy onlyCounterfactual required
We Can Learn More from Nature’s Biological DesignstierProxy onlyVerification contested
Uncertainty and Noise in the Science of Room-Temperature Superconductivitykind of workCoordination and institutionsPrediction and modeling
Our Platforms for Civic Engagement and Democratic Decision-Making Don’t Take Advantage of 21st Century Scalable Technologykind of workCoordination and institutionsReading and synthesis
Robust and Compact Plasma Confinement for Fusion is Still Not Solvedkind of workPrediction and modelingDesign search
We Have a Limited Ability to Acquire, Concentrate and Substitute Chemical Elements in Processeskind of workDesign searchRunning experiments
Biological Life is Our Only Working Example of Complex Evolved Computationkind of workPrediction and modelingDesign search
Intervening in Earth Systems at Scale is Largely Untestedkind of workPhysical buildPrediction and modeling
AI Could Be Misusedkind of workCoordination and institutionsReading and synthesis
Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation kind of workReading and synthesisCoordination and institutions
Quantum Gravity is Experimentally Hard to Constrain kind of workPhysical buildDesign search
Difficulty Delivering Physical Probes for Imaging into Living Cellskind of workDesign searchMeasurement and sensing
Synthetic Biology Platforms Are Over-Reliant on Evolved Cells That We Don’t Fully Understand or Controlkind of workDesign searchRunning experiments
Poor Scalability of Bioreactors Limits Biomanufacturingkind of workPhysical buildPrediction and modeling
Progress indicators, built and not proposed

Eight gaps across all four tiers. Every value was read off a page that was actually fetched. Four verify against Crossref or arXiv; two are reachable but not scholarly-verified and are recorded that way.

GapTierQuantityCurrentSource
Silicon-Based Electronics Face Fundamental Limits in Dimensional ScalingDirectly measurableContacted gate pitch, most recent publicly disclosed volume-production node45 nmIEDM 2022 – TSMC 3nm
Searching Through the Vast, Underexplored Space of Materials is Slow and ExpensiveDirectly measurableNovel inorganic compounds experimentally realised per day of autonomous laboratory operation2.4 compounds per dayAn autonomous laboratory for the accelerated synthesis of novel materials
Frontier Telescopes Are Expensive and Take Decades to BuildDirectly measurableElapsed time from first concept study to launch, flagship space observatory32 yearsMission Timeline — James Webb Space Telescope
Quantum Gravity is Experimentally Hard to Constrain Verification contestedHighest benchmarked quantum-gravity figure-of-merit eta across mechanical quantum-control platforms1.1e-7 dimensionless (eta)Agafonova, Rossello, Mekonnen & Hosten, 'One-milligram torsional pendulum toward experiments at the quantum-gravity interface', arXiv:2408.09445v4 (20 December 2025), published as Communications Physics 9, 80 (30 January 2026). Table S1 of the supplement benchmarks eta across thirteen platforms.
Fraud in the Scientific LiteratureProxy onlyRetractions per 10,000 published articles, averaged over 2000-202411.82 retractions per 10,000 articlesZhou, Lou, Shen & Li, 'Mapping Academic Integrity: Global Retraction Trends Explored through a Topic Lens', arXiv:2511.21176v2 (6 July 2026). Preprint, not peer reviewed.
Clinical Trials Are Poorly Optimized for Evidence GatheringDirectly measurableMedian estimated cost of a pivotal clinical trial supporting a new FDA approval19.0 million USDEstimated Costs of Pivotal Trials for Novel Therapeutic Agents Approved by the US Food and Drug Administration, 2015-2016
A Limited Set of Rigid Organizational Structures for Organizing and Funding Research Constrains the Forms of R&D That Get DoneCounterfactual requiredProgress toward less rigid organizational and funding structures for researchnone found
Most Brain Circuitry is Still InvisibleDirectly measurableVolume of brain tissue reconstructed at synaptic resolution in a single dataset1.0 mm³Functional connectomics spanning multiple areas of mouse visual cortex

A Limited Set of Rigid Organizational Structures for Organizing and Funding Research Constrains the Forms of R&D That Get Done. NULL CORROBORATED by Gate A, and this is now the strongest row in the sample. A second searcher, working from the gap description alone and without sight of the queries below, ran eight independent searches of its own wording and also found no maintained series measuring which forms of R&D get done. Two searchers failing with different wording is a materially different claim from one searcher failing, and it is the reason to trust this row. The original six searches (recorded in research-log/searches/phase-3.json, phase 3, this gap — the search_log database table is empty, so that file is the provenance): researcher time spent on administration; metrics for diversity of research funding mechanisms; counterfactual measurement of funding structure against discoveries; outcome metrics for new organizational forms including FROs; indicators of institutional innovation in science funding; and randomised experiments in grant allocation. Near-misses, all failing for stateable reasons. (1) Administrative burden: US federally funded principal investigators report spending roughly 42% of their research time on administration. The gate objected that the gap's own sentence opens 'scientists are spending a lot of time not doing science', so this measures the first clause and was rejected against the second. The objection is fair and the rejection still stands, but on narrower grounds than before: the 42% is a real measure of one symptom, and cutting it to 20% would be a genuine gain that left the gap's actual claim — which forms of R&D get done — untouched. It is the wrong quantity, not a trivial one, and the FDP survey behind it last ran in 2018. (2) NIH's median age at first R01-equivalent award, an official series updated annually, has held near 42 across FY21-25 (mean 43-44). It is maintained, current and genuinely about rigidity, which makes it a stronger near-miss than either of the two originally listed — and it still measures who gets funded rather than which forms of research become possible. (3) The randomised evidence on funding allocation, including grant lotteries, is the right instrument but is a set of individual studies, not a series, and its own conclusion is that counterfactual assessment of funding structures is usually impossible because researchers have alternative sources. This is what Counterfactual required means: the quantity of interest is the research a different structure would have produced and this one did not, and no observation of the world we are in contains it. The null is not a search failure. It is the tier being correct.

Distributions across all 103 gaps

Measurability tier

Directly measurable
72 · 70%
Proxy only
19 · 18%
Verification contested
11 · 11%
Counterfactual required
1 · 1%

Kind of work in the way

Prediction and modeling
22 · 21%
Measurement and sensing
19 · 18%
Design search
16 · 16%
Coordination and institutions
15 · 15%
Reading and synthesis
10 · 10%
Running experiments
9 · 9%
Physical build
8 · 8%
Real-time control
4 · 4%

Tier against kind of work

Kind of workDirectly measurableProxy onlyVerification contestedCounterfactual requiredTotal
Coordination and institutions482115
Design search15·1·16
Measurement and sensing1612·19
Physical build8···8
Prediction and modeling1453·22
Reading and synthesis442·10
Real-time control4···4
Running experiments711·9

Tier against field

FieldDirectly measurableProxy onlyVerification contestedCounterfactual requiredTotal
Astrophysics31··4
Biophysics5·2·7
Biosecurity32··5
Cellular and Molecular Biology3···3
Chemistry7···7
Computation413·8
Ecology121·4
Geophysics and Climate321·6
Global Health3···3
Immunology2···2
Materials Science4···4
Mechanical Engineering7···7
Metascience22·15
Nanoscale Fabrication3···3
Neuroscience31··4
Physics5·3·8
Physiology and Medicine51··6
Social Science261·9
Space Engineering21··3
Synthetic Biology5···5
Social Science is the outlier: 6 of its 9 gaps are proxy-only, the highest share of any field.
What it cost to build
PhaseKindUnitsElapsedWhat happened
phase-0agent46 minRepo scaffold, schema, importer, additive guardrail, integrity report, search client, baseline snapshot.
phase-1agent10311 minFull-coverage labeling of all 103 gaps across 20 field batches: outcome, AI type and maturity, measurability tier.
phase-2agent3619 minStratified blind audit across 3 independent auditors, adjudication, findings report.
phase-2bagent10355 minTaxonomy revision (8th type, frame dimension), full independent blind relabel of all 103 gaps by 5 labelers, mechanical adjudication, withdrawal of the maturity gradient.
phase-3agent89 minprogress indicators: 8 rows across 4 tiers, 2 honest nulls, 29 logged searches. Local-session setup before this point (rebuild, network verification, ingest plumbing) is not counted.
phase-4agent46 minnew gaps: 4 proposed, 1 candidate dropped; dedup over the full export plus funding checks.
phase-5agent27 mintwo critical paths, 15 steps, plus the cross-field intersection. Expectations committed before the analysis in a separate commit.
phase-6agent127 minartifact (Next.js static export), CSV keyed on their id and slug, findings summary, cover note draft; found and fixed the adjudication rebuild defect.
revision-1agent73 minrebuilt around the argument in Convergent's voice, after a blind cold review
revision-2agent50 minsecond cold review: replotted the chart, fixed three dead links, stated three holes
revision-3agent41 mincut page one to a fifth, split into five pages, rebuilt the map on their components
revision-4agent22 minnew opening, all 103 outcome sentences rewritten, indicators page
revision-5agent33 minthird cold review: two reasoning errors corrected, outcomes deticked, attributes page
revision-6agentopenrobotics as the AI analogue for physical build, time ledger, drafting cost corrected
allhuman-review244 minDavid's own time reviewing, directing and rewriting, over the same window as the revision cycle. His estimate rather than an instrumented figure, and roughly equal to the agent's. The build phases above had no human review at all.

Agent time and human time are tracked separately, because a single blended number would be the first thing worth objecting to. Phase 0’s start was never instrumented, so it counts as zero and the 181 minutes for the build is a lower bound.

Calls that could have gone the other way

31 of them, each with the runner-up and what would reverse it. The runner-up was usually the more flattering option.

Maturity is scored relative to the specific gap, not to the capability class in general.
phase-2, confidence high
Two blind auditors independently flagged the taxonomy as ambiguous between the two readings, which give different answers. Both auditors and the original labeler had already converged on the for-this-gap reading, so fixing the wording changes no existing label but makes the maturity disagreement rate interpretable.
Runner-up
Score maturity per capability class globally, which is simpler to apply but collapses the distinction between autonomous experimentation in chemistry and in orbital assembly.
What would reverse it
If a future batch shows labelers cannot apply the for-this-gap reading consistently, switch to global class maturity and record the loss.
Keep a single primary AI type and a single tier per gap, and name composite gaps as a stated limitation rather than splitting them.
phase-2, confidence medium
Two gaps bundle several unrelated research programmes whose sub-components would take different types and tiers. Splitting them would mean creating gap records Convergent did not write, which violates the additive-only constraint and is exactly the schema redesign the brief forbids.
Runner-up
Allow multiple tiers per gap, or decompose composite gaps into sub-gaps.
What would reverse it
If Convergent respond that they would welcome decomposition, split them in a later version.
Publish the audit disagreement rate as measured, including the population-weighted figure, without re-adjudicating labels to improve it.
phase-2, confidence high
The brief requires the artifact to be auditable rather than confident. A tuned agreement rate would defeat the purpose of running the audit, and a near-zero rate would be evidence the audit was not independent rather than evidence the labels are good.
Runner-up
Adjudicate every disagreement to a single answer and report only the post-adjudication labels.
What would reverse it
Never — if this is reversed the audit stops measuring anything.
Downgrade every 'Proxy only' tier to confidence 'guess', including gaps the auditor never sampled.
phase-2, confidence high
The stratum ran 78 percent disagreement against 0 percent for Directly measurable and 9 percent for Verification contested, and all three auditors independently reported the category was repeatedly the nearest alternative and almost never won. A category that two independent labelers applying the same written definition cannot agree on has not earned confident anywhere, and downgrading only the sampled failures would let an unreliable category keep its confidence wherever the auditor happened not to look.
Runner-up
Downgrade only the eight sampled disagreements, which understates the problem and implies the unsampled Proxy only labels are sound.
What would reverse it
If Proxy only is redefined sharply enough that a re-run audit brings its disagreement rate near the other strata, restore confidence on relabeled rows.
Do not add a 'Control of physical systems' category to the AI type taxonomy in this version; record it as a finding instead.
phase-2, confidence medium
A blind auditor found that closed-loop control of a physical system has no home: fusion plasma control chooses no experiments so it is not autonomous experimentation, and builds nothing so it is not physical build. The case is real, but adding an eighth category mid-run would invalidate the 103 labels already applied and the audit measured against them. Naming it is the contribution; acting on it belongs in a later version.
Runner-up
Add the eighth category now and relabel all 103 gaps against it.
What would reverse it
If a re-run is scheduled anyway, add the category first and relabel from scratch.
Withdraw the working-now maturity gradient as a finding. Report maturity per gap with its disagreement rate, and make no aggregate claim resting on it.
phase-2b, confidence high
An independent relabel of all 103 gaps did not reproduce it. Coordination and institutions inverted from 7 percent working-now to 67 percent, physical build stopped being zero, and maturity agreement between the two passes was 63 of 103 with 23 of 40 disagreements moving the same direction. The cause is an unresolved ambiguity between availability and efficacy readings of Working now, which diverge completely for institutional capability. The gradient was the most striking thing this analysis produced, which is exactly why publishing it after it failed replication would be indefensible.
Runner-up
Keep the gradient with a caveat, on the grounds that v1 applied the efficacy reading consistently. Rejected: consistency within one labeler is not reliability, and a reader cannot tell which reading produced a published number.
What would reverse it
If a forced-choice maturity rubric is written and a third independent pass reproduces an ordering, the gradient can be reinstated on that measurement.
Adjudicate relabel disagreements mechanically: agreement keeps the label as confident, disagreement takes v2 and is flagged guess. No case-by-case adjudication.
phase-2b, confidence high
I authored v1, so I am not a neutral adjudicator and picking winners case by case would reintroduce exactly the bias the blind pass was built to remove. v2 wins ties because it alone had the complete eight-category taxonomy. A guess flag is the honest record of two independent labelers failing to converge on the same text.
Runner-up
Adjudicate each of the 57 disagreements on the merits, which yields more confident labels but launders my prior judgement through a process that looks independent and is not.
What would reverse it
If a third independent labeler is run, majority vote across three passes would be a legitimate adjudication rule and would restore confidence to the cases where two of three agree.
Define the ai-as-object frame by a criterion (the gap would still exist if AI did not) rather than by enumeration, and add Labor-Replacing AI Could Lead to Human Disempowerment as a fourth member.
phase-2b, confidence high
The first revision listed three gaps instead of stating a test. A blind relabeler immediately found a fourth meeting the same description and correctly noted that by the letter of the rule it stayed ai-as-instrument. An enumeration cannot be applied to a gap nobody thought of, which is the thing a labeler has to do.
Runner-up
Keep the enumeration and extend it as cases arise, which fails the same way again on the next unanticipated gap.
What would reverse it
If the criterion admits gaps that are clearly not about AI, tighten it rather than reverting to a list.
Recorded the quantum gravity gap as an honest null even though a real, published, improving quantity exists (the lower bound on the QG energy scale from photon dispersion: Fermi-LAT E_QG,1 > 7.6 E_Planck, GRB 090510, 2013; extended by LHAASO on GRB 221009A, 2024).
phase-3, confidence medium
The bound constrains linear-in-energy Lorentz violation, which most leading quantum-gravity programmes do not predict, and the community has no target and no shared claim about what any particular value of it would establish. A number that improves without agreement on what its improvement means is not a progress indicator for the gap. This is precisely what the Verification contested tier asserts, so recording it as an indicator would have contradicted our own tier label.
Runner-up
Record the Fermi/LHAASO bound as the indicator with confidence 'guess'. This would have been the more flattering choice — a tier-3 gap with a real number is a more interesting headline than a null — and it is why the decision is logged rather than assumed.
What would reverse it
A published community statement, roadmap or review that names a target value for E_QG or an equivalent parameter and says what reaching it would settle. Reverse immediately if one exists.
Used a 2025 arXiv preprint as the source for the retraction rate rather than a peer-reviewed paper or the Nature news analysis.
phase-3, confidence medium
Nature's site returns an authentication redirect and could not be fetched, and the plan forbids sourcing a number from a search snippet. The preprint gives an explicit denominator (Web of Science, 2000-2024) which the widely quoted figures do not, and arXiv is checkable by engine/validate-indicators.mjs, which verified the title. The row is marked 'guess' and its rationale states that published estimates disagree by roughly a factor of three.
Runner-up
Leave the fraud gap without an indicator. Rejected because a contested number with its dispute stated is more useful to a reader than a blank, and because the dispute is the substance of the Proxy only tier.
What would reverse it
A peer-reviewed rate with a stated denominator becomes fetchable; replace the source and re-run the validator.
Left target_value NULL on five of the six non-null indicators, populating it only where a programme or roadmap states one (brain volume, silicon scaling).
phase-3, confidence high
The plan requires a target_basis for every target and rules out invented round numbers. For elapsed telescope time, materials-per-day, trial cost and retraction rate, no funder, roadmap or community body has published a figure. In two of those cases a plausible-looking number was available and rejected on inspection: the 2.5 million USD median for US-funded phase 3 trials is a different population from industry pivotal trials, and a lower retraction rate is not unambiguously better because the quantity measures detection.
Runner-up
Derive targets from the best observed case in each series. Rejected: it manufactures an improvement factor out of a sampling difference and would read as analysis rather than as the arithmetic it is.
What would reverse it
A funder or roadmap publishes a target for any of these quantities.
Persisted the Phase 3 search transcript to research-log/searches/phase-3.json, and added ingest scripts so indicators, new gaps, critical paths, decisions and runs all rebuild from files.
phase-3, confidence high
db/gapmap.sqlite and research-cache/ are both gitignored, by deliberate earlier design. Without this, every Phase 3-6 row and every search behind the two nulls would exist only in an untracked binary and an untracked cache, and the claim that the nulls are provable rather than asserted would be false in the repository a stranger actually receives.
Runner-up
Un-ignore research-cache/. Rejected: it commits several megabytes of third-party search results to hold a few hundred kilobytes of evidence, and the titles and URLs are the part that matters.
What would reverse it
None expected. If cache contents themselves become disputed, commit the cache.
Dropped the fifth candidate gap — queue and turnaround time at shared nanofabrication user facilities — rather than shipping four with one weak.
phase-4, confidence high
Five logged searches turned up no published wait-time or turnaround data for NNCI or comparable facilities, so the gap could not be grounded in an observed rate limit. It also had the weakest funding check of the five: NSF's National Nanotechnology Coordinated Infrastructure exists precisely to provide open access across 16 sites, so proposing access latency as an unowned gap would have required arguing against a programme built for it, on no evidence.
Runner-up
Include it with confidence 'guess' and an honest funding check. Rejected: the plan asks for 3-5 gaps, four are well grounded, and an ungrounded fifth costs more credibility than the count gains.
What would reverse it
Facility-level turnaround statistics become available, from NNCI reporting or a user survey.
Proposed the coating thermal noise gap even though coatings research is funded, and said so in the funding check rather than around it.
phase-4, confidence medium
The NSF LSC Center for Coatings Research and Italy's ETIC project both fund this work as a detector subsystem inside one instrument programme. The gap proposed is the cross-instrument materials framing — the same loss angle bounds optical clocks and cavity-stabilised lasers, and no programme owns it at that level. Writing 'not clear of funding' in the funding_check and stating exactly what is funded is more useful than a claim of novelty a reader could puncture in one search.
Runner-up
Drop it as already funded. Rejected: it would discard the strongest tier-1 candidate over a framing question, and the framing is the contribution.
What would reverse it
A programme is found that funds low-mechanical-loss optical coatings as a materials target across instrument classes.
Added ai_type, maturity and tier columns to new_gaps rather than writing proposed gaps into gap_ai_types and gap_measurability.
phase-4, confidence high
Those two tables hold foreign keys into gm_gaps. A proposed gap is deliberately not one of theirs, and giving it a row in the same tables would make the two indistinguishable in every downstream query and export. The plan requires new gaps to carry the same three labels; this carries them without blurring the boundary the whole project rests on.
Runner-up
Relax the foreign keys so both kinds of gap share the label tables. Rejected: it trades the clearest structural guarantee in the schema for a small convenience in the export.
What would reverse it
None expected.
The Notion connector, unavailable at the start of this run, was authorised mid-run, so chain 2 is built from its named source of record after all.
phase-5, confidence high
Recorded because a decision to substitute public sources was logged earlier in this same run and then reversed; the ledger should show the reversal rather than quietly drop it. Public sources gathered before the connector returned are kept and cited alongside the Notion page, since they are citable in the artifact and the page is not.
Runner-up
Proceed on public sources only. No longer necessary.
What would reverse it
None.
Gave the chain 1 headline two accountings rather than one, and reported the weaker of the two as the honest figure.
phase-5, confidence high
The expectation recorded in advance was that closing every cognitive link would change the total duration 'very little'. Counting JWST's seven-year science-case period as compressible cognitive work makes the saving nine and a half of thirty-two years, which is under a third but not 'very little'. Counting it as community consensus formation — which is what the milestone record shows it was — makes the saving two and a half years. Both are stated. Tuning the classification to protect the prediction would have been the easy move and would have destroyed the value of recording the prediction in advance.
Runner-up
Report only the realistic accounting. Rejected: the generous one is the number a skeptical reader would compute, and pre-empting it is worth more than winning on it.
What would reverse it
Evidence that the 1989-1996 period was analysis-limited rather than consensus-limited.
Split chain 2's reviewer recruitment link into matching and willingness inside a single link rather than making them two links.
phase-5, confidence medium
They are sequentially inseparable — an editor cannot recruit before identifying — so two links would imply an ordering that does not exist. But the whole finding of the chain lives in the split: matching is a prediction problem AI is applicable to and, on the published evidence, not even best at; willingness is labour supply that no matching system touches. The split is carried in the blocker and rationale fields of one link.
Runner-up
Two links, 'matching' then 'recruitment'. Rejected as a false serialisation.
What would reverse it
A venue is found where identification and invitation are genuinely separate stages with separate durations.
Verified Aaron Tohuvavohu against both the export and the live site before relying on the association.
phase-5, confidence high
The plan flagged it as needing checking. Confirmed twice: resource 1c3cb37e-2a00-80a1-8ddf-fb19d0b8b0ee, type Individual, is cited by the capability 'Space Telescope Factory', which is attached to the telescope gap; and the name appears in the acknowledgments list on gap-map.org/about.
Runner-up
Rely on the plan's statement. Rejected — the plan itself asked for the check.
What would reverse it
The site's acknowledgments change.
Cut all three 3ie figures from the cover note rather than softening them: the count of evidence gap maps, the Development Evidence Portal totals, and the absolute-gap versus synthesis-gap terminology.
phase-6, confidence high
The plan flagged them as needing verification and they did not verify. 3ie's own gap maps page states no total; their own blog posts give portal figures that disagree by roughly a factor of three (3,745 impact evaluations in one, 'more than 11,000' in another) and nothing found states 21,800 or 1,700; and neither the gap maps page nor the working paper page uses the terms 'absolute gap' or 'synthesis gap'. The substance of the distinction is on their page verbatim and is kept as a quote in the notes, so a later comparison has a defensible form.
Runner-up
Keep the figures with an 'approximately'. Rejected outright: the cover note's entire argument is that this work checks things, and an unverified number in it would be the one thing a reader could puncture.
What would reverse it
3ie publishes a stated total, or the Development Evidence Portal exposes counts to a non-JavaScript client.
Rendered the two chains as HTML/SVG in the artifact and shipped the Mermaid source alongside, rather than running a Mermaid renderer in the page.
phase-6, confidence high
Each diagram is eight boxes and an arrow. A client-side Mermaid runtime would have been the single largest dependency in the build, for output that is less accessible than the HTML version — which also carries the binding flag as a word in the link table, so the distinction never rests on colour. The Mermaid source is in docs/critical-paths.md and in a copy block under each chain, so nothing is lost for anyone who wants to paste it elsewhere.
Runner-up
Add mermaid as a dependency and render client-side. Rejected on build weight and accessibility, not on effort.
What would reverse it
A reviewer wants the diagrams to match Convergent's own tooling exactly.
Wired engine/adjudicate.mjs into engine/rebuild.mjs after finding that a clean rebuild silently dropped every Phase 2 confidence downgrade.
phase-6, confidence high
adjudicate.mjs writes confidence flags directly to SQLite and had only ever been run by hand. Because db/gapmap.sqlite is gitignored, a rebuild from files produced 5 tier guesses and 10 AI-type guesses where the published findings say 24 and 20 — including the blanket downgrade of every 'Proxy only' assignment, which is one of the project's headline results. The artifact was reading those wrong numbers before this was caught. The script is idempotent, so running it inside rebuild is safe.
Runner-up
Bake the downgrades into the label files. Rejected: the labels are what the labeler wrote, and overwriting them would erase the fact that adjudication changed something.
What would reverse it
None. This was a defect.
Kept the JAMA article page as the source URL for the clinical trial indicator even though it returns HTTP 403 to programmatic clients, and recorded that in the row's rationale.
phase-6, confidence medium
It is the page the number was actually read off, the DOI verifies against Crossref, and a human clicking the link sees the paper. Swapping to a URL that passes an automated check but is not where the number was read would make the link check pass and the provenance worse.
Runner-up
Point at the PubMed record, which returns 203. Rejected: the PubMed abstract does not carry the IQR or the design breakdown the rationale relies on.
What would reverse it
JAMA stops blocking, or an open-access copy with the same figures appears.
Replaced mechanical adjudication with adjudication on the merits for the maturity field only, re-deciding 17 of 54 reviewed gaps in research-log/relabel-adjudication.json. Type is still adjudicated mechanically by taking v2.
maturity-repair, confidence high
The rule "on disagreement take v2" was chosen so the author of v1 could not launder his own judgement, and it worked for that. But a rule that cannot be wrong removes the bias and every check on validity together, and v2 read "Working now" as availability rather than efficacy, so the rule adopted the wrong reading 25 times in one direction. It produced "AI Could Be Misused" as Coordination and institutions / Working now, which no reader would accept. Efficacy is now pinned in methodology/taxonomy.md with the DNA-synthesis-screening example, and each re-adjudicated entry carries its own per-gap reasoning in its note, so the judgement is auditable row by row rather than resting on a rule.
Runner-up
Keep the mechanical rule and disclaim maturity in the artifact. Rejected: the labels are wrong, not merely uncertain, and a caveat a reader cannot act on is noise. The second runner-up was to re-run a third blind pass under the pinned definition, which is the more rigorous fix and was rejected on cost — it would relabel 103 gaps to correct a defect that provably reaches only 54, and every change here landed on a gap where the two existing passes already disagreed.
What would reverse it
A third independent pass under the pinned efficacy definition disagrees with this adjudication on more than about a fifth of the 54 reviewed gaps, or David rules the other way on the escalations in research-log/maturity-repair/escalations.md, in which case the affected entries revert individually rather than the method being abandoned.
Made agreement between the two labelling passes an overridable default in engine/apply-relabel.mjs, rather than a branch that always takes v1. An explicit entry in relabel-adjudication.json now wins even where both passes agreed, and the rebuild logs the override count separately.
maturity-repair, confidence high
David ruled on four escalations and two of them — Ephemeral Societal Data and Inadequate Emergency Climate Interventions — were gaps both passes had agreed on, which the code had no way to express. Unoverridable agreement is the same defect as the unoverridable 'on disagreement take v2' rule, one layer down: two independent passes sharing a misreading is exactly how 'AI Could Be Misused' came out as Working now. Agreement is evidence, not proof. Logging the override count on its own line keeps it from happening quietly.
Runner-up
Edit the v1 maturity in research-log/labels/*.json so those gaps fall into the disagreement branch. Rejected: it would falsify the record of what the first pass actually said, and move the v1-v2 agreement figure from 63/103 to 61/103, a number both docs/relabel-report.md and methodology/taxonomy.md cite.
What would reverse it
The override count printed by rebuild stops being small. If explicit overrides on agreed gaps become routine rather than exceptional, the problem is in the labelling passes and not the adjudication layer, and the fix belongs upstream.
Did not add a '5-10 years' maturity value, despite 63 of 103 gaps now sitting in '2-5 years'. Recorded as a proposal for a later pass, together with a proposal to separate the time question from the has-a-path question that 'Speculative' actually asks.
maturity-repair, confidence medium
David raised it against 'AI is Still Narrow', and the diagnostic supports him — three fifths of the map in one bucket is barely a label, and 'Speculative' conflates 'slow' with 'no known route'. But adding a fourth value means re-reviewing all 63 gaps currently at '2-5 years', because a value nobody has applied to the whole set is worse than three honest ones: those 63 would silently mean '2-5 or 5-10, unexamined'. That is a full relabel, and this branch is a repair.
Runner-up
Add the value and apply it only to the gaps this pass touched. Rejected for exactly that reason — a partially applied enum value makes the distribution unreadable, which is a worse artifact than the pileup it fixes.
What would reverse it
A pass is commissioned with a blind second reader to re-review all 63 gaps at '2-5 years', at which point the split should be done on both axes rather than by adding one bucket to one of them.
Recorded, and did not build, a blocked-on-adoption versus blocked-on-capability dimension. Eleven of this pass's downgrades turn on adoption rather than on whether the technique exists.
maturity-repair, confidence medium
David's note on the clinical-trials ruling. DNA synthesis screening exists and is not adopted; archiving works and permission is withheld; adaptive trials run and the field does not take them up. Maturity absorbs all of this and reports 'not ready', which is the wrong diagnosis for a funder, because the intervention for an unadopted capability is not more research. This is a genuine addition to Convergent's map rather than a correction to our own labels.
Runner-up
Encode it now as a fourth dimension on the augmentation tables. Rejected: it would need its own definition, its own labelling pass over 103 gaps and its own audit, and it would arrive in the same artifact as a repair — which is how a contribution turns into a rewrite of someone else's map.
What would reverse it
Convergent asks what would be worth adding next, in which case this is the strongest candidate on the evidence this pass produced.
Added docs/future-work.md and an app page at /missing/ ("What's missing") listing the nine pieces of open work, and did NOT regenerate app/public/data.json to go with it. The page reads every number it prints out of data.json instead of hardcoding them.
maturity-repair, confidence high
The artifact regeneration is its own track and is meant to run once after both the maturity repair and the review gates land; regenerating a large derived JSON now would conflict with the review-gate branch, which changes indicators, new-gaps and critical-paths. Reading from data.json means the page cannot contradict the rest of the site today, and it picks up the repaired maturity distribution automatically when the artifact track runs — the middle-bucket share it cites goes from 50% to 61% with no edit to the page.
Runner-up
Hardcode the post-repair numbers in the page copy. Rejected: it would make the site contradict itself until the artifact regenerates, which is worse than being briefly out of date, and it would need a second edit later that nobody would remember to make.
What would reverse it
The artifact track runs and the page still reads wrong, which would mean a number it needs is not in the data.json summary and should be added to engine/export-artifact.mjs rather than typed into the page.
Rename all eight AI-capability categories for display so every one names a kind of work, not a kind of model. Reading and synthesis, Prediction and modeling, Design search, Measurement and sensing, Running experiments, Real-time control, Physical build, Coordination and institutions.
revision-7, confidence high
Six of the eight were named after the AI that would do the work and two after the work itself, so a column that means 'what stands in the way' read as 'which AI would do it'. That made Coordination and institutions look like a category error when it was one of the two naming the thing consistently. David raised it; the diagnosis is his. Applied at the export boundary in engine/export-artifact.mjs, which walks the emitted object rewriting values, object keys and category references inside rationale strings, so summary buckets, cross-tabs, chain links, per-gap labels and the CSV all move together and no render site can be missed by hand.
Runner-up
Add a second attribute naming which AI capability could accelerate each gap, and show blocker and accelerator as two sections. Rejected for now because that attribute has never been labelled: it would need a full pass over all 103 gaps plus an audit, and for coordination gaps the honest answer is often 'none'. Recorded in docs/todo.md as an open question rather than closed.
What would reverse it
One constant in engine/export-artifact.mjs. The stored enum, the CHECK constraints and every research-log judgment are unchanged, so reverting is deleting the map.
AI Could Be Misused is Speculative, not 2-5 years.
revision-7, confidence high
David ruled on 2026-08-25. The gap has many partial mitigations -- evaluation regimes, deployment policy, convening -- and the maturity repair had moved it to 2-5 years on the strength of those existing as mechanisms. But the capability that would actually close the gap is alignment, and a substantial part of the field holds it may not be solvable at all. A capability whose feasibility is itself contested is what Speculative is for. This is the same distinction the maturity repair was built on: availability is not efficacy.
Runner-up
2-5 years, on the grounds that partial mitigation mechanisms exist and are being built. Rejected because partial mitigations do not close this gap, and the taxonomy asks what would move the gap rather than what is being attempted.
What would reverse it
A change in the state of alignment research such that the field stops treating the core problem as possibly unsolvable. Recorded in research-log/relabel-adjudication.json for this gap.
Dropped is_binding from the telescope chain entirely, and renamed what survives on the publishing chain to "carries the cost".
revision-6, confidence high
A blind cold reviewer pointed out that the telescope chain is strictly sequential, so removing any step shortens the total and every step is on the critical path. Marking four of eight as binding claimed a distinction the structure does not contain, and the marked set turned out to be exactly the steps no AI capability acts on, which made the headline finding restate its own labelling. The textbook sense of binding needs parallel paths and neither chain has them. What survives is the cost sense on chain 2, where three of seven steps genuinely account for a disproportionate share of reviewer and editor labor. That is a claim about distribution rather than about slack, so it gets its own words.
Runner-up
Keep the vocabulary on both chains and explain the difference in a footnote. Rejected because the circularity was real rather than presentational: no wording rescues a flag that has no discriminating power on the chain it is applied to. The arithmetic that replaced it -- AI acts on 9.5 of the 32.5 years -- is stronger and needs no vocabulary at all.
What would reverse it
A chain is built with genuine parallel structure, where some steps run alongside others. is_binding then recovers its classic meaning on a time axis and methodology/critical-path.md needs a third case.