{"generated_at":"2026-09-06","source":{"name":"Convergent Research — Fundamental Development Gap Map","version":"v1.0","url":"https://www.gap-map.org/","snapshot":"2026-07-29"},"summary":{"n_gaps":103,"n_fields":20,"n_capabilities":369,"n_edges":389,"n_new_gaps":2,"n_indicators":8,"n_indicator_nulls":1,"tier":{"Directly measurable":72,"Verification contested":11,"Proxy only":19,"Counterfactual required":1},"ai_type":{"Prediction and modeling":22,"Real-time control":4,"Measurement and sensing":19,"Running experiments":9,"Design search":16,"Physical build":8,"Reading and synthesis":10,"Coordination and institutions":15},"maturity":{"2-5 years":62,"Working now":26,"Speculative":15},"tier_by_ai_type":{"Prediction and modeling":{"Directly measurable":14,"Proxy only":5,"Verification contested":3},"Real-time control":{"Directly measurable":4},"Measurement and sensing":{"Directly measurable":16,"Verification contested":2,"Proxy only":1},"Running experiments":{"Directly measurable":7,"Verification contested":1,"Proxy only":1},"Design search":{"Directly measurable":15,"Verification contested":1},"Physical build":{"Directly measurable":8},"Reading and synthesis":{"Directly measurable":4,"Verification contested":2,"Proxy only":4},"Coordination and institutions":{"Proxy only":8,"Verification contested":2,"Directly measurable":4,"Counterfactual required":1}},"tier_by_field":{"Computation":{"Directly measurable":4,"Verification contested":3,"Proxy only":1},"Physics":{"Directly measurable":5,"Verification contested":3},"Chemistry":{"Directly measurable":7},"Synthetic Biology":{"Directly measurable":5},"Nanoscale Fabrication":{"Directly measurable":3},"Materials Science":{"Directly measurable":4},"Mechanical Engineering":{"Directly measurable":7},"Geophysics and Climate":{"Proxy only":2,"Directly measurable":3,"Verification contested":1},"Astrophysics":{"Directly measurable":3,"Proxy only":1},"Ecology":{"Verification contested":1,"Proxy only":2,"Directly measurable":1},"Space Engineering":{"Directly measurable":2,"Proxy only":1},"Biosecurity":{"Directly measurable":3,"Proxy only":2},"Social Science":{"Proxy only":6,"Verification contested":1,"Directly measurable":2},"Metascience":{"Proxy only":2,"Directly measurable":2,"Counterfactual required":1},"Global Health":{"Directly measurable":3},"Biophysics":{"Verification contested":2,"Directly measurable":5},"Physiology and Medicine":{"Directly measurable":5,"Proxy only":1},"Cellular and Molecular Biology":{"Directly measurable":3},"Immunology":{"Directly measurable":2},"Neuroscience":{"Directly measurable":3,"Proxy only":1}},"maturity_by_ai_type":{"Prediction and modeling":{"2-5 years":18,"Working now":3,"Speculative":1},"Real-time control":{"2-5 years":2,"Speculative":1,"Working now":1},"Measurement and sensing":{"2-5 years":9,"Speculative":1,"Working now":9},"Running experiments":{"Working now":5,"2-5 years":4},"Design search":{"Working now":2,"Speculative":3,"2-5 years":11},"Physical build":{"Speculative":4,"2-5 years":3,"Working now":1},"Reading and synthesis":{"Working now":5,"2-5 years":5},"Coordination and institutions":{"Speculative":5,"2-5 years":10}},"confidence":{"outcome":{"confident":102,"guess":1},"tier":{"confident":79,"guess":24},"primary_ai_type":{"guess":48,"confident":55}}},"gaps":[{"id":"1fccb37e-2a00-8071-859f-f25fac2df35c","slug":"silicon-based-electronics-face-fundamental-limits-in-dimensional-scaling","name":"Silicon-Based Electronics Face Fundamental Limits in Dimensional Scaling","description":"For over five decades, silicon-based CMOS technology has driven unprecedented progress in computing and information technology through dimensional scaling following Moore's Law. This miniaturization has led to exponential increases in transistor density, performance, and energy efficiency. However, as transistor channel dimensions shrink below a few nanometers, silicon and conventional bulk semiconductors (e.g., SiGe, III-V materials) are encountering insurmountable fundamental physical and material limits (heat dissipation, short-channel effects, etc.).","field":"Computation","outcome":"Compute keeps getting cheaper per operation after silicon channel physics gives out.","outcome_rationale":"Their description is unusually specific that the limits are fundamental and material, not engineering.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Gate length, carrier mobility and on-off ratio at manufacturable yield are direct device metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":0,"rationale":"Process development for a new channel material is an iterative fabricate-and-measure loop, which is what closed-loop experimentation automates.","confidence":"confident"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Design search vs Prediction and modeling; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Design search vs Prediction and modeling; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[{"id":5,"quantity":"Contacted gate pitch, most recent publicly disclosed volume-production node","current_value":"45","unit":"nm","as_of":"2022-12-31","target_value":"12","target_basis":"IRDS 2023 More Moore: physical channel length is projected to saturate near 12 nm from worsening electrostatics, with roughly 14 nm of width reserved for the device contact, and 'after 2031 there is no room for 2D geometry scaling, where 3D very large scale integration (VLSI) of circuits and systems using sequential/stacked integration approaches will be necessary.' A physical and roadmap limit, not an aspiration. The 2031 date is the roadmap's, and is recorded here rather than in target_value because this row's unit is nm.","source_title":"IEDM 2022 – TSMC 3nm","source_url":"https://semiwiki.com/semiconductor-manufacturers/tsmc/322688-iedm-2022-tsmc-3nm/","source_doi":null,"source_checked":"unchecked","reads_as":"TSMC's N3 ran a 45 nm contacted gate pitch, disclosed at IEDM in December 2022. This is the last figure in the public record, not the tightest pitch running today.","direction":"lower is better","context":"IRDS 2023 projects no room left for 2D geometry scaling after 2031, which makes this one of the last few steps rather than a point on a trend line. TSMC's N2 entered volume production in Q4 2025 and Intel 18A has shipped since, so at least two production nodes have passed without a comparable public pitch disclosure.","caveat":"Two problems, and the second is the larger. The number comes from conference reporting rather than a primary TSMC disclosure that could be fetched. And it is four years and two production nodes old: it is the most recent figure that is public, not the current state of the art, and no CPP figure for N2 could be found in the public record.","is_null_result":0,"rationale":"REVISED after Gate A, which graded this row stale. The 45 nm value and the IRDS quotation both verify against their sources and are unchanged. What was wrong was the tense: the row read as 'the tightest contacted gate pitch in volume production' while carrying as_of 2022-12-31, and TSMC's N2 (volume production Q4 2025) and Intel's 18A have shipped since. The claim has been narrowed to what the evidence supports — the most recent publicly disclosed figure — rather than re-dated, because no CPP disclosure for a later node could be found.\n\nThe gate also noted that target_value held a date statement in a row whose unit is nm. The target is now the 12 nm channel-length saturation the IRDS projects, and the 2031 date has moved into target_basis where it reads as the roadmap statement it is.\n\nTSMC disclosed the 45 nm contacted gate pitch for N3 at IEDM 2022, with a minimum metal pitch of 23 nm for N3E. Contacted gate pitch is the quantity that makes Convergent's gap statement — 'insurmountable fundamental physical and material limits' — checkable, because it is the pitch that stops scaling. Still a guess: the figure is Scotten Jones's conference reporting, TSMC's IEDM papers sit behind IEEE access, and the IRDS ground-rule tables are published as images with no gate-pitch series readable from the text.","confidence":"guess"}],"capabilities":[{"id":"1c2cb37e-2a00-801f-bf91-fb2cfe757da1","name":"Next-Gen 3D Integration of Discrete Components","description":"New logic and memory technologies based on CNTs or other structures, 3D integration with fine-grained connectivity, and new architectures for computation immersed in memory"},{"id":"1f0cb37e-2a00-80cb-a436-c2dfcbfbf3a6","name":"Atomically Thin 2D Semiconductors for Integrated Circuits","description":"Two-dimensional (2D) semiconductors, especially transition metal dichalcogenides (TMDs), have potential to shrink transistors beyond the scaling limits of silicon. Unlike silicon, which suffers from degraded performance at channel thickness < 12 nm, 2D materials are \"dangle-bond-free,\" meaning their surfaces are naturally stable and less prone to defects. This allows them to maintain high carrier mobility and low leakage currents even at atomically thin channel thickness.\r\n\r\nWafer-scale growth of monolayer TMD material has been demonstrated and prototype transistors have shown feasibility of outperforming conventional transistors. However, their industrial adoption requires optimization for industrial manufacturing, integration with semiconductor foundry processes, and standardized methods for characterizing material properties and device performance."}]},{"id":"1f0cb37e-2a00-803f-80fe-cb4166e4f94b","slug":"artisanal-nature-of-experimental-physics-platforms","name":"Artisanal Nature of Experimental Physics Platforms","description":"There is a lack of open and repeatable tooling to spread experimental physics into new areas, e.g., can ultracold atoms be more readily leveraged by people outside a small set of quantum physics labs to pave the way for more applied uses?","field":"Physics","outcome":"A materials group with no quantum optics background can run an ultracold atom experiment. Techniques currently locked inside the labs that invented them spread to applied fields.","outcome_rationale":"Their description asks precisely this: can these platforms be leveraged by people outside a small set of quantum physics labs.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Count of groups outside the originating labs operating a standardised platform is a direct measure of the stated goal.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"2-5 years","is_primary":0,"rationale":"Turning bespoke apparatus into a reproducible instrument is hardware engineering with a manufacturability objective.","confidence":"confident"},{"ai_type":"Real-time control","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Coordination and institutions vs Real-time control; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Real-time control","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Coordination and institutions vs Real-time control; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1f0cb37e-2a00-80b3-8818-f60e6fd1a9d2","name":"Standardized Open Experimental Physics Platform Tools","description":"Applying more modern digital fabrication and hardware engineering techniques could make experimental physics knowledge more transferable and applicable."}]},{"id":"1c1cb37e-2a00-80fb-9659-f73deeec29c5","slug":"inability-to-image-materials-atom-by-atom","name":"Inability to Image Materials Atom by Atom","description":"Many current imaging techniques lack the resolution to image materials on an atomic scale, limiting our understanding of material properties at the most fundamental level.","field":"Chemistry","outcome":"You can look at where the atoms actually are, instead of inferring arrangement from bulk averages, and tie a material's properties to its real structure.","outcome_rationale":"Their description frames the loss as understanding properties at the most fundamental level.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Whether individual atoms are resolved, and at what dose and field of view, is a direct instrument specification.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-805b-a538-ea600b3838c4","name":"Atom-by-Atom Imaging with Nanoscale Laser Ionization","description":"Utilize nanoscale laser ionization techniques to achieve atom-by-atom imaging of materials, providing unprecedented resolution of atomic structures."},{"id":"1c1cb37e-2a00-8047-a449-d52fa641cfa1","name":"Compact X-Ray Lasers","description":"Make X-ray lasers more accessible and compact"},{"id":"1c1cb37e-2a00-80af-9096-c2035d27e930","name":"Increasing the Range of Biopolymer Genetic Systems","description":"Use new enzymes to enable the use of synthetic genetic polymers.\nNote: don’t make mirror life https://www.science.org/doi/10.1126/science.ads9158 "}]},{"id":"1c1cb37e-2a00-80f9-8c4d-fa70c0da8ada","slug":"limited-understanding-of-the-chemical-reaction-space","name":"Limited Understanding of the Chemical Reaction Space","description":"Our overall knowledge of the chemical reaction space, including the catalysts that drive these reactions, is still rudimentary. We also lack detailed the large materials synthesis and processing datasets needed to enable highly predictive models.","field":"Chemistry","outcome":"Synthesis routes get planned rather than discovered. Most of chemical space has never been tested, and reaction outcomes across it turn predictable.","outcome_rationale":"Their description names the missing ingredient as large synthesis and processing datasets enabling predictive models.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Held-out reaction yield and product prediction accuracy, and coverage of reaction classes, are countable.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"Working now","is_primary":0,"rationale":"Reaction prediction models already work where data exists, which is exactly why data generation is the binding link.","confidence":"confident"},{"ai_type":"Running experiments","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Running experiments","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-805e-ad84-fa38834e0d21","name":"Comprehensive Mapping and Modeling of Chemical Space","description":"Map and model chemical reactions and catalysts more comprehensively to better understand reaction mechanisms and discover novel catalysts."},{"id":"1c8cb37e-2a00-80d4-8f78-ec0d82bca65a","name":"Open Synthesis Database","description":"A community-wide, structured repository (PDB-equivalent) of how materials are made, including processing parameters, environments, recipes, and results. Procedure logs and outcomes including failed experiments. "},{"id":"1c8cb37e-2a00-8049-8754-e7df26db4bbc","name":"Multi-Modal Chemical Data to Enable AlphaChem","description":"An open dataset that links multiple modes of chemical characterization, integrating existing databases that are currently siloed or behind paywalls. \n\nThe dataset should include structure, synthetic route, NMR/ IR spectra, and bioactivity to enable truly holistic chemical AI that can predict synthesis routes and spectra based on structure."}]},{"id":"1c1cb37e-2a00-8069-8477-e0e62e56db05","slug":"we-dont-have-easy-programmable-synthesis-of-bio-polymers-other-than-nucleic-acids","name":"We Don’t Have Easy Programmable Synthesis of Bio Polymers Other Than Nucleic Acids\n","description":"While long-chain nucleic acid synthesis is advancing rapidly, the programmable synthesis of other polymers remains underdeveloped, limiting our capacity to design and produce diverse synthetic polymers.","field":"Chemistry","outcome":"The design-build-test loop that transformed DNA and protein work extends to the rest of polymer chemistry.","outcome_rationale":"Their description sets up the contrast explicitly: nucleic acid synthesis advancing rapidly, everything else underdeveloped.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Achievable chain length, per-step coupling efficiency and monomer diversity are direct synthesis metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"2-5 years","is_primary":0,"rationale":"Coupling chemistry and protecting-group strategy for new monomer classes is a search problem over reaction conditions.","confidence":"confident"},{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Automation is not the binding constraint. Automated flow peptide synthesis already works; for glycans, polyketides and non-natural backbones there is no general programmable coupling chemistry for a self-driving lab to run. The loop is mature and has nothing to iterate on. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Running experiments","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Automation is not the binding constraint. Automated flow peptide synthesis already works; for glycans, polyketides and non-natural backbones there is no general programmable coupling chemistry for a self-driving lab to run. The loop is mature and has nothing to iterate on. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-80ab-8962-fd3b0b579e8f","name":"Solid-Phase Synthesizers for Other Polymers","description":"Build solid-phase synthesizers capable of the universal, programmable synthesis of polymers such as proteins, peptides, spiroligomers, carbohydrates, and RNA mimetics."}]},{"id":"1c1cb37e-2a00-8043-bb0a-f10b4d7f65c8","slug":"we-cant-yet-replicate-animal-olfaction-synthetically-as-a-sensing-and-classification-modality","name":"We Can’t Yet Replicate Animal Olfaction Synthetically as a Sensing and Classification Modality","description":"We currently lack a comprehensive model explaining how biological systems decode and classify chemical signals through olfaction. Understanding this process is critical for applications ranging from flavor science to disease diagnostics to understanding and harnessing animal communication.","field":"Chemistry","outcome":"Smell turns into a measurement. Predict perceived odour from structure and chemical sensing becomes diagnostic rather than descriptive.","outcome_rationale":"Their description names flavour science, disease diagnostics and animal communication as the downstream unlocks.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Predicted versus panel-rated odour descriptors on held-out molecules is a direct and already-run evaluation.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":0,"rationale":"Receptor-binding maps require multiplexed assay readout at a scale current sensing does not reach.","confidence":"confident"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-80d0-8540-df20b02c0e3e","name":"Build an Olfaction Decoding Model","description":"Develop an integrative model that explains how the olfactory system decodes and classifies complex chemical stimuli, linking molecular features to perceived odors."},{"id":"1c1cb37e-2a00-8001-9a69-edf60684a8c1","name":"Map of Odorant Receptor Binding","description":"Humans have ~400 odorant receptor genes that encode functional GPCR proteins. Binding ligands have been identified for ~80. Mapping the binding profiles of olfactory receptors from humans and other animals would enable novel biosensors and reveal novel therapeutic targets."}]},{"id":"1c1cb37e-2a00-8026-a3fb-ffd7b5bf54f1","slug":"manual-and-laborious-nature-of-chemical-synthesis","name":"Manual and Laborious Nature of Chemical Synthesis","description":"Chemical synthesis remains largely manual, limiting throughput and reproducibility. The field requires robust automation to accelerate discovery and production of new molecules.","field":"Chemistry","outcome":"How many compounds get tested each year stops depending on how many synthetic chemists there are.","outcome_rationale":"Their description names throughput and reproducibility as the losses from manual synthesis.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Compounds synthesised per instrument-day and human hours per compound are direct throughput measures.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"2-5 years","is_primary":0,"rationale":"General-purpose manipulation of arbitrary glassware and solids remains the limit on what automation can cover.","confidence":"confident"},{"ai_type":"Running experiments","maturity":"Working now","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Running experiments","primary_maturity":"Working now","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-80a1-9b70-f4ddebf822b1","name":"Chemputers for Automated Synthesis","description":"Implement a “chemputer” system to automate chemical synthesis processes, reducing human intervention and increasing reproducibility."},{"id":"1c1cb37e-2a00-8041-9f1e-d03852660262","name":"Broader Chemistry Automation","description":"Advance the field of chemistry automation through additional robotics and high-throughput platforms, enabling scalable synthesis processes."},{"id":"1c1cb37e-2a00-806b-9cb7-e4441ea649ea","name":"Modular Synthesis Using Improved Building Blocks","description":"Develop and utilize a better set of standardized building blocks to enable modular synthesis, making the assembly of complex molecules more efficient and scalable."},{"id":"1c1cb37e-2a00-8020-aaa7-ca63427edc46","name":"Generative Model\n  Based Parallel Library Synthesis","description":"Synthesize massive biopolymer libraries according to the statistics of a generative model"},{"id":"1c1cb37e-2a00-807e-9d70-e8d9ac9aef86","name":"Cheap Enzymatic DNA Synthesis","description":"Direct long DNA synthesis could still be cheaper. \nBiosecurity consideration: implementation should be governed by security measures such as:\n• Trustless DNA synthesis screening\n• Hardware lock for DNA synthesizer\n\nSee these capabilities under the “Risks of Malicious Bioengineering” bottleneck."},{"id":"1c1cb37e-2a00-8014-916a-f6ba1b6c9b15","name":"Cheap Long Peptide\n  Synthesis","description":"Fast on-demand complex peptide manufacturing"},{"id":"1c1cb37e-2a00-8087-8b0c-c7eec822f9e7","name":"In-Vivo Externally Programmable DNA/RNA Synthesis","description":"Physics based control across the cell membrane of DNA synthesis inside the cell as a route to maximize “bandwidth across the cell membrane”"}]},{"id":"1c1cb37e-2a00-803c-a293-e7aa6da74419","slug":"protein-design-has-been-limited-to-static-bio-mimetic-structures","name":"Protein Design Has Been Limited to Static, Bio-mimetic Structures","description":"Protein engineering has largely focused on designing static structures that closely mimic natural proteins. This narrow approach limits the creation of truly novel or highly functional enzymes.","field":"Synthetic Biology","outcome":"Enzymes get designed for jobs evolution never had to solve, rather than adapted from something that already exists.","outcome_rationale":"Their description names the constraint as static, biomimetic design limiting truly novel or highly functional enzymes.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Catalytic efficiency of designed enzymes and design success rate per attempt are direct and published.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"Working now","is_primary":0,"rationale":"Design-build-test-learn cycles for protein variants are already automated at high throughput.","confidence":"confident"},{"ai_type":"Design search","maturity":"Working now","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 Working now / v2 2-5 years; mechanical rule had given 2-5 years). RFdiffusion and ProteinMPNN are design search, and they produce folds with no natural template today; designed switches and hinges have escaped the static case too. Flagged guess because designed dynamics is far behind designed structure and the gap names both. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Design search","primary_maturity":"Working now","primary_rationale":"Maturity re-adjudicated on the merits (v1 Working now / v2 2-5 years; mechanical rule had given 2-5 years). RFdiffusion and ProteinMPNN are design search, and they produce folds with no natural template today; designed switches and hinges have escaped the static case too. Flagged guess because designed dynamics is far behind designed structure and the gap names both. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-8082-8e27-ef315ef6c142","name":"Protein Carpentry","description":"Develop techniques for “protein carpentry” that allow for the precise construction and remodeling of protein structures to yield dynamic, functional enzymes."},{"id":"1c1cb37e-2a00-80a2-a2d5-f1cb9572214c","name":"AI Enzyme Design and De Novo Design of Novel Protein Functions","description":"Leverage generative AI to design new enzymes with novel functions by predicting active conformations and optimizing catalytic activity. Especially for complex redox reactions and reactions that go beyond classic biological catalysis."},{"id":"1c1cb37e-2a00-8085-9370-c37f84819bab","name":"Non-Canonical Amino Acids","description":"Expand the chemical diversity of proteins by incorporating non-canonical amino acids, thereby enabling functions beyond those accessible with natural amino acids."},{"id":"1c1cb37e-2a00-807b-ab7a-c9a9eddf0a33","name":"Enzyme Stabilization in Extreme Environments","description":"Expand the operating range of biological enzymes to new solvents, enabling anhydrous enzymes, gas phase and vacuum biochemistry"},{"id":"1c1cb37e-2a00-800a-9d14-dc9f4a12e5cd","name":"Biosynthesis of Inorganic Materials","description":"Biosynthesis of C allotropes, e.g. (strong) polyynes and chiral CNTs; Biosynthesis of Si allotropes e.g. crystalline Si from SiO2 for photovoltaics"}]},{"id":"1c1cb37e-2a00-805f-8b53-dd343406ab07","slug":"limited-microbial-hostschassis-organisms","name":"Limited Microbial Hosts/Chassis Organisms","description":"Scientists are constrained to a small number of microbial hosts for bioproduction, limiting the diversity and efficiency of engineered biological systems. Expanding the repertoire of microbial hosts could unlock novel biochemical pathways, enabling the production of a wider array of biomolecules and improving the efficiency of biosynthetic processes. It is important to address any biosafety and biosecurity risks associated with developing such technologies.","field":"Synthetic Biology","outcome":"Bioproduction escapes the handful of domesticated microbes it currently runs on, and with it the biochemistry those few hosts cannot perform.","outcome_rationale":"Their description names novel biochemical pathways and wider biomolecule range as the unlock.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Number of chassis with production-grade toolkits, and achievable titre per host, are direct and countable.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Design search vs Running experiments; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Running experiments","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Design search vs Running experiments; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1f0cb37e-2a00-8067-b51a-d86a4e3393b9","name":"New Microbial Chassis","description":"Create platforms that simplify the process of identifying and adapting new microbial hosts, providing recipes for their use in synthetic biology applications.\n\nThis could take advantage of biocontainment approaches to enhance safety."}]},{"id":"1c1cb37e-2a00-807e-b924-fe555253617a","slug":"inability-to-program-complex-organisms-and-developmental-pathways","name":"Inability to Program Complex Organisms and Developmental Pathways","description":"Current genetic tools primarily enable modification of simple organisms. Programming more complex organisms and orchestrating entire developmental pathways remains a major challenge.","field":"Synthetic Biology","outcome":"Genetic engineering moves up from single genes in simple organisms to whole body plans and the developmental programmes that build them.","outcome_rationale":"Their description names the boundary directly: simple organisms tractable, complex organisms and developmental pathways not.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Whether the specified developmental outcome occurs, and at what penetrance, is directly observable.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"Speculative","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Design search","primary_maturity":"Speculative","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-805a-a009-e07d5b06675b","name":"Easier Genetic Access to Plants","description":"Develop advanced technologies to facilitate genetic manipulation in plants, broadening the range of programmable organisms. "},{"id":"1c1cb37e-2a00-80d1-9d3f-fed0fd7943d2","name":"Developmental Biology Read-Write Platforms","description":"Create integrated platforms that allow for real-time monitoring, modeling and precise manipulation of developmental processes, including with bioelectric and other novel control layers. \n\nGoals could include: Genome encoding of symmetry (e.g. 2 to 8-fold radial), accurate size ratios, Branching pattern codes, Natural & synbio eutely and other counting mechanisms"},{"id":"1c1cb37e-2a00-80eb-a77e-cab539293187","name":"Map of Cell Adhesion Codes","description":"Decode the molecular signals governing how cells communicate and organize to form complex 3D structures, a fundamental process in tissue formation, organ development, immune cell targeting, etc."}]},{"id":"1c1cb37e-2a00-8008-8543-ca3099dae993","slug":"lack-of-applied-synthetic-biology-platforms","name":"Lack of Applied Synthetic Biology Platforms","description":"Applied synthetic biology is underutilized in applications such as building sustainable food systems and repairing the environmental damage caused by conventional agriculture and industry. Despite advances in tools and chassis engineering, there are few robust platforms that translate synthetic biology into scalable, field-ready solutions. This includes not only the production of low-impact proteins and agricultural inputs but also bioremediation technologies for legacy pollutants—such as pesticide-laden soils, heavy metals, and nutrient runoff—that degrade ecosystems and constrain land use. A new generation of synthetic biology platforms is needed to address both sides of the problem: replacing harmful production methods and cleaning up their long-term consequences.","field":"Synthetic Biology","outcome":"Synthetic biology starts acting on real land and real pollution, not just on demonstrations in a lab.","outcome_rationale":"Their description names the shortfall as few robust platforms translating to scalable, field-ready solutions.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Tonnes produced, hectares remediated and cost per unit against the conventional alternative are direct.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"2-5 years","is_primary":0,"rationale":"Bioremediation pathway and chassis design remain genuine design problems for legacy pollutants.","confidence":"confident"},{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Physical build vs Running experiments; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Running experiments","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Physical build vs Running experiments; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-806a-af47-f7c449b19a11","name":"Synthetic Meat Production via Fungus Engineering","description":"Develop synthetic meat production processes based on fungus engineering to create affordable, ethical, and sustainable protein sources."},{"id":"1c1cb37e-2a00-806c-ade2-c903dd306344","name":"Ethical Food Production Interventions","description":"Implement alternative strategies that enhance the ethical aspects of food production without compromising cost-performance."},{"id":"1c8cb37e-2a00-8045-9056-da1570c17f64","name":"Bioremediation","description":"Engineered microbes and plants for environmental remediation (e.g., superfund and landfill clean-up and mining) that can survive in toxic conditions and degrade diverse classes of pollutants while concentrating and mining valuable elements.\n\nBiocontainment risks need to be addressed. "}]},{"id":"1c1cb37e-2a00-80cd-9ae4-ee6fa18c02ed","slug":"poor-scalability-of-bioreactors-limits-biomanufacturing","name":"Poor Scalability of Bioreactors Limits Biomanufacturing","description":"Current bioreactor designs are inefficient when scaling up production processes, limiting the ability to produce bioproducts at industrial scales.","field":"Synthetic Biology","outcome":"Bioprocess economics currently fall apart at exactly the volume where they start to matter commercially. Yields hold from bench to tank.","outcome_rationale":"Their description names inefficiency on scale-up as the specific failure.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Titre, rate and yield at stated working volume, and capital cost per litre, are the industry's own metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Physical build vs Prediction and modeling; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Physical build vs Prediction and modeling; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-809a-bb03-e09175225115","name":"Modularization of Bioreactors","description":"Develop modular bioreactor designs that can be easily scaled up, offering flexibility and improved efficiency in industrial bioproduction."},{"id":"1c1cb37e-2a00-80e6-beab-cd8b014b0e22","name":"Novel Materials and Sterilization of Bioreactors ","description":"Replace steel steam sterilized bioreactors with something more scalable"}]},{"id":"1c1cb37e-2a00-8057-ae20-c47fd74df08d","slug":"synthetic-biology-platforms-are-over-reliant-on-evolved-cells-that-we-dont-fully-understand-or-control","name":"Synthetic Biology Platforms Are Over-Reliant on Evolved Cells That We Don’t Fully Understand or Control","description":"We currently perform synthetic biology using naturally evolved (“kludgy”) cells rather than truly bottom-up engineered cells. This bottleneck limits our ability to design fully customizable biological systems.","field":"Nanoscale Fabrication","outcome":"Cells get built from specified parts. Synthetic biology stops inheriting mechanisms that nobody has characterised.","outcome_rationale":"Their description names the reliance on kludgy evolved cells as the limit on fully customisable biological systems.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Whether a bottom-up construct self-replicates, and how many functions are supplied by designed rather than borrowed components, are direct and countable.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"Speculative","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Design search","primary_maturity":"Speculative","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-8038-b86c-ec0adcf9d0d8","name":"Bottom-Up Synthetic Cells","description":"Develop entirely synthetic cells constructed from the ground up, rather than relying on evolved cell systems. This approach would enable precise control over cellular functions and properties, although risks must also be considered, e.g., https://www.science.org/doi/10.1126/science.ads9158"}]},{"id":"1c1cb37e-2a00-80d4-944d-f9dd2104cc6e","slug":"inability-to-perform-chemistry-with-direct-positional-control","name":"Inability to Perform Chemistry with Direct Positional Control","description":"Our current methods do not allow precise control over the positional placement of atoms or groups during chemical synthesis, limiting our ability to build molecules with atomic precision. A general-purpose approach to atomically precise fabrication was envisioned by Drexler in the 1980s and Feynman in the late 1950s. DNA origami made a leap in 2006, but DNA is in some key ways a much less precise and versatile nanoscale building material than proteins/peptides. A promising path would extend “DNA origami” to “protein carpentry” by adapting Beta Solenoid proteins, or other modular protein components with programmable binding properties, as lego-like building blocks and then using the latter to construct massively parallel protein-based 3D printers for lego-like covalent assembly of a restricted set of chemical building blocks. This one is riskier: how programmably can we really control protein assembly, and could we bootstrap from initial crappy prototype protein-carpentry-and-or-DNA-origami-based molecular 3D printers to genuinely useful ones? ","field":"Nanoscale Fabrication","outcome":"Put atoms where you want them rather than where chemistry prefers. Arbitrary molecular structures turn buildable instead of merely discoverable.","outcome_rationale":"Their description traces the ambition from Feynman and Drexler through DNA origami to protein carpentry, and the unlock is positional control itself.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Positional placement accuracy and per-step assembly yield are direct and unambiguous, which is why their own text can be candid about the risk of failing.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"Speculative","is_primary":0,"rationale":"Molecular 3D printing and vacuum mechanosynthesis are fabrication capabilities with no working instance.","confidence":"confident"},{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (maturity: Speculative vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (maturity: Speculative vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-8029-82a2-ff0e872388dc","name":"Cell-Free Systems","description":"Utilize cell-free platforms that enable synthetic biology outside of living cells, thereby bypassing the limitations of evolved cellular machinery."},{"id":"1c1cb37e-2a00-8086-aa50-cf6d74a26b7e","name":"Designer Enzymes from Quantum Chemistry and Protein Design","description":"Create enzymes specifically engineered via quantum chemical methods and de novo protein design, which can precisely catalyze reactions at defined positions."},{"id":"1c1cb37e-2a00-8043-a0a8-ed6dec92fbca","name":"Direct 3D Specification of Protein-Like Molecules","description":"Design polymers that are not limited to amino acids and are  directly specified in three dimensions, enabling precise positional control in synthesis and potentially broader or more robust functions than proteins."},{"id":"1c1cb37e-2a00-8050-b4e3-c5ccf7cfa171","name":"Vacuum Mechanosynthesis Exploration","description":"Investigate the feasibility of vacuum mechanosynthesis—a process that uses mechanical forces under vacuum conditions to construct molecules with high positional precision."},{"id":"1c1cb37e-2a00-8042-acea-fe39a87e2605","name":"Molecular 3D Printing","description":"Explore methods for molecular-scale 3D printing, which would enable the precise assembly of molecules layer by layer.\n \nThis would in principle move us towards a general-purpose approach to atomically precise fabrication as envisioned by Drexler in the 1980s and Feynman in the late 1950s. DNA origami made a leap in 2006, but DNA is in some key ways a much less precise and versatile nanoscale building material than proteins/peptides. A promising path would extend “DNA origami” to “protein carpentry” by adapting Beta Solenoid proteins, or other modular protein components with programmable binding properties, as lego-like building blocks and then using the latter to construct massively parallel protein-based 3D printers for lego-like covalent assembly of a restricted set of chemical building blocks. This one is riskier: how programmably can we really control protein assembly, and could we bootstrap from initial crappy prototype protein-carpentry-and-or-DNA-origami-based molecular 3D printers to genuinely useful ones? \n\nSafety consideration: https://iopscience.iop.org/article/10.1088/0957-4484/15/8/001 \n\nStrategy consideration: https://www.effectivealtruism.org/articles/ea-global-2018-paretotopian-goal-alignment "},{"id":"1c1cb37e-2a00-80b8-a789-e201a5975aad","name":"Silicon-Protein Interfaces","description":"Current fabrication methods allow us to work at macroscopic scales (10^0 m) down to the nanometer scale (10^-8 m) with photolithography, and further down to the atomic scale (10^-10 m) with proteins. However, directly bridging from macroscopic to atomic scales (10^0 m to 10^-10 m) for nanotechnology applications remains a significant challenge. A key obstacle is the lack of effective interfaces between single addressable electrodes and proteins."}]},{"id":"1c1cb37e-2a00-809a-a069-c93e72813eb6","slug":"current-chip-fabrication-methods-are-extremely-expensive-and-hard-to-change","name":"Current Chip Fabrication Methods are Extremely Expensive and Hard to Change","description":"Modern chip fabs are enormous, multi-billion-dollar facilities with limited versatility in what they can produce. This bottleneck restricts the ability to create assemblies with diverse molecular components on a small scale.","field":"Nanoscale Fabrication","outcome":"Nanoscale device research stops being rationed by access to a handful of multi-billion-dollar fabs. Small batches and odd materials become fabricable.","outcome_rationale":"Their description names versatility and small-scale assembly as what current fabs foreclose.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Capital cost per wafer start, minimum economic batch size and feature size achievable are direct facility metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"Speculative","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Speculative). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Physical build","primary_maturity":"Speculative","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Speculative). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80f6-98f0-eb07caafe9b7","name":"Expand the Scope of Chip Fabs","description":"Broaden the capabilities of existing chip fabrication facilities to produce a wider range of devices, potentially reducing costs and expanding functionality."},{"id":"1c2cb37e-2a00-802f-8eac-c62f94af69c3","name":"Shrink Molecularly Diverse 3D Assemblies for Fabrication","description":"Develop methods for producing smaller, more diverse assemblies using patterned techniques that allow for greater molecular variability."},{"id":"1c2cb37e-2a00-808f-ad46-e2a8407062c1","name":"Nano Modular Electronics","description":"Develop modular electronic components at the nanoscale, enabling flexible, low-cost assemblies with high molecular diversity."},{"id":"1c2cb37e-2a00-8083-bb51-e286cbd1c06b","name":"Randomized DNA-Based Chip Assembly","description":"Explore methods to assemble chips using randomized DNA as a templating or assembly tool, allowing for scalable production with inherent molecular diversity."},{"id":"1c2cb37e-2a00-80e5-8b02-c75fdf251d66","name":"Digital Assembly","description":"Use digital techniques to plan and assemble chip components, integrating computational design with physical fabrication."},{"id":"1c2cb37e-2a00-80e1-97d4-e0a987ddea8b","name":"Laser Direct Write","description":"Employ laser direct writing techniques to pattern chips with high precision, offering an alternative to traditional lithographic methods."}]},{"id":"1c1cb37e-2a00-80cb-8db6-e28439779875","slug":"searching-through-the-vast-underexplored-space-of-materials-is-slow-and-expensive","name":"Searching Through the Vast, Underexplored Space of Materials is Slow and Expensive","description":"“New materials create fundamentally new human capabilities. And yet…new materials-enabled human capabilities have been rare in the past 50 years.” The core challenge lies in our inability to reliably design and manufacture materials that meet specific engineering requirements–and to do so at an industrial scale and reasonable cost. \n\nIdentifying promising new materials is hampered by the slow pace of exploration. The integration of machine learning, physics-based property prediction, and self-driving laboratories could dramatically accelerate this process. A significant opportunity lies in modeling the vast, unexplored space of potential materials in silico.","field":"Materials Science","outcome":"Specify the engineering property you need and find the material that has it, at industrial cost and scale. Today the order runs the other way.","outcome_rationale":"Their description states the core challenge as inability to reliably design and manufacture materials meeting specific engineering requirements.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Candidates screened per unit time, hit rate against a property specification and cost per validated material are all countable.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (type: Design search vs Running experiments; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Running experiments","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (type: Design search vs Running experiments; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[{"id":2,"quantity":"Novel inorganic compounds experimentally realised per day of autonomous laboratory operation","current_value":"2.4","unit":"compounds per day","as_of":"2023-11-29","target_value":null,"target_basis":null,"source_title":"An autonomous laboratory for the accelerated synthesis of novel materials","source_url":"https://www.osti.gov/biblio/2281696","source_doi":"10.1038/s41586-023-06734-w","source_checked":"verified","reads_as":"An autonomous lab made 41 new inorganic compounds in 17 days of continuous running, which is 2.4 a day.","direction":"higher is better","context":"The same run attempted 21 experiments a day at a 71% success rate, with no chemist at the bench.","caveat":"A reanalysis by Palgrave and Schoop argues most of the 41 were ordered versions of compounds already known. The rate is solid; what counts as a new compound is contested.","is_null_result":0,"rationale":"The A-Lab paper reports 41 novel compounds realised from 58 targets over 17 days of continuous operation, which is 2.4 per day. The rate, not the total, is the thing to watch: Convergent's gap is that the search is slow, and this is the only published figure that puts experimentally validated materials on a clock. Marked a guess for a specific and important reason, not as hedging: a reanalysis by Robert Palgrave (UCL) and Leslie Schoop (Princeton), reported by Chemistry World on 16 January 2024, argues that most of the 41 were ordered versions of already-known disordered compounds and that the automated Rietveld refinement was inadequate — Palgrave's summary is that 'it's likely they didn't make any discoveries'. Gerbrand Ceder's response was that the demonstration was of autonomous capability rather than of refinement quality. The denominator of this indicator is therefore solid and the numerator is contested, which is worth more to a reader than a clean number would be. No target: nobody has published a compounds-per-day goal, and the honest reading of the dispute is that the field does not yet agree on what counts as one compound.","confidence":"guess"}],"capabilities":[{"id":"1c2cb37e-2a00-8040-a2d0-dc1a77a1282c","name":"ML & Physics-Based Property Prediction and Iterative Self-Driving Lab","description":"Leverage machine learning models combined with physics-based property prediction to iteratively explore the materials space using automated, self-driving laboratory platforms, to find things like higher temperature superconductors or topological materials.\n \nNew designs are needed to minimize large capital expenditures and integrate flexible, modular components that can be rapidly repurposed for new experiments and are robust to variations and error handling. "},{"id":"1c2cb37e-2a00-809b-8afb-c1baea4e399a","name":"Build Better Assay Platforms for Materials","description":"Develop more efficient assay platforms to test the properties of materials, thus enabling faster feedback and iteration in materials discovery."},{"id":"1c8cb37e-2a00-80fa-b355-c60a4435c377","name":"Materials Property Bank","description":"Large open dataset of experimentally determined mechanical, thermal, electrical properties of millions of samples that consolidates published and crowdsourced data to enable ML models.\n\nThis would augment initiatives like the Materials Project and OQMD, which are simulation-heavy. "},{"id":"1f0cb37e-2a00-807f-ba7a-f6e827399665","name":"In-Silico Modeling of Material Space","description":"Utilize deep learning and computational modeling to predict and discover millions of new materials, expanding our understanding of what can exist."}]},{"id":"1c1cb37e-2a00-809e-8e26-e1c9f14f5143","slug":"many-molecules-cant-easily-be-crystallized","name":"Many Molecules Can’t Easily Be Crystallized","description":"Crystallization is crucial for determining molecular structure, yet many molecules resist forming crystals. Improved computational models of crystal growth are needed to guide experimental efforts.","field":"Materials Science","outcome":"Crystallisation gets planned instead of attempted. The trial-and-error step that currently gates structure determination for hard molecules goes away.","outcome_rationale":"Their description names computational models of crystal growth as the route to guiding experiment.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Crystallisation success rate and predicted-versus-observed polymorph are directly scorable, with blind tests already run in the field.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"Working now","is_primary":0,"rationale":"High-throughput crystallisation screening is already automated in structural biology facilities.","confidence":"confident"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80d2-9452-e0c2e86eb31f","name":" Computational Crystal Growth","description":"Develop integrated computational frameworks that combine quantum chemistry, reaction potential modeling, experimental data, and ML generative models to describe and direct crystal growth at the atomic level."}]},{"id":"1c1cb37e-2a00-80f5-b3d1-d19cab7aeec3","slug":"limited-ability-to-design-and-scalably-synthesize-macroscale-materials","name":"Limited Ability to Design and Scalably Synthesize Macroscale Materials","description":"While many promising materials have been discovered in the lab, current synthesis methods are often too expensive to produce these materials in sufficient quantities. Some examples of novel materials that would be highly enabling include: \n\n• Low activation, thermally conductive materials that are resistant to radiation damage are needed to enable fusion reactors (the first wall material is currently a limitation), spacecraft, etc.\n• Materials that emit at the transparency window of the atmosphere (that were easy to apply like paint) to drastically diminish solar earth heating (example)\n• Hyper-efficient thermoelectrics that could directly turn heat into electricity. \n• Materials that autonomously heal to improve our infrastructure and prevent system failures due to material defects.\n• Materials that have the insulating properties and high melting points of ceramics but the formability and ductility of metals for jet engines and atmospheric reentry vehicles.","field":"Materials Science","outcome":"Fusion first walls, radiative cooling coatings and high-temperature ductile ceramics get made in tonnes rather than in grams.","outcome_rationale":"Their description lists concrete blocked applications and names cost of scaled synthesis, not discovery, as the obstacle.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Cost per kilogram at a stated property specification is the direct and decisive quantity.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":0,"rationale":"Process-parameter optimisation for scale-up is an iterative make-and-measure loop that closed-loop platforms can run.","confidence":"confident"},{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Physical build vs Design search; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Physical build vs Design search; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80a5-8f6c-d94fb6d29f51","name":"Novel Design and Processing of Textile Fibers","description":"Next generation, high performance protein-based fibers created through new spinning processes that can align and control the molecular assembly of the final fiber, taking advantage of protein’s unique capabilities."},{"id":"1c2cb37e-2a00-805c-bb14-d715232049dc","name":"Macroscale Bioprinting","description":"Methods to build large-scale structures from cells and proteins."},{"id":"1c2cb37e-2a00-8015-b5c6-f9de263ac76b","name":"Scalable Synthesis of Carbon-Based Materials ","description":"Carbon fiber could potentially replace steel in many situations and sequester atmospheric carbon instead of creating it if we could make enough of it cheaply enough.\n\nArbitrarily long carbon nanotubes would enable tethers with tensile strength near the limits of physics which unlock things like space elevators.Scaling the production of conductive carbon materials could potentially replace copper."}]},{"id":"1c1cb37e-2a00-80b4-a5fa-cd97b7b1bed6","slug":"we-have-a-limited-ability-to-acquire-concentrate-and-substitute-chemical-elements-in-processes","name":"We Have a Limited Ability to Acquire, Concentrate and Substitute Chemical Elements in Processes","description":"The cost of materials is often dominated by the cost to obtain their constituent elements. What presents commercially as the “critical minerals problem” masks a larger scientific bottleneck on how we acquire, concentrate, and substitute chemical elements.","field":"Materials Science","outcome":"Separation gets cheap and constrained elements get designed around. What a material costs stops being set by what its elements cost.","outcome_rationale":"Their description reframes the critical minerals problem as a scientific bottleneck in acquisition, concentration and substitution.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Separation factor, energy per kilogram and magnet energy product without neodymium are direct performance quantities.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":0,"rationale":"Lanthanide separation ligand screening is a bench-scale formulation loop of exactly the kind closed-loop platforms run.","confidence":"confident"},{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8060-ad08-d56f1148c660","name":"Efficient Chemical Separation of Lanthanides","description":"Create an industrial center of excellence focused on the practical separation of Lanthanides to distribute this knowledge as a public good."},{"id":"1c2cb37e-2a00-80bd-9452-d7bced369af6","name":"High-Strength, Non-Neodymium Magnets","description":"Develop high-strength permanent magnets not made of rare-earth elements. Currently the high-strength magnets underpinning many technologies (e.g., hard disk drives, mobile phones, electric vehicle motors) are all made out of neodymium, a rare earth element at risk of supply chain shortages and environmental issues. "},{"id":"1c2cb37e-2a00-8062-901b-f19fc4788be7","name":"Recovering Metals and Rare Minerals from Waste","description":"E-waste represents a significant opportunity to recapture and reuse rare earth elements that are in short supply. "}]},{"id":"1c1cb37e-2a00-80a7-86dd-d2672b519acf","slug":"designing-manufacturing-systems-is-hard","name":"Designing Manufacturing Systems is Hard","description":"Modern manufacturing system design remains complex, with traditional methods relying on outdated processes. AI-based design approaches have the potential to reimagine these systems without relying on the legacy of humanoid robots.","field":"Mechanical Engineering","outcome":"A new production line gets designed around the task it has to do, rather than inheriting the layout of the last one.","outcome_rationale":"Their description explicitly rejects designing around humanoid-robot legacy assumptions.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Throughput, capital cost per unit capacity and changeover time are the standard manufacturing figures.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8045-bcb1-e3ba89e6e21d","name":"AI-Based Design of Manufacturing Systems","description":"Develop AI-driven design tools that optimize manufacturing systems from the ground up, moving beyond traditional approaches to enable more efficient, automated production processes."}]},{"id":"1c1cb37e-2a00-80c4-b958-d7161b660b3e","slug":"designing-buildings-is-hard","name":"Designing Buildings is Hard","description":"Architectural design and construction planning are complex and labor-intensive. Advanced computational design and AI-driven optimization have the potential to revolutionize how buildings and construction plans are generated.","field":"Mechanical Engineering","outcome":"Architects can explore many building designs cheaply and pick on cost, energy and buildability, instead of settling the design early and living with it.","outcome_rationale":"Their description names labour-intensity as the constraint that computational design would lift.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Design hours per project, alternatives evaluated and resulting cost or energy performance are all directly recorded.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80b0-a4b4-eebe80189d9d","name":"AI Design of Buildings and Construction Plans","description":"Utilize generative AI to automatically generate and optimize building designs and construction plans, streamlining the design process and reducing manual effort."}]},{"id":"1c1cb37e-2a00-8039-a26b-ea2d3b4f9bd6","slug":"bioengineering-is-still-done-manually","name":"Bioengineering is Still Done Manually","description":"Despite advances in automation, many bioengineering processes remain highly manual, limiting throughput and reproducibility in laboratory settings.","field":"Mechanical Engineering","outcome":"Biology experiments run without trained hands at the bench. Throughput rises, and results are reproducible by construction.","outcome_rationale":"Their description names throughput and reproducibility as the two losses from manual work.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Experiments per instrument-day, human hours per experiment and replicate variance are direct operational metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"Working now","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 Working now / v2 2-5 years; mechanical rule had given 2-5 years). Biofoundries and cloud labs — Ginkgo, the DOE Agile BioFoundry, Emerald — run design-build-test-learn cycles on strains today, and the throughput gain over manual work is the gap's own stated loss. Same standing as automated chemical synthesis, which both passes called Working now. Type unchanged from the mechanical adjudication.","confidence":"confident"}],"primary_ai_type":"Running experiments","primary_maturity":"Working now","primary_rationale":"Maturity re-adjudicated on the merits (v1 Working now / v2 2-5 years; mechanical rule had given 2-5 years). Biofoundries and cloud labs — Ginkgo, the DOE Agile BioFoundry, Emerald — run design-build-test-learn cycles on strains today, and the throughput gain over manual work is the gap's own stated loss. Same standing as automated chemical synthesis, which both passes called Working now. Type unchanged from the mechanical adjudication.","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80f6-a207-db619471de63","name":"AI-Assisted Programming of Lab Robots","description":"Implement AI-assisted systems for programming and controlling lab robots to automate bioengineering workflows and reduce the reliance on manual processes."},{"id":"1c2cb37e-2a00-8047-9af0-c0787f96115e","name":"Bio Lab of the Future","description":"Develop an integrated \"bio lab of the future\" that combines robotics, AI, and real-time data analysis to fully automate bioengineering tasks."}]},{"id":"1c1cb37e-2a00-8007-8f03-ea967970cb8f","slug":"modeling-mechanical-systems-is-hard","name":"Modeling Mechanical Systems is Hard","description":"The simulation and modeling of complex mechanical systems is challenging due to the intricate interplay of multiple physical phenomena. Improved computational models can enhance design and optimization.","field":"Mechanical Engineering","outcome":"Mechanical design iterates in simulation rather than in hardware, because coupled multiphysics behaviour can be predicted without a bespoke model per system.","outcome_rationale":"Their capability is a transferrable multiphysics foundation model, and transferability is the whole point.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Prediction error against high-fidelity simulation or experiment at a stated compute budget is a direct benchmark.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"Working now","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 Working now / v2 2-5 years; mechanical rule had given 2-5 years). Learned surrogates for CFD and structural analysis ship in commercial engineering tools and cut design-iteration time on exactly the coupled-multiphysics problems the gap names. Flagged guess because they do not yet stand in for high-fidelity solvers where certification is required. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"Working now","primary_rationale":"Maturity re-adjudicated on the merits (v1 Working now / v2 2-5 years; mechanical rule had given 2-5 years). Learned surrogates for CFD and structural analysis ship in commercial engineering tools and cut design-iteration time on exactly the coupled-multiphysics problems the gap names. Flagged guess because they do not yet stand in for high-fidelity solvers where certification is required. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80ad-8cef-fa475ca23767","name":"Transferrable Multiphysics Foundation Model","description":"Develop a comprehensive multiphysics foundation model that can be applied across various mechanical systems, integrating physics-based models, experimental data, and machine learning for broad transferability."}]},{"id":"1c1cb37e-2a00-8050-bea1-ffeab17c0e7b","slug":"robot-hardware-and-software-is-still-clunky","name":"Robot Hardware and Software is Still Clunky","description":"Robots have the potential to revolutionize manufacturing, logistics, and many other industries—but only if they are both affordable and capable of high performance. Today’s robotic hardware is often prohibitively expensive and built using legacy designs that do not prioritize cost reduction, modularity, or scalability. Moreover, many robots struggle with dexterity and tactile sensing, and current design practices decouple hardware and software, preventing a co-evolution that could unlock new performance regimes. Overcoming these limitations requires a rethinking of both robot morphology and control, with an emphasis on integrated design, cost-effective production, and enhanced functionality.\r\n","field":"Mechanical Engineering","outcome":"Robots get cheap and dexterous at the same time. Physical tasks currently reserved for human hands become worth automating.","outcome_rationale":"Their description makes affordability and capability jointly necessary, and names dexterity and tactile sensing as the specific shortfalls.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Cost per degree of freedom, standard dexterity benchmark scores and mean time between failures are directly comparable.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"2-5 years","is_primary":0,"rationale":"Their co-design of morphology and control capability is a joint search over body and controller.","confidence":"confident"},{"ai_type":"Physical build","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Physical build","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80e0-91bd-c2d4a563a09f","name":"Dexterous Robots","description":"Create advanced robot bodies that leverage novel materials, innovative design principles, and improved manufacturing and control techniques to enhance dexterity while reducing cost."},{"id":"1c2cb37e-2a00-80b7-bd46-e8c46c6476f0","name":"Non-Humanoid Mobile Robots","description":"Design mobile robots with streamlined, efficient designs. Reduce unnecessary degrees of freedom for simpler, cheaper, and more reliable mobility platforms."},{"id":"1c2cb37e-2a00-8008-b97e-d82df77ffaca","name":"Underactuated Robots","description":"Advance the use of underactuated robotic systems, which use fewer actuators than degrees of freedom, to create compliant, adaptable, and significantly cheaper robots. "},{"id":"1c2cb37e-2a00-8005-a5f5-e96e732da83a","name":"Nanoscale Robots","description":"Employ lithographic techniques and advanced nanofabrication to create tiny robots with high precision. Scalable production of nanoscale robotic systems could enable breakthroughs in medicine (e.g., targeted drug delivery) and materials science."},{"id":"1c2cb37e-2a00-8078-b591-f7c5100dc41c","name":"Improved Robot Actuation","description":"Develop novel actuator technologies that combine hybrid or mode-switching capabilities with power-dense magnetic actuation. These improvements would allow robots to seamlessly transition between compliant and stiff modes, mimicking biological muscle performance while enhancing energy efficiency."},{"id":"1c2cb37e-2a00-8000-9598-f1314fecc2ff","name":"Co-Design of Morphology and Control","description":"Integrate hardware and software design processes to co-evolve robot bodies alongside control policies. This approach reduces inefficiencies caused by decoupled design methods and can unlock entirely new performance regimes."},{"id":"1c2cb37e-2a00-8027-99bd-e307904f0f84","name":"Robot Tactile Perception","description":"Develop multimodal electronic skin (e-skin) that enables robots to detect force vectors, slippage, and temperature across large surface areas. Enhanced tactile perception will facilitate fine-grained control and more adaptive interactions with the environment."}]},{"id":"1c1cb37e-2a00-8018-8a8c-f42f50d9229a","slug":"many-methods-are-stuck-in-20th-century-fabrication-paradigms","name":"Many Methods Are Stuck in 20th Century Fabrication Paradigms","description":"Modern manufacturing systems largely rely on paradigms developed in the last century where large machines produce components smaller than themselves. This approach is increasingly limited by scaling challenges and cost inefficiencies. To meet future demands, we need to reimagine manufacturing by developing universal robotic construction systems and low-capital, high-energy manufacturing solutions that leverage emerging technologies such as advanced robotics, precision machining, and renewable energy integration. These innovations could, for example, dramatically lower the cost of machining high-performance materials like titanium or enable widespread automation in sectors like desalination.","field":"Mechanical Engineering","outcome":"Manufacturing no longer needs machines bigger than the parts they make, so high-performance fabrication can be set up cheaply and close to where it is used.","outcome_rationale":"Their description names the paradigm precisely and gives titanium machining and desalination as blocked examples.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Capital cost per unit capacity and cost per part in a named material are direct and comparable.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Physical build","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80e6-b691-e5ba37ae3436","name":"Universal Robotic Construction Systems","description":"Build the robots that build the robots. \n\nThe current paradigm of all automated manufacturing is for machines (from robot arms to presses) to make things smaller than themselves. This quickly runs into scaling limits etc. "},{"id":"1c2cb37e-2a00-80c7-a4cf-ee6416d47604","name":"Cheap, Automatic Five Axis Electrochemical Machining ","description":"Electrochemical machining (ECM) creates complex shapes with high precision, but is costly. Lower cost ECM could make machined titanium as cheap as aluminum."},{"id":"1c2cb37e-2a00-8011-9ab5-eb6b4c9859df","name":"Automated Defect and Weld Inspection","description":"Automation of welding requires improved ability to verify welds."},{"id":"1c2cb37e-2a00-804c-98ce-ea4a6aabab5f","name":"Low-Capital Manufacturing Systems ","description":"If we could create more manufacturing systems with low capex but high energy needs, we could take advantage of drastically cheaper solar. \n\nOne particularly useful example would be desalination."}]},{"id":"1c1cb37e-2a00-80c5-921e-de69ffaca475","slug":"sim-to-real-transfer-for-robots-is-hard","name":"Sim-to-Real Transfer for Robots is Hard","description":"Bridging the gap between simulated robot behavior and real-world performance remains a significant challenge, particularly for tactile interactions and complex environments.","field":"Mechanical Engineering","outcome":"A policy trained in simulation works on real hardware first time. Robot capability stops being gated by the cost of collecting real-world data.","outcome_rationale":"Their description identifies the simulation-to-reality boundary as the specific point of loss, especially for contact and tactile interaction.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Task success rate on real hardware versus in simulation is a direct, standard and unambiguous comparison.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":0,"rationale":"Tactile perception is the modality where the reality gap is worst and where data collection is the named capability.","confidence":"confident"},{"ai_type":"Physical build","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Prediction and modeling vs Physical build; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Physical build","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Prediction and modeling vs Physical build; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8038-8f26-f48b3dfe8c12","name":"Combine Physics Models and Generative AI","description":"Integrate detailed physics models with generative AI techniques to improve the accuracy of simulations and facilitate effective transfer of robotic behavior from simulation to reality."},{"id":"1c2cb37e-2a00-8000-95c5-dcc4265f3cf9","name":"Large-Scale Data Collection for Tactile Interaction","description":"Systematically collect and curate large training datasets focusing on tactile interactions to enhance simulation accuracy and real-world performance."},{"id":"1c2cb37e-2a00-80ca-b646-e604885bc8d7","name":"Robot Foundation Model","description":"Robot foundational models can address major obstacles in robot learning and enable training on action-free data including video. This is essential for enabling reasoning about novel situations and robustly handling real-world variability."}]},{"id":"1c1cb37e-2a00-80ed-8131-ede48bd0d654","slug":"silicon-compute-is-massively-energy-intensive-compared-to-biological-brains","name":"Silicon Compute is Massively Energy Intensive Compared to Biological Brains","description":"Modern deep learning and general computation demand enormous energy, limiting scalability and sustainability. Addressing energy efficiency is critical for the next generation of computing platforms, though it also supports potential proliferation of advanced AI and should be advanced alongside AI safety and governance considerations.","field":"Computation","outcome":"Energy stops setting the ceiling on how much inference and simulation science can afford to run. Orders of magnitude more compute per joule.","outcome_rationale":"Their description names scalability and sustainability as the losses, and energy is the shared constraint across all eleven capabilities.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Joules per operation at a stated workload is a direct and universally reported figure.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"Speculative","is_primary":0,"rationale":"Neuromorphic, reversible and thermodynamic computing all require fabricating device physics that no production process supports.","confidence":"confident"},{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Learned floorplanning and architecture search are deployed and deliver percent-scale efficiency gains. The gap sets its target against biological brains, which is orders of magnitude, and no design-search application applied today approaches that. Type unchanged from the mechanical adjudication.","confidence":"confident"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Learned floorplanning and architecture search are deployed and deliver percent-scale efficiency gains. The gap sets its target against biological brains, which is orders of magnitude, and no design-search application applied today approaches that. Type unchanged from the mechanical adjudication.","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80c4-b7fa-fd22d4e73b7a","name":"Lower Energy Architectures for Deep Learning","description":"Develop novel hardware architectures optimized for deep learning and artificial intelligence that dramatically reduce energy consumption compared to current systems."},{"id":"1c2cb37e-2a00-8098-9702-c91a8180fbcf","name":"Neuromorphic Systems","description":"Leverage brain-inspired neuromorphic hardware to perform computation more efficiently, emulating the low-energy operation of biological neural networks."},{"id":"1c2cb37e-2a00-80ae-9f20-ffb59fd13825","name":"Superconducting Virtual Brains","description":"Explore superconducting hardware to achieve brain-inspired computing with drastically reduced energy consumption, scaling to large networks."},{"id":"1c2cb37e-2a00-80a8-b5aa-caaad648d2d7","name":"Millivolt Switching","description":"Develop switching technologies that operate at millivolt levels, significantly reducing the energy required for signal processing and computation."},{"id":"1c2cb37e-2a00-8041-b5eb-c43fda8642aa","name":"Reversible Computing","description":"Create computing architectures that use reversible logic, theoretically allowing computation with near-zero energy dissipation by avoiding information loss."},{"id":"1c2cb37e-2a00-8033-b4c2-e928ef582b51","name":"Probabilistic Computing Hardware","description":"Hardware for probabilistic computation, which can perform certain tasks more energy-efficiently by embracing uncertainty."},{"id":"1c2cb37e-2a00-8004-92d5-e1eea5e05513","name":"Thermodynamic Computing","description":"Computing paradigms based on thermodynamic principles, where computation is driven by energy gradients and can operate at lower energy costs by harnessing reversible and low-energy processes."},{"id":"1c2cb37e-2a00-8056-9a63-d5a196b591c1","name":"Living Computers","description":"Biologically inspired or living computer systems that use biological components to perform computation at very low energy levels."},{"id":"1c2cb37e-2a00-80f7-9469-fe53fb428ee5","name":"AI-Based Design of AI Chips","description":"Use AI to design the next generation of hardware for AI."},{"id":"1c2cb37e-2a00-801f-bf91-fb2cfe757da1","name":"Next-Gen 3D Integration of Discrete Components","description":"New logic and memory technologies based on CNTs or other structures, 3D integration with fine-grained connectivity, and new architectures for computation immersed in memory"},{"id":"1f0cb37e-2a00-8031-8696-d662eb107ba5","name":"Denser and More Robust Longer Term Data Storage","description":"Enable new modalities of long term, ultra dense data storage such as via DNA polymers"}]},{"id":"1c1cb37e-2a00-8012-aeb4-c0c2e9c20f46","slug":"proving-math-theorems-is-challenging-for-both-humans-and-ai","name":"Proving Math Theorems is Challenging for Both Humans and AI","description":"Both human mathematicians and current AI systems struggle with proving complex math theorems. Enhancing theorem proving through interactive and automated methods could push the boundaries of mathematical reasoning.","field":"Computation","outcome":"Formal verification stops costing more than the proof it verifies, and machine-assisted proof turns routine at research level.","outcome_rationale":"Their capability is reinforcement learning from theorem-prover feedback, aimed at making proof search practical.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"A proof either checks in a formal system or it does not, and benchmark theorem sets with formalisation rates already exist.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"Working now","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"Working now","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80f2-a6e3-e4b3d8d9f275","name":"Reinforcement Learning from Interactive Theorem Prover Feedback","description":"Close the loop between AI and human-guided interactive theorem provers by using reinforcement learning to refine proofs based on feedback from proof assistants."}]},{"id":"1c1cb37e-2a00-80a7-ad9a-c209405e462d","slug":"ai-could-go-rogue","name":"AI Could Go Rogue","description":"The potential for AI systems to behave unpredictably or dangerously (“go rogue”) is a critical concern. Ensuring safe and controllable AI architectures is essential for reliable operation. \n\nSee also: \n• https://www.lesswrong.com/posts/fAW6RXLKTLHC3WXkS/shallow-review-of-technical-ai-safety-2024\n• https://deepmind.google/discover/blog/taking-a-responsible-path-to-agi/","field":"Computation","outcome":"You can check a claim about an AI system against its internals, rather than infer it from behaviour on tests the system may recognise.","outcome_rationale":"Their capability list leads with automated interpretability and guaranteed-safe architectures, both aimed at making assurance verifiable.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Evaluations and interpretability probes exist and are run, but there is no agreement in the field that any of them would settle whether a system is safe. This is a verifier problem rather than a measurement problem.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":0,"rationale":"Hardware governance appears in their own capability list and is a control-and-allocation mechanism, not a technique.","confidence":"confident"},{"ai_type":"Reading and synthesis","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80dc-ba94-eca26e0d568a","name":"Guaranteed Safe AI Architectures","description":"Develop and implement AI architectures with separable, auditable world models; where safety can be specified in terms of the state space of the model; and proposed AI outputs come with proofs that the output does not leave the safe region of the world model’s state space."},{"id":"1c2cb37e-2a00-8014-a3c5-e93d5674bf09","name":"Understanding Neural Design Principles of Social Instincts","description":"Study the neural basis of human social instincts to inform AI design, ensuring that AI systems can safely interpret and emulate human social behavior."},{"id":"1c2cb37e-2a00-8036-895f-f2b612bcc2d9","name":"Automate AI Interpretability","description":"Use AI to enhance the interpretability of other AI systems, creating tools that automatically explain and verify AI behavior."},{"id":"1c2cb37e-2a00-808b-bcce-d83a3801b683","name":"Hardware Governance","description":"Develop hardware-level governance mechanisms to enforce safety and compliance in AI systems, ensuring robust operational constraints. This includes tamper-proof hardware. "},{"id":"1c2cb37e-2a00-80af-a74e-e6a053db7a32","name":"Understand AI Psychology without Assuming Human-Like Psychology","description":"Observing emergent AI decision-making processes and cognitive patterns with fewer anthropomorphic assumptions."},{"id":"1c2cb37e-2a00-8038-ac41-da6272887452","name":"Mitigating (Indirect) Data Poisoning","description":"Robust strategies for data integrity, anomaly detection, and defensive training protocols to mitigate situations where indirect data poisoning could lead to intentionally misaligned AI systems (not unlike “sleeper agents”)."},{"id":"1c2cb37e-2a00-8074-b15a-dfa64892aeeb","name":"Secure and Privacy-Preserving Local AI Enclaves","description":"Digital fortresses that enable sensitive data to be processed in a controlled, privacy-preserving environment. "},{"id":"1f0cb37e-2a00-80ef-9202-fda6711ee9f5","name":"Non-Agentic AI Scientists","description":"Build non-agentic AI scientists that act as oracles but don’t have long-term states or goals "}]},{"id":"1c1cb37e-2a00-800b-8598-ea123934c801","slug":"ai-could-be-misused","name":"AI Could Be Misused","description":"The risk of AI being misused—whether through malicious intent or unintended consequences—necessitates robust safeguards and countermeasures.","field":"Computation","outcome":"Releasing a capable model stops meaning releasing its worst uses along with it.","outcome_rationale":"Their two capabilities target robustness and decentralised control, both aimed at decoupling capability from misusability.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Jailbreak success rates on benchmark suites are measurable, but misuse actually averted is unobserved, so the gap itself is measured only through adjacent indicators.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"Working now","is_primary":0,"rationale":"Adversarial robustness training and automated red-teaming are working techniques already in production.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"Speculative","is_primary":1,"rationale":"Maturity re-adjudicated twice. v1 said 2-5 years, v2 said Working now, the mechanical rule took Working now, and the maturity repair moved it to 2-5 years on the grounds that evaluation regimes and deployment policy exist as mechanisms even if none is in force at a scale that constrains misuse. David ruled Speculative on 2026-08-25 and the ruling stands: misuse has many partial mitigations, but the one that would actually close this gap is alignment, and a substantial part of the field holds that it may not be solvable at all. A capability whose feasibility is itself contested is what Speculative is for. Convening on AI safety is available this afternoon; that is the availability reading of 'working now' the repair was built to reject. Type unchanged.","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"Speculative","primary_rationale":"Maturity re-adjudicated twice. v1 said 2-5 years, v2 said Working now, the mechanical rule took Working now, and the maturity repair moved it to 2-5 years on the grounds that evaluation regimes and deployment policy exist as mechanisms even if none is in force at a scale that constrains misuse. David ruled Speculative on 2026-08-25 and the ruling stands: misuse has many partial mitigations, but the one that would actually close this gap is alignment, and a substantial part of the field holds that it may not be solvable at all. A capability whose feasibility is itself contested is what Speculative is for. Convening on AI safety is available this afternoon; that is the availability reading of 'working now' the repair was built to reject. Type unchanged.","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80df-abe5-df20e6fedd75","name":"Jailbreak Resistance","description":"Develop techniques to ensure AI systems are resistant to \"jailbreaking\" or circumvention of built-in safety and control protocols."},{"id":"1c2cb37e-2a00-80ec-81b3-fa62f2466b2f","name":"Decentralized Training","description":"Maintain decentralized control over large neural network training, like SETIatHome/Wikipedia for large language models, to equalize access. However, this also introduces some AI proliferation and control related risks."}]},{"id":"1c1cb37e-2a00-80b3-aab7-d6e108c54123","slug":"our-compute-stack-is-insecure-but-fundamentally-doesnt-have-to-be","name":"Our Compute Stack is Insecure but Fundamentally Doesn’t Have To Be","description":"Insecure software can lead to vulnerabilities that undermine the reliability and safety of computational systems. Formal methods and rigorous verification are needed to synthesize secure software.","field":"Computation","outcome":"Whole classes of vulnerability stop being possible, rather than being found and patched one at a time.","outcome_rationale":"Their description names formal methods and verification as the route, and the unlock is elimination rather than detection.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Proportion of a critical codebase carrying machine-checked proofs, and vulnerability counts in verified versus unverified components, are directly countable.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80fa-bc8b-e009571ab054","name":"Formally Verified Software Synthesis","description":"Utilize formal verification techniques to synthesize software that is provably secure, reducing vulnerabilities and enhancing system robustness. Traditional programs in these areas have tended to assume that AI capabilities are saturating, leaving important avenues neglected. \n\nEfforts could break down into several key components: \n• Specification generation and validation tools\n• Code and proof generation systems\n• Tools that integrate formal verification into existing engineering workflows \n• Practical formalization structures that facilitate real-world adoption"}]},{"id":"1c1cb37e-2a00-8072-a36f-f7ea8adace81","slug":"ai-is-still-narrow-in-its-reasoning-and-planning","name":"AI is Still Narrow in its Reasoning and Planning","description":"Current AI systems exhibit narrow reasoning and planning capabilities compared to human cognition. Broadening AI training methods to include holistic, brain-inspired architectures and cognitive frameworks can advance general intelligence (flagging that there is an AI safety risk here).","field":"Computation","outcome":"An AI system handles a domain it was not trained on without someone rebuilding the scaffolding around it.","outcome_rationale":"Their description contrasts narrow current capability with human cognition and names general intelligence as the target.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Reasoning benchmarks exist and are run constantly, but the field openly disputes whether benchmark performance measures general reasoning or contamination and pattern coverage. A measurement can always be taken; what it shows is what is argued about.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"2-5 years","is_primary":1,"rationale":"Maturity decided by David, 2026-08-23. Speculative means no clear path, and the narrowness is receding quickly enough that the claim will not survive contact with a reader. Recorded as contested: David notes many informed people would say otherwise, and the honest label may be a 5-10 year bucket this taxonomy does not have. See research-log/maturity-repair/notes.md.","confidence":"confident"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"2-5 years","primary_rationale":"Maturity decided by David, 2026-08-23. Speculative means no clear path, and the narrowness is receding quickly enough that the claim will not survive contact with a reader. Recorded as contested: David notes many informed people would say otherwise, and the honest label may be a 5-10 year bucket this taxonomy does not have. See research-log/maturity-repair/notes.md.","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8095-8480-ccb23b3fc612","name":"Mimic Human Evolution in Silico","description":"Use evolutionary algorithms and generative models to simulate human evolution processes in computational systems, driving more robust and adaptable AI."},{"id":"1c2cb37e-2a00-8039-bd56-dd7149a858b2","name":"Bayesian Cognitive Architectures  and Program Synthesis","description":"Implement Bayesian cognitive models and probabilistic programming and program synthesis techniques to endow AI with more human-like reasoning, planning, and decision-making capabilities."},{"id":"1c2cb37e-2a00-8031-813c-ee8051dce625","name":"Improve Out of Context Reasoning in Large Language Models","description":"Improve the ability of large language models or systems based on them to synthesize disparate information outside their context windows to formulate new scientific hypotheses"},{"id":"1c2cb37e-2a00-8076-b14c-f53f763fe525","name":"Developmentally Realistic AI Training","description":"Train AI models on data similar to those developing humans or animals actually experience"},{"id":"1f0cb37e-2a00-80d0-8192-fa2f0e3f43ae","name":"Holistic Brain-Inspired Architectures","description":"Develop AI architectures that are inspired by the human brain's structure and functionality, enabling more flexible and general reasoning and planning."}]},{"id":"1c1cb37e-2a00-808e-bcac-eb5bcdc8f23c","slug":"biological-life-is-our-only-working-example-of-complex-evolved-computation","name":"Biological Life is Our Only Working Example of Complex Evolved Computation","description":"Biological systems are the sole example we have of complex, evolved computation. Replicating this level of complexity in digital systems could unlock entirely new computational paradigms.","field":"Computation","outcome":"We get a second example of complex evolved computation to generalise from. At present biology is the only one.","outcome_rationale":"Their description states the problem as biology being the sole example; the unlock is a second one.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Complexity and open-endedness metrics exist and can be computed on any simulation, but there is no agreement that any of them distinguishes genuine evolved complexity from an artifact of the fitness landscape.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"Speculative","is_primary":1,"rationale":"Independent passes disagreed (type: Prediction and modeling vs Design search; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Design search","primary_maturity":"Speculative","primary_rationale":"Independent passes disagreed (type: Prediction and modeling vs Design search; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-802a-9420-c1f9887b4301","name":"Virtual Life Enabled by Learning and GPUs","description":"Develop virtual life simulations powered by machine learning and modern GPU infrastructure to mimic the complex, evolved computations seen in biological systems."}]},{"id":"1c1cb37e-2a00-80ff-986f-d66a6d646d7d","slug":"insufficient-integrated-earth-climate-models","name":"Insufficient Integrated Earth Climate Models","description":"Current models struggle to accurately predict climate tipping points due to the intricate interplay of diverse climatic factors, hindering proactive intervention efforts. Additionally, designing optimal climate control strategies is challenging because of the nonlinear and multifaceted interactions among economic, technological, and social factors.","field":"Geophysics and Climate","outcome":"Climate interventions get evaluated before anyone commits to them, because Earth system models can represent tipping behaviour and the human feedbacks that drive it.","outcome_rationale":"Their description names both halves: tipping-point prediction and the nonlinear economic and social coupling.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Hindcast skill against the historical record is measurable, but the quantity of interest is prediction of tipping events that have not occurred, so validation is indirect.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-804b-87d4-e3aa9f80b675","name":"Build a Predictive System for Tipping Points","description":"Develop advanced predictive systems that integrate diverse climate data to forecast tipping points more accurately."},{"id":"1c2cb37e-2a00-8051-a4a1-d8976472caf4","name":"Next Generation Earth System Model","description":"Build open, composable Earth System Simulation infrastructure based on high-resolution data. De-silo climate data across ESMs, IAMs, observational data to increase collaborative potential.\n\nTraditional ESMs are built using legacy programming languages, which introduce a barrier for new entrants to the field and impede the usage of hybrid machine-learning techniques and modern computing architectures.\n\nCollect high-resolution earth system data: Today’s global models and reanalyses are at tens of kilometers resolution. Build an open dataset of ultra-high-resolution simulations or merged observations (e.g. <1 km, resolving clouds, storms, and local topography). It could train AI to capture fine-scale processes (convection, urban heat islands, etc.) that current models miss."},{"id":"1c2cb37e-2a00-801e-92a0-cab2f0d40c0e","name":"Develop Integrated Assessment Models","description":"Create more robust integrated assessment models that minimize ungrounded economic assumptions and better capture sensitive intervention points and amplification mechanisms in socioeconomic and political systems."}]},{"id":"1c1cb37e-2a00-8058-9458-d04c0edfd1b8","slug":"insufficient-monitoring-and-modeling-of-climate-processes-and-control-paths","name":"Insufficient Monitoring and Modeling of Climate Processes and Control Paths","description":"We have limited capacity to predict key disruptive events, such as solar flares that threaten power grids and communications, alongside an incomplete understanding of natural processes (atmospheric, ocean, etc.) that underpin climate models. We need better monitoring tools for characterizing phenomena that impact climate dynamics, such as aerosol-cloud interactions, and assessing potential interventions such as marine cloud brightening. These issues underscore the need for enhanced observational tools and more sophisticated models of climate processes.","field":"Geophysics and Climate","outcome":"Aerosol-cloud interaction and ocean heat transport dominate climate uncertainty. Observe them at model resolution and they stop having to be parameterised.","outcome_rationale":"Their description ties poor prediction directly to insufficient observation of specific named processes.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Observational coverage and retrieval uncertainty for each named process are direct, reported quantities.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (type: Measurement and sensing vs Prediction and modeling; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (type: Measurement and sensing vs Prediction and modeling; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80bd-8beb-f9e1465f2225","name":"Improve Solar Flare Models","description":"Enhance predictive models for solar flares using advanced data analytics and observation techniques to better forecast solar activity.\n\nNote: NASA spends $0.8 bn/yr on heliophysics. "},{"id":"1c2cb37e-2a00-8089-a597-c502a63392fd","name":"Integrated Methane and N2O Monitoring in Natural Systems","description":"Deploy networks of in-situ and remote sensors to monitor emissions across key ecosystems like thawing permafrost, peatlands, and tropical forests. "},{"id":"1c2cb37e-2a00-8068-8f89-edeb057c4367","name":"Equip Ships and Planes as Sensors or Modulators of Climate","description":"Retrofit existing vessels and aircraft with advanced sensors to systematically measure aerosol–cloud interactions in situ, improving our understanding and models."},{"id":"1c2cb37e-2a00-8050-8de1-cbcfe62c3020","name":"Improved Remote Sensing of Aerosol Particle Size and Composition","description":"Develop improved sensor technologies for airborne particulates, e.g., hyperspectral and lidar-based remote sensing for aerosol particle size, type, and radiative forcing potential"},{"id":"1c2cb37e-2a00-8067-8b23-d9f381214260","name":"Stratospheric Observation Platforms","description":"Deploy specialized platforms in the stratosphere to gather high-resolution data on atmospheric processes and composition."},{"id":"1c2cb37e-2a00-80d5-b69b-cf1991caba0a","name":"Starlink Atmospheric Tomography","description":"Explore the concept of using satellite constellations (like Starlink) to perform atmospheric tomography, thereby building a 3D picture of atmospheric dynamics."},{"id":"1c2cb37e-2a00-8082-bf36-f98b9dd238e4","name":"Dynamic Ecosystem Feedback Modeling","description":"Develop dynamic models that incorporate microbial, hydrological, and climate-driven processes to better capture methane/N₂O feedback loops."},{"id":"1c2cb37e-2a00-80d3-b023-fccdefd31b4c","name":"Ocean Heat and Circulation Modeling","description":"Global ARGO-like sensors for deep ocean currents. We have relatively sparse ocean data compared to atmospheric data. Initiatives like Argo floats ( ~4,000 drifting sensors) have collected over two million ocean profiles of temperature and salinity, providing a crucial 3D view of the oceans.\n\nExpanding such efforts (more floats, deeper measurements, biogeochemical sensors) and releasing the data in unified formats could enable AI to model ocean currents, carbon uptake, and climate patterns like El Niño with greater skill. A gap remains in high-resolution, full-depth ocean data that AI models could exploit for improved climate forecasts."},{"id":"1c3cb37e-2a00-8086-b204-e2ae9e3f46b5","name":"Efficient Ocean Exploration Systems","description":"Innovate and deploy more efficient and robust systems for deep ocean exploration, enabling comprehensive study of deep-sea environments and their unique biology."}]},{"id":"1c1cb37e-2a00-80fa-9903-df046affff97","slug":"inadequate-emergency-climate-interventions-and-response","name":"Inadequate Emergency Climate Interventions and Response","description":"There is a critical need for more precise, rapid, and localized climate intervention strategies. Current approaches lack the fine-grained models and rapid response mechanisms required to adapt to diverse climate impacts, such as heatwaves, which demand swift and effective action. The ability to control local weather phenomena—including cloud formation and hurricanes—could help mitigate climate risks.","field":"Geophysics and Climate","outcome":"A local climate hazard gets forecast while there is still time to act on it, not attributed afterwards.","outcome_rationale":"Their description emphasises precision, speed and locality, all of which matter only if response is possible inside the event window.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Forecast skill and response latency are measurable, but the gap is about intervention capability whose efficacy at scale has no observation set.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"Speculative","is_primary":0,"rationale":"Hurricane diversion and cloud seeding at effective scale require deployed physical infrastructure that does not exist.","confidence":"confident"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"Maturity decided by David, 2026-08-23, overriding a label both independent passes agreed on. The ML is real and is already covered by the AI for Good gap map; what is missing is the alerting and action layer that turns a forecast into a faster response. Prediction without an action path does not move an emergency-response gap.","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"Maturity decided by David, 2026-08-23, overriding a label both independent passes agreed on. The ML is real and is already covered by the AI for Good gap map; what is missing is the alerting and action layer that turns a forecast into a faster response. Prediction without an action path does not move an emergency-response gap.","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80fd-a2db-ff18db8be27a","name":"Fine-Grained Impact Modeling and Response","description":"Develop detailed models that capture the local impacts of climate change (e.g., heatwaves) and implement responsive strategies to minimize harm and adapt to changing conditions."},{"id":"1c2cb37e-2a00-80bd-8beb-f9e1465f2225","name":"Improve Solar Flare Models","description":"Enhance predictive models for solar flares using advanced data analytics and observation techniques to better forecast solar activity.\n\nNote: NASA spends $0.8 bn/yr on heliophysics. "},{"id":"1c2cb37e-2a00-8068-8f89-edeb057c4367","name":"Equip Ships and Planes as Sensors or Modulators of Climate","description":"Retrofit existing vessels and aircraft with advanced sensors to systematically measure aerosol–cloud interactions in situ, improving our understanding and models."},{"id":"1c2cb37e-2a00-8050-8de1-cbcfe62c3020","name":"Improved Remote Sensing of Aerosol Particle Size and Composition","description":"Develop improved sensor technologies for airborne particulates, e.g., hyperspectral and lidar-based remote sensing for aerosol particle size, type, and radiative forcing potential"},{"id":"1c2cb37e-2a00-8067-8b23-d9f381214260","name":"Stratospheric Observation Platforms","description":"Deploy specialized platforms in the stratosphere to gather high-resolution data on atmospheric processes and composition."},{"id":"1c2cb37e-2a00-80d5-b69b-cf1991caba0a","name":"Starlink Atmospheric Tomography","description":"Explore the concept of using satellite constellations (like Starlink) to perform atmospheric tomography, thereby building a 3D picture of atmospheric dynamics."},{"id":"1c2cb37e-2a00-8082-bf36-f98b9dd238e4","name":"Dynamic Ecosystem Feedback Modeling","description":"Develop dynamic models that incorporate microbial, hydrological, and climate-driven processes to better capture methane/N₂O feedback loops."},{"id":"1c2cb37e-2a00-80d3-b023-fccdefd31b4c","name":"Ocean Heat and Circulation Modeling","description":"Global ARGO-like sensors for deep ocean currents. We have relatively sparse ocean data compared to atmospheric data. Initiatives like Argo floats ( ~4,000 drifting sensors) have collected over two million ocean profiles of temperature and salinity, providing a crucial 3D view of the oceans.\n\nExpanding such efforts (more floats, deeper measurements, biogeochemical sensors) and releasing the data in unified formats could enable AI to model ocean currents, carbon uptake, and climate patterns like El Niño with greater skill. A gap remains in high-resolution, full-depth ocean data that AI models could exploit for improved climate forecasts."},{"id":"1c2cb37e-2a00-8068-9fed-df83fafa97c1","name":"Backup Power Transformer Protection","description":"Establish protocols and infrastructure for banking backups of power transformers to mitigate the impact of solar flare-induced disruptions."},{"id":"1c2cb37e-2a00-8051-af8d-cd113b3f827c","name":"Cloud Seeding","description":"Assess the feasibility of using  cloud seeding techniques to stimulate precipitation and modulate local weather conditions in a controlled manner."},{"id":"1c2cb37e-2a00-8090-8bab-c3e792a185f2","name":"Hurricane Diversion or Mitigation","description":"Explore strategies to divert or mitigate the impact of hurricanes using advanced atmospheric control methods."},{"id":"1cfcb37e-2a00-80c8-9a7d-deb9ed5cef96","name":"Subduction Zone Observation","description":null},{"id":"1cfcb37e-2a00-80a2-a53c-f8602d0a4f8c","name":"Earthquake Prediction","description":null}]},{"id":"1c1cb37e-2a00-80ce-ba28-fe5226876f66","slug":"intervening-in-earth-systems-at-scale-is-largely-untested","name":"Intervening in Earth Systems at Scale is Largely Untested","description":"There are currently no direct interventions to address climate tipping points such as glacier melt, leaving some critical processes unmitigated. The fundamental science and engineering principles behind emergency climate interventions remain largely untested at relevant scales, limiting our preparedness for rapid climate change.\n\nSee: https://www.outlierprojects.org/","field":"Geophysics and Climate","outcome":"Emergency climate intervention is either a real option or ruled out, decided at meaningful scale and in advance rather than under duress.","outcome_rationale":"Their description frames the loss as preparedness: untested principles mean no usable option in a rapid-change scenario.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Small-scale field trials are technically possible and some have been proposed, but there is no agreement that results at trial scale would establish planetary-scale efficacy, and active dispute over whether testing is permissible at all. The blocker is agreement on what would count as a verdict.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Physical build vs Prediction and modeling; maturity: Speculative vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Physical build vs Prediction and modeling; maturity: Speculative vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8007-b08b-d682fc365833","name":"Glacial Climate Intervention","description":"Develop and test intervention strategies aimed at stabilizing glaciers and mitigating associated climate impacts."},{"id":"1c2cb37e-2a00-80fb-a999-c3565dccfc13","name":"Safely Testing Emergency Atmospheric Interventions","description":"Conduct controlled, scaled experiments to validate the basic science and engineering principles behind emergency climate interventions such as sunlight reflection modification and atmospheric methane removal."},{"id":"1c2cb37e-2a00-80b9-82cc-c2e7affad1b0","name":"Solar Radiation Modification Fundamentals","description":"Measurement reporting and verification, foundation modeling and interventions for solar radiation modification."},{"id":"1c2cb37e-2a00-80e7-be9f-fd49fdd8cf7a","name":"Planetary Sunshade Fundamentals","description":"Understand whether a planetary sunshade could be viable"},{"id":"1c2cb37e-2a00-8080-a736-dbb60e9d31f5","name":"Removal of Methane and N2O from the Atmosphere","description":"Develop systems to oxidise diffuse atmospheric methane and nitrous oxide."},{"id":"1c8cb37e-2a00-807a-8c8a-f352bb0eaba5","name":"Arctic Interventions","description":"Understand whether specific interventions in the Arctic, from Mixed-Phase Cloud Thinning to Sea Ice Thickening can counter extreme warming and tipping-points in the region. "},{"id":"1c8cb37e-2a00-80cc-8cbd-d7f587ecea40","name":"Glacier Measurement Tools","description":"New technologies to collect data about ice sheets and improve the efficiency with which we can use that data. For example, UAV-borne ice-penetrating radar systems."},{"id":"1c8cb37e-2a00-80cb-9d04-eb04f661a70d","name":"Protecting Ocean Ecosystems","description":"Develop approaches to protect ocean ecosystems. E.g. Studies have shown that shading corals for the 4 hottest hours of the day can significantly reduce bleaching."}]},{"id":"1c1cb37e-2a00-8037-9a36-c2081290cac8","slug":"inadequate-interventions-for-greenhouse-gas-removal","name":"Inadequate Interventions for Greenhouse Gas Removal","description":"We need more effective approaches to removing greenhouse gases from the atmosphere to mitigate climate impacts. However, challenges remain in harnessing natural carbon removal systems—due to difficulties in accurately measuring their environmental impact—and in reducing methane emissions from sources like the cow rumen. Innovative strategies, including modifying cow microbiomes and deploying scalable measurement and validation platforms, are essential to advance greenhouse gas removal efforts.\n\nSee also:  https://www.bezosearthfund.org/news-and-insights/bezos-earth-fund-releases-global-roadmap-to-scale-greenhouse-gas-removal-technologies and https://gaps.frontierclimate.com/ ","field":"Geophysics and Climate","outcome":"Removal gets paid for by verified result. Right now the inability to prove that carbon stayed removed is what limits capacity.","outcome_rationale":"Their description names measurement difficulty as the specific obstacle to harnessing natural removal systems.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Tonnes of CO2-equivalent removed and cost per tonne are the units the entire field already transacts in.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"2-5 years","is_primary":0,"rationale":"Modifying the cow microbiome to suppress methanogenesis is a biological design objective.","confidence":"confident"},{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8092-8b67-de1d9531246d","name":"Develop Scaled Platforms for Validation & Measurement","description":"Create computational and experimental platforms tailored to validate, measure, report, and verify carbon removal and its environmental impacts in natural systems such as the ocean or soil.\n\nBeyond MRV, methods to valorize CDR at scale are needed."},{"id":"1c2cb37e-2a00-8089-a597-c502a63392fd","name":"Integrated Methane and N2O Monitoring in Natural Systems","description":"Deploy networks of in-situ and remote sensors to monitor emissions across key ecosystems like thawing permafrost, peatlands, and tropical forests. "},{"id":"1c2cb37e-2a00-806e-85d2-d0591dae93a7","name":"Modify the Cow Microbiome","description":"Reduce methane production from cows by modifying or removing the methanogen microbes in their rumen: \n• Microbe-targeting vaccines\n• Gene engineered cow microbiome\n• Highly specific antibacterials"},{"id":"1c2cb37e-2a00-80ad-bddd-de59a8608c69","name":"Precise Methane Production Measurement","description":"Implement high-precision monitoring systems to measure methane production from livestock, enabling optimized agricultural practices."},{"id":"1c2cb37e-2a00-8080-a736-dbb60e9d31f5","name":"Removal of Methane and N2O from the Atmosphere","description":"Develop systems to oxidise diffuse atmospheric methane and nitrous oxide."}]},{"id":"1c1cb37e-2a00-80fe-9540-c14ef07e4d85","slug":"outdated-and-fragmented-recycling-cleanup-and-bioremediation-systems","name":"Outdated and Fragmented Recycling, Cleanup and Bioremediation Systems","description":"We need to improve our management of natural systems. The world's oceans suffer from extensive pollution, undermining marine ecosystems and disrupting global climate processes. Inefficient wildfire management is a prime example of problematic management of natural fire cycles.","field":"Geophysics and Climate","outcome":"Materials and pollutants come back out at something like the rate they go in. The loops that currently run one way start closing.","outcome_rationale":"Their description names recycling, ocean cleanup and fire-cycle management as instances of the same management failure.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Recycling rates by material, tonnes of pollutant removed and acreage under managed burn are all directly reported.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (type: Coordination and institutions vs Physical build; maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Physical build","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (type: Coordination and institutions vs Physical build; maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8011-855f-cda13bd4d846","name":"Efficient Ocean Cleaning Technologies","description":"Innovate and deploy efficient, scalable methods for cleaning up ocean pollution to restore marine ecosystems and improve ocean health."},{"id":"1c2cb37e-2a00-80f4-a338-db68225eba34","name":"Advanced Recycling Capabilities","description":"Develop and scale advanced recycling methods, especially for plastics."}]},{"id":"1c1cb37e-2a00-8008-9ca1-cf81d4b44116","slug":"frontier-telescopes-are-expensive-and-take-decades-to-build","name":"Frontier Telescopes Are Expensive and Take Decades to Build","description":"The current model for building space telescopes is cost-prohibitive and slow, often requiring decades of development. New approaches that exploit reduced launch costs and modular assembly are needed to accelerate telescope construction and reduce costs, and there needs to be the organizational structure and hunger to adopt such methods.","field":"Astrophysics","outcome":"Direct spectroscopy of an Earth-like exoplanet atmosphere arrives within one generation of astronomers instead of two, because large apertures fit inside a single funding cycle.","outcome_rationale":"The unlock is aperture per dollar per decade; everything downstream in observational astrophysics scales with it.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Cost per telescope and elapsed years from concept study to first light are published for every major observatory.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"Working now","is_primary":0,"rationale":"Optical and structural design optimisation already works, which is precisely why it is not the binding link.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"Speculative","is_primary":0,"rationale":"Their own description closes on needing 'the organizational structure and hunger to adopt such methods'.","confidence":"confident"},{"ai_type":"Physical build","maturity":"Speculative","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Physical build","primary_maturity":"Speculative","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[{"id":1,"quantity":"Elapsed time from first concept study to launch, flagship space observatory","current_value":"32","unit":"years","as_of":"2021-12-25","target_value":null,"target_basis":null,"source_title":"Mission Timeline — James Webb Space Telescope","source_url":"https://www.stsci.edu/jwst/about-jwst/history/mission-timeline","source_doi":null,"source_checked":"unchecked","reads_as":"From the 1989 workshop on what became JWST to its launch took 32 years.","direction":"lower is better","context":"Rubin ran about 33 years from first planning to first light, the ELT about 31. Three flagship observatories, three decades each.","caveat":"No agency publishes a target for this, so there is nothing official to measure the 32 against.","is_null_result":0,"rationale":"STScI's own timeline states plainly: 'From conception to launch of the James Webb Space Telescope took 32 years.' The clock starts at the September 1989 NGST workshop at STScI and stops at the 25 December 2021 launch. This is the quantity Convergent's gap statement is about — their sentence is that the current model 'often requir[es] decades of development' — so the indicator is the decade count itself rather than a proxy for it. The comparison case runs the same way: Rubin Observatory's own history page dates the first planning to the early 1990s and first images to 23 June 2025, and dates the National Academies recommendation to 2001, giving 24 years from formal recommendation to first light. No target is recorded because no agency has published one. Astro2020's remedy for exactly this problem is a Great Observatories technology-maturation programme, which is a mechanism proposal, not a duration commitment, and inventing a round number here would be worse than leaving it blank.","confidence":"confident"}],"capabilities":[{"id":"1c2cb37e-2a00-807b-bd1e-d673f8b6586b","name":"Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes","description":"Leverage the reduced launch costs enabled by vehicles like Starship and employ modular assembly techniques to build telescopes more rapidly and economically."},{"id":"1c2cb37e-2a00-80f2-902f-f0c88b808f31","name":"Space Telescope Factory","description":"Standardize a broadly capable design, and create a space telescope factory to produce modular components at scale, dramatically reducing the cost and time required for mission development."},{"id":"1c2cb37e-2a00-8029-81c3-e9ecd615a1fb","name":"Leveraging Commercial Component Advances","description":"Utilize advances in commercial components and software to accelerate the development cycle of telescopes."}]},{"id":"1c1cb37e-2a00-8072-976d-fb7598c7401e","slug":"major-planetary-science-and-astrobiology-missions-are-not-realized-by-existing-government-space-agencies","name":"Major Planetary Science and Astrobiology Missions Are Not Realized by Existing Government Space Agencies","description":"Some of the most important planetary science and astrobiology missions remain unrealized by traditional government agencies like NASA. Alternative, independent initiatives are needed to explore these high-priority scientific questions.","field":"Astrophysics","outcome":"Europa and Enceladus plume sampling fly on a schedule set by mission design rather than by position in an agency queue.","outcome_rationale":"Their framing is explicitly about missions remaining unrealized by existing agencies, so the unlock is realisation, not capability.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Decadal-survey priorities left unflown after N years is countable, but the science forgone by not flying them is not observable.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"2-5 years","is_primary":0,"rationale":"Commercial launch and spacecraft manufacture have moved fast enough that hardware is no longer the constraint they describe.","confidence":"guess"},{"ai_type":"Coordination and institutions","maturity":"Speculative","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"Speculative","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80da-903f-fb877a408b45","name":"Independent Planetary Missions","description":"Develop and launch independent planetary science missions to explore key astrobiological questions. Relying solely on traditional government agencies like NASA unnecessarily limits planetary science and astrobiology research."}]},{"id":"1c1cb37e-2a00-80ed-8762-c51aee94a680","slug":"limited-detection-of-gravitational-waves-across-the-frequency-spectrum","name":"Limited Detection of Gravitational Waves Across the Frequency Spectrum","description":"Detecting gravitational waves allows us to observe cosmic events like black hole mergers and neutron star collisions that are invisible through traditional telescopes. Current gravitational wave detectors are primarily sensitive to audio-band signals. Some phenomena, including speculative ones such as high-frequency emissions from advanced propulsion systems, might only be detectable with novel approaches.","field":"Astrophysics","outcome":"Gravitational-wave astronomy hears below the audio band. Intermediate-mass black hole mergers and long early inspirals become observable for the first time.","outcome_rationale":"Each new frequency band in GW detection has opened a distinct source population; the unlock is the population, not the instrument.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Strain sensitivity as a function of frequency is the instrument's defining curve, and detection counts per band follow from it.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Physical build vs Measurement and sensing; maturity: Speculative vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Physical build vs Measurement and sensing; maturity: Speculative vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-801a-9fce-de57e79bf6b3","name":"Explore High-Frequency Gravitational Wave Detection","description":"Investigate new methodologies to detect high-frequency gravitational waves, potentially unveiling phenomena that are invisible to current detectors."},{"id":"1c2cb37e-2a00-802d-8b6f-ebd6dd675f08","name":"Space-Based Gravitational Wave Detection","description":"Space-based Laser Interferometer Gravitational-Wave Observatories (LIGOs) allows an extremely large detector to study regions of the gravitational wave spectrum that are inaccessible from Earth."},{"id":"1c2cb37e-2a00-80f6-8ca0-fac461706d54","name":"Decihertz (~0.1 Hz) Gravitational Wave Detection","description":"There are an exceptionally large number of compelling signals that live in this band, including: the elusive intermediate mass black holes, white dwarf mergers (putative SN Ia progenitor), tidal disruption events, high precession and eccentricity systems, neutron star merger early warning, and more."}]},{"id":"1c1cb37e-2a00-8042-a5bd-ff3dc13b3998","slug":"higher-resolution-views-of-the-universe-are-roadblocked-by-formation-flying-technology","name":"Higher-Resolution Views of the Universe Are Roadblocked by Formation Flying Technology","description":"Space telescopes offer vastly superior sensitivity to ground-based systems, but enhancing their resolution requires spacecraft with sub-micron precision. \r\n\nAngular resolution of telescopes is limited by the size of the primary optic. Coherent aperture synthesis (interferometry) gets around this by coherently combining signal from separated telescopes where the resolution is proportional to the baseline separation. This has been very successful in the radio (see Event Horizon Telescope) but in the optical regime requires extremely difficult optomechanics and controls, and the sensitivity on the ground is inherently limited by the coherence time of the atmosphere. \r\n","field":"Astrophysics","outcome":"Resolution comes from how far apart the spacecraft fly rather than how big the mirror is. Surface features on nearby exoplanets get resolved.","outcome_rationale":"Their description states the physics directly: resolution is proportional to baseline once coherent combination works.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Achieved baseline, station-keeping precision in microns and delivered angular resolution are all direct instrument metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":0,"rationale":"Metrology and coherent signal combination are tractable learned-estimation problems; radio interferometry already does the analogous reconstruction.","confidence":"confident"},{"ai_type":"Real-time control","maturity":"Speculative","is_primary":1,"rationale":"Independent passes disagreed (type: Physical build vs Real-time control; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Real-time control","primary_maturity":"Speculative","primary_rationale":"Independent passes disagreed (type: Physical build vs Real-time control; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-808a-8117-d8d4cf2fdc2d","name":"Advance Formation Flying Technology for Spacecraft","description":"Even small optical interferometers in space could vastly exceed the sensitivity of ground based systems, and directly image Earth-like planets around Sun-like stars, but this requires advances in precision (<1 micron) formation flying technology for spacecraft. [This is both an engineering bottleneck, and a scientific bottleneck] "}]},{"id":"1c1cb37e-2a00-803c-84ea-fa89d39a8929","slug":"particle-accelerators-are-large-and-expensive","name":"Particle Accelerators Are Large and Expensive","description":"Traditional particle accelerators are enormous and costly, limiting experimental flexibility. Compact, benchtop accelerators could democratize high-energy physics and open new avenues in applications such as medical isotope production.","field":"Physics","outcome":"High-energy beams on a benchtop. Accelerator-dependent research and medical isotope production stop queueing for facility time.","outcome_rationale":"Their description names democratisation of high-energy physics and isotope production as the two unlocks.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Beam energy per metre, achieved current and cost per GeV are direct and comparable across designs.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"Working now","is_primary":0,"rationale":"Learned control and tuning of accelerator parameters already works and is not the constraint.","confidence":"confident"},{"ai_type":"Real-time control","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Learned tuning is in operations at existing accelerators and makes them run better, not smaller. The gap is compactness and cost, which means laser-wakefield machines, where closed-loop optimisation is producing research results and no usable instrument. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Real-time control","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Learned tuning is in operations at existing accelerators and makes them run better, not smaller. The gap is compactness and cost, which means laser-wakefield machines, where closed-loop optimisation is producing research results and no usable instrument. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80e6-af8b-c49074411b80","name":"Benchtop Particle Accelerators","description":"Develop compact particle accelerators that can be operated on a benchtop scale, reducing both cost and size while retaining necessary performance for scientific and medical applications."}]},{"id":"1c1cb37e-2a00-80ad-a913-d3a908406535","slug":"inadequate-imaging-of-material-structures","name":"Inadequate Imaging of Material Structures","description":"Many materials’ internal structures are difficult to image with current technologies, limiting our understanding of their properties at the nanoscale.","field":"Physics","outcome":"Look inside a material at the nanoscale and trace its bulk behaviour to the internal features that cause it.","outcome_rationale":"Their description names understanding properties at the nanoscale as the loss from inadequate imaging.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Achieved spatial resolution and contrast for a named material class are direct instrument specifications.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8074-901a-db2673b0bad3","name":"Neutron Microscopy","description":"Develop and deploy neutron microscopy techniques to image material structures at high resolution, benefiting from neutrons’ deep penetration and sensitivity to light elements."}]},{"id":"1c1cb37e-2a00-8066-824e-d4ec0d434c3d","slug":"robust-and-compact-plasma-confinement-for-fusion-is-still-not-solved","name":"Robust and Compact Plasma Confinement for Fusion is Still Not Solved","description":"Stable plasma confinement is a major obstacle in achieving practical fusion energy. Advanced control systems and novel confinement techniques are needed.","field":"Physics","outcome":"Plasma stays confined for as long as the reactor needs, not as long as the instability allows. That is the last physics obstacle before sustained operation.","outcome_rationale":"Their description names stable confinement as the major obstacle to practical fusion energy.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Confinement time, energy gain factor and sustained stable duration are direct and universally reported fusion metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"2-5 years","is_primary":0,"rationale":"Magneto-inertial approaches require building confinement hardware that does not yet exist at the required performance.","confidence":"confident"},{"ai_type":"Real-time control","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (type: Prediction and modeling vs Real-time control; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Real-time control","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (type: Prediction and modeling vs Real-time control; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8066-89a7-c8aba5d79014","name":"AI-Driven Control Systems for Plasma Confinement","description":"Implement AI-based control systems to dynamically stabilize and confine plasma during fusion reactions, improving overall efficiency and stability."},{"id":"1c2cb37e-2a00-80b9-929b-d8e2eec3099a","name":"Magneto-Inertial Confinement","description":"Develop magneto-inertial confinement strategies that combine magnetic fields with inertial forces to better confine plasma for fusion."}]},{"id":"1c1cb37e-2a00-800c-aafb-ec36c16025a8","slug":"uncertainty-and-noise-in-the-science-of-room-temperature-superconductivity","name":"Uncertainty and Noise in the Science of Room-Temperature Superconductivity","description":"There remains significant uncertainty over whether metallic hydrogen can exhibit room-temperature superconductivity at reasonable pressures, and measurements of other systems have been irreproducible and fragmented.","field":"Physics","outcome":"A superconductivity claim can be settled in weeks rather than argued for years and retracted.","outcome_rationale":"Their capability is stated as conducting rigorous, non-fraudulent experiments, which makes adjudication the unlock rather than discovery.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Measurements have repeatedly been taken and repeatedly retracted; the field disputes which signature, resistance drop or Meissner expulsion under pressure, actually establishes superconductivity. The instrument is not the blocker, agreement about the verdict is.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":0,"rationale":"Automated, instrumented synthesis and measurement would remove the human discretion where irreproducibility and fraud enter.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). A replication consortium and shared measurement standards could be convened today, but getting the high-pressure community to adopt them — and resolving the underlying physics of metallic hydrogen — is what the gap asks for and neither follows from convening. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). A replication consortium and shared measurement standards could be convened today, but getting the high-pressure community to adopt them — and resolving the underlying physics of metallic hydrogen — is what the gap asks for and neither follows from convening. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8086-adf9-ee95725c38f8","name":"Conduct Rigorous, Non-Fraudulent Experiments","description":"Encourage and fund rigorous experimental investigations into metallic hydrogen superconductivity, ensuring data integrity and reproducibility."}]},{"id":"1c1cb37e-2a00-8027-8eda-e6f4a472c3ad","slug":"quantum-gravity-is-experimentally-hard-to-constrain","name":"Quantum Gravity is Experimentally Hard to Constrain\n","description":"Quantum gravity remains elusive, with experimental constraints hindered by the need for extremely large-scale or prohibitively expensive experiments.","field":"Physics","outcome":"Quantum gravity gets an experimental constraint that the field agrees is one, and becomes a partly empirical subject.","outcome_rationale":"The unlock is not any single measurement but the existence of a result the community treats as evidence.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Candidate discriminators exist and are actively pursued, including gravitationally induced entanglement between mesoscopic masses, tabletop interferometry and Lorentz-invariance violation searches. What is contested is whether observing them would settle anything: there is live literature arguing gravity-mediated entanglement can be reproduced by classical gravitational fields, so reading it as proof requires further assumptions about local mediators. The blockage is the absence of agreement about what counts as a verdict, not the absence of an instrument.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Speculative","is_primary":1,"rationale":"Independent passes disagreed (type: Physical build vs Measurement and sensing; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Speculative","primary_rationale":"Independent passes disagreed (type: Physical build vs Measurement and sensing; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[{"id":7,"quantity":"Highest benchmarked quantum-gravity figure-of-merit eta across mechanical quantum-control platforms","current_value":"1.1e-7","unit":"dimensionless (eta)","as_of":"2025-12-20","target_value":"1","target_basis":"A physical threshold, not an aspiration. Eq. 1 of the source defines eta^2 = xi^2|grad F|/(hbar*gamma) >= 1 as the condition for gravitationally induced entanglement to outrun decoherence, where xi is the coherence length, |grad F| the force gradient between the masses and gamma the oscillators' total mechanical energy dissipation rate. eta >= 1 is what has to be true for the experiment to detect anything at all.","source_title":"Agafonova, Rossello, Mekonnen & Hosten, 'One-milligram torsional pendulum toward experiments at the quantum-gravity interface', arXiv:2408.09445v4 (20 December 2025), published as Communications Physics 9, 80 (30 January 2026). Table S1 of the supplement benchmarks eta across thirteen platforms.","source_url":"https://arxiv.org/abs/2408.09445","source_doi":"10.1038/s42005-026-02514-w","source_checked":"verified","reads_as":"The most sensitive apparatus anywhere is about seven orders of magnitude short of the point at which gravity could be shown to entangle two masses.","direction":"higher is better","context":"Table S1 benchmarks thirteen platforms spanning more than twenty orders of magnitude of mass. The highest eta is 1.1e-7, a microgram-scale superconducting microsphere; LIGO's 10 kg pendulums reach 2.8e-8; the source's own 1 mg torsional pendulum reaches 4.3e-9 (quoted as 4e-9 in its text), roughly on par with ton-scale bar resonators. The source projects three improvement levels for its own platform, the furthest reaching 1.3e-4 — still four orders short.","caveat":"eta is necessary but not sufficient. Entanglement also requires the purity P to approach 1, and the platform holding the highest eta has P = 2e-9, so a single number understates the distance. This is one group's compilation, not a community-maintained series, and the comparison depends on assumed separations and geometries set out in their Supplementary Section 2.3.","is_null_result":0,"rationale":"REVISED after Gate A, which found the previous version of this row to be a FALSE NULL — the row asserted that no usable indicator existed, and one does. That is worse than a missing number, because the artifact made a claim about it.\n\nWhat survives from the original reasoning, because the gate endorsed it: the lower bound on the quantum-gravity energy scale from photon time-of-flight dispersion (Fermi-LAT, E_QG,1 > 7.6 E_Planck from GRB 090510; extended by LHAASO on GRB 221009A) is still rejected as the indicator for this gap, on the original two grounds. It constrains one class of model — linear-in-energy Lorentz violation — that most leading quantum-gravity programmes do not predict, and the community has no shared claim about what any particular bound would settle. That judgment was right. The error was generalising it into a claim that the field has no such quantity at all.\n\nIt has one. eta is defined in the source as the ratio that decides whether gravitationally induced entanglement survives decoherence, it has a threshold everyone in that literature agrees on (eta = 1), it has a direction (higher), and Table S1 benchmarks it across thirteen platforms — which is what makes it an indicator rather than one laboratory's number.\n\nThree of the gate's own specifics were wrong and were corrected here against the source PDF, read directly: the value is 4e-9 for the torsional pendulum and not 2.8e-9 (that figure is from an earlier arXiv version; 2.8e5 is the phonon occupation); the denominator of eta^2 is hbar*gamma, the dissipation rate, not hbar*omega_0; and the pendulum is not 'second only to LIGO' — it ranks fourth in Table S1, behind a superconducting microsphere (1.1e-7), LIGO (2.8e-8) and a bar resonator (5.1e-9), which the paper's own text states as 'roughly on par with ton-scale bars and 10-kg-scale pendulums'.\n\nThe value recorded is the field maximum (1.1e-7), not the source paper's own apparatus, because the gap is about whether the field can constrain quantum gravity, not about one group. That choice is a judgment and is stated here so it can be reversed. The publication was confirmed through Crossref: Communications Physics 9, 80, issued 2026-01-30, DOI 10.1038/s42005-026-02514-w.","confidence":"confident"}],"capabilities":[{"id":"1c2cb37e-2a00-80e3-b8ec-e600634fff91","name":"Targeted Experiments for Quantum Gravity","description":"Design and execute key experiments that probe quantum gravity phenomena without requiring massive accelerators, leveraging innovative, cost-effective approaches."}]},{"id":"1c1cb37e-2a00-801a-bbfd-cd37c0419eca","slug":"incomplete-resolution-of-the-possibility-of-low-energy-nuclear-reactions","name":"Incomplete Resolution of the Possibility of Low-Energy Nuclear Reactions","description":"While low-energy nuclear reactions (LENRs) have received substantial attention and there is no good evidence they exist, there may still be other mechanisms or parameter combinations that are underexplored.","field":"Physics","outcome":"The question closes, either way, and stops absorbing attention.","outcome_rationale":"Their description is careful that there is no good evidence but some parameter combinations remain underexplored; the unlock is resolution, not confirmation.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Calorimetry can be performed and has been, but the field's history is that results are disputed rather than accepted, so no agreement exists that a given measurement settles the claim. Arguably tier 1 with sufficiently rigorous controls, which is why this is marked a guess.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Running experiments","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8025-8dbc-d924ee7ef0ea","name":"Possibility of LENR Experiments","description":"Design and implement experiments to rigorously test low-energy nuclear reaction principles, providing clear data on reaction mechanisms and viability.\n\nThis is highly speculative and at this point unlikely to yield practical LENR, however, see: https://coldfusionblog.net/2019/03/13/the-case-against-cold-fusion-experiments/ "}]},{"id":"1c1cb37e-2a00-80c0-80e2-d417b7597a77","slug":"inability-to-model-turbulence","name":"Inability to Model Turbulence","description":"Modeling turbulence remains one of the most challenging problems in physics due to its nonlinear and chaotic nature.","field":"Physics","outcome":"Turbulence gets predicted at engineering cost. That removes the largest single source of uncertainty in climate, aerospace and fusion modelling at once.","outcome_rationale":"Turbulence is a shared upstream dependency, so the downstream unlock spans several fields at once.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Prediction error against direct numerical simulation and experiment at a stated Reynolds number and compute budget is a direct benchmark.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-8058-8e8b-eb1956e3292a","name":"Develop New Modeling Frameworks for Turbulence","description":"Create and implement novel mathematical models and computational frameworks that can more accurately simulate and predict turbulent flows."}]},{"id":"1c1cb37e-2a00-804b-bc6b-de7f4bc3adce","slug":"inability-to-anticipate-or-prevent-ecosystem-tipping-points","name":"Inability to Anticipate or Prevent Ecosystem Tipping Points","description":"We lack the models and infrastructure to monitor and predict how ecosystems behave under stress or when they might collapse, as well as what metrics to use to determine that restoration is effective. We need to understand the underlying dynamics, feedback loops, and thresholds that lead to ecosystem degradation/ collapse, as well as to secure essential systems like pollination.","field":"Ecology","outcome":"An ecosystem under stress gives a usable warning before it collapses, and management turns pre-emptive rather than reactive.","outcome_rationale":"Their capability list leads with detecting early-warning signals of collapse, which is the whole value of anticipation.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Candidate early-warning indicators exist and are computed, critical slowing down among them, but the field disputes whether they reliably precede collapse or only appear so in hindsight. Their own text concedes it is unclear what metrics would show restoration worked.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":0,"rationale":"Pollination-network and biodiversity monitoring depend on automated detection from sensor streams, which works today.","confidence":"confident"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80df-a152-c3a9cb9e0122","name":"Foundational Datasets and Models of Biodiversity and Species Interaction","description":"Develop comprehensive, high-quality datasets and predictive models to better understand and forecast animal movements and biodiversity shifts.\n\nUse advanced computational and theoretical models that capture how species interact, cascade through ecosystems, and ultimately influence stability or collapse. These models will help identify key feedback loops and thresholds, enabling targeted intervention before degradation accelerates."},{"id":"1c2cb37e-2a00-8072-8084-ef21873adfa6","name":"Detect Early-Warning Signals of Collapse","description":"Develop monitoring techniques to detect early-warning indicators—such as critical slowing down—that suggest an ecosystem is approaching a tipping point. \n\nEstablish robust, scalable networks for real-time ecological monitoring using integrated technologies. Deploy sensor arrays that combine soundscapes, environmental DNA (eDNA) sampling, and satellite data fusion to continuously assess ecosystem health across diverse regions."},{"id":"1c2cb37e-2a00-80ff-a32a-ccdab49abf1f","name":"Pollination Network Monitoring and Recovery","description":"Secure and restore pollination services critical to food systems and biodiversity. This involves:\n\n• Building ecological models that map plant–pollinator interactions and forecast vulnerability or collapse points.\n• Deploying global pollinator monitoring systems using visual, acoustic, and eDNA sensors paired with AI to track pollinator diversity and behavior.\n• Designing landscape-level interventions such as habitat corridors, floral resource planning, and pesticide regulations to boost wild pollinator recovery.\n\nPollination is critical for food systems and biodiversity, yet global pollinator populations are in sharp decline, and we lack robust ways to track, model, or supplement pollination services at scale. We need to: \n\n• Build ecological models that map plant–pollinator interactions and predict vulnerability or collapse points.\n• Deploy global pollinator monitoring infrastructure: networks of sensors (visual, acoustic, eDNA) and AI models to monitor pollinator presence, diversity, and behavior across ecosystems and crop systems\n• Design and deploy landscape-Level Interventions for Pollinator Recovery: large-scale habitat corridors, floral resource planning, and pesticide regulations to recover wild pollinators."}]},{"id":"1c1cb37e-2a00-800c-90aa-fcb685877ee5","slug":"challenges-in-tracking-and-restoring-resilient-ecosystems","name":"Challenges in Tracking and Restoring Resilient Ecosystems","description":"Regenerating degraded environments and designing self-sustaining systems require a unified understanding of ecological dynamics. Our current models fall short in predicting complex interactions—such as feedback loops and stability thresholds—that determine ecosystem behavior. \n\nTo close this gap, we need better datasets and models of biodiversity and animal movements, as well as tools to predict and contain invasive species. We also need the ability to experiment with restoration strategies, and validate approaches ranging from rewilding to engineering de-extinction technologies.","field":"Ecology","outcome":"Test a restoration strategy against a predictive model before betting a landscape on it.","outcome_rationale":"Their description asks for the ability to experiment with and validate restoration strategies, which is what a testbed plus model buys.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Species counts and movement data are measurable, but ecosystem resilience is inferred from those indicators rather than observed directly.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":0,"rationale":"Biodiversity and animal-movement datasets come from camera traps, acoustics and satellite imagery, where automated classification already works.","confidence":"confident"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80f8-b488-cbdf3bab9c56","name":"Tools to Predict and Contain Invasive Species ","description":"Track and forecast the spread of invasive species and simulate containment strategies. "},{"id":"1c2cb37e-2a00-80b3-9520-f283b2744e56","name":"Study Closed Ecosystems","description":"Conduct detailed studies of existing closed ecosystems to understand the interactions and feedback loops that enable self-sustainability."},{"id":"1c2cb37e-2a00-8067-b11f-c90242b46a0e","name":"Build Robust Ecosystem Testbeds (Closed & Open)","description":"Create well-designed testbeds that simulate key ecosystems, allowing for controlled experimentation and development of restoration strategies.\n\nDevelop experimental testbeds that simulate closed ecosystems, allowing for controlled experimentation and refinement of life-support strategies."},{"id":"1c2cb37e-2a00-80e1-b6bc-f977fe66bce1","name":"Leverage Stem Cell and Genome Editing Technologies for De-Extinction","description":"Utilize cutting-edge stem cell and mammalian genome editing techniques to facilitate de-extinction and promote genetic diversity in repopulating ecosystems.\n\nAre we preserving the genomes of critically endangered species to enable future de-extinction?"},{"id":"1c8cb37e-2a00-80cb-9d04-eb04f661a70d","name":"Protecting Ocean Ecosystems","description":"Develop approaches to protect ocean ecosystems. E.g. Studies have shown that shading corals for the 4 hottest hours of the day can significantly reduce bleaching."}]},{"id":"1c1cb37e-2a00-8044-b547-ddba3202ee63","slug":"much-of-the-biosphere-remains-uncharted-and-vulnerable-to-information-loss","name":"Much of the Biosphere Remains Uncharted and Vulnerable to Information Loss","description":"Much of Earth's biosphere—from the deep ocean to atmospheric bioaerosols—remains unexplored, with the microbial majority largely uncharted. Advancing new exploration technologies and systematically cataloging the Earth's microbiome could unlock discoveries of new life forms and biological insights that could impact health, climate, geoengineering, agriculture, and fundamental biology.","field":"Ecology","outcome":"Most of Earth's biology is microbial and uncatalogued. It gets recorded before it is lost, and turns into something searchable.","outcome_rationale":"Their description names discovery of new life forms with impact across health, climate and agriculture as the unlock.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Taxa catalogued, genomes assembled and ocean volume sampled are direct counts against a known-unknown denominator.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (type: Physical build vs Measurement and sensing; maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (type: Physical build vs Measurement and sensing; maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c2cb37e-2a00-80eb-83a5-dd256b4bb042","name":"Genomic Maps of Global Microbiomes","description":"Develop comprehensive genomic mapping initiatives for global microbiomes to catalog species essential to Earth's biosphere and inform conservation efforts."},{"id":"1c3cb37e-2a00-8034-a5b8-c53075ab7d76","name":"Map Atmospheric Extremophiles","description":"Atmospheric bio-aerosols could have applications for cloud seeding or other topics. "},{"id":"1c3cb37e-2a00-80ef-81be-ce8d7ff459db","name":"Dense Ocean DNA Sampling","description":"Implement dense, systematic DNA sampling in diverse ocean environments—especially areas with sharp gradients in temperature, depth, salinity, and pH (such as reefs, deep trenches, and hydrothermal vents)—to uncover new species and biological insights."},{"id":"1c3cb37e-2a00-8086-b204-e2ae9e3f46b5","name":"Efficient Ocean Exploration Systems","description":"Innovate and deploy more efficient and robust systems for deep ocean exploration, enabling comprehensive study of deep-sea environments and their unique biology."}]},{"id":"1c1cb37e-2a00-80ea-93f0-c8ed8e7f2155","slug":"we-can-learn-more-from-natures-biological-designs","name":"We Can Learn More from Nature’s Biological Designs","description":"Nature’s blueprints span from the unseen nanoworld to the enigmatic origins of life. Despite the incredible diversity of nanostructures, many remain hidden due to current imaging limitations. Likewise, the mysteries of animal communication—crucial for decoding behavioral cues and social structures—await breakthrough insights. Moreover, the primordial conditions of our planet, essential for understanding life’s genesis, are obscured by the absence of rocks from before 4.1 Ga and fossils before 3.5 Ga (although life almost certainly established itself on our planet before that). Addressing these challenges can unlock new understandings in biology, ecology, and more.","field":"Ecology","outcome":"Biological nanostructures, animal communication and the chemistry of early Earth become readable as engineering precedent rather than admired as curiosities.","outcome_rationale":"Their description bundles several searches under one heading; the shared unlock is converting nature's existing solutions into usable knowledge.","outcome_confidence":"guess","tier":"Proxy only","tier_rationale":"Each sub-question has its own countable output, such as Hadean zircons screened, but the gap as stated bundles four unrelated searches and has no single observable.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"2-5 years","is_primary":0,"rationale":"Decoding animal communication is a sequence-modelling problem and is the one component where language-model methods apply directly.","confidence":"confident"},{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-8006-bdb3-ec0051f89b70","name":"Advanced Imaging of Natural Nanostructures","description":"Deploy cutting-edge imaging techniques to capture high-resolution details of natural nanostructures, revealing new insights into their form and function."},{"id":"1c3cb37e-2a00-8007-9ae2-c88136304ae8","name":"Understand Ancient Proteins","description":"Systematically testing functions of ancient proteins 3800 to 50 Mya to understand ancient biochemical capabilities and environmental constraints, informing evolutionary ecology."},{"id":"1c3cb37e-2a00-8058-bbdf-f1aa464a199e","name":"Leverage Modern Machine Learning to Communicate with Animals","description":"Use advanced machine learning techniques to decode animal vocalizations and behavioral signals, enabling meaningful communication and insights into animal cognition.\n\nNote: Many animals (e.g. insects, cephalopods, amphibians) use non-vocal communication (bioluminescence, gestures, electroreception, olfaction). For olfaction see: https://www.osmo.ai/ \n "},{"id":"1c3cb37e-2a00-80a8-81d1-f9356323833d","name":"Massive-Scale Search and Screening of Hadean Zircons","description":"Almost all our knowledge of Hadean Earth comes from Hadean zircons, the “black-box recorder” of minerals. Zircons can incorporate small amounts of the ancient ocean . All known Hadean zircons come from a single site, and they are gathered artisanally by individual graduate students. The solution is to search for more Hadean zircons with the steady, systematic approach of (for example) a diamond company, and then screen massive numbers of zircons for ocean and atmosphere data."},{"id":"1c3cb37e-2a00-80a5-83fa-e62903176489","name":"Rake Ancient Lunar Regolith - “Earth’s Attic” - to Find Fragments of Hadean Rock from Earth","description":"No rock survives from the Hadean - but that is because Earth has plate tectonics. Statistically, there must be Hadean rocks waiting on ancient Lunar terrains for us to find (just as there are abundant Lunar meteorites on Earth). A plausible Earth mineral has already been identified in Apollo samples (https://www.sciencedirect.com/science/article/abs/pii/S0012821X19300202). They will be easy to spot due to their distinct mineralogy (and color). However, raking through the Lunar regolith to find them is tedious work, beyond the patience of astronauts - but not beyond the patience of robots. This is a NASA science goal, but it is undervalued, due to the intense focus on Lunar resources."},{"id":"1c3cb37e-2a00-808b-9c70-e34a843a8bcf","name":"Develop Missions to Search for Life on Europa and/or Enceladus","description":"Design and implement missions specifically aimed at exploring Europa and/or Enceladus for signs of life, leveraging advanced detection and sampling technologies."}]},{"id":"1c1cb37e-2a00-80c6-88ec-c73d91a62be4","slug":"outdated-space-station-construction","name":"Outdated Space Station Construction","description":"Only one new space station (天宫) has been launched this century, due to high costs and reliance on traditional, government-led megaprojects.","field":"Space Engineering","outcome":"Pressurised volume in orbit stops arriving one facility per generation, because habitats can be built outside government megaproject structures.","outcome_rationale":"Their description makes the point with a single statistic: one new station this century.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Cost per habitable cubic metre in orbit and stations launched per decade are direct and countable.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":0,"rationale":"Their own framing blames reliance on traditional government-led megaprojects, which is a procurement structure.","confidence":"confident"},{"ai_type":"Physical build","maturity":"Speculative","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Speculative). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Physical build","primary_maturity":"Speculative","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Speculative). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-8078-a20c-e22c7b8f585b","name":"Leverage Commercial Approaches","description":"Utilize emerging commercial spaceflight companies to drive down costs and accelerate the development of new space stations through innovative design and manufacturing."},{"id":"1c3cb37e-2a00-808b-ba63-f14fb8ff8cf1","name":"New Construction Approaches for Space Stations","description":"Develop novel construction methods—such as inflatable or modular architectures—that can be assembled in orbit to create new space stations more efficiently."},{"id":"1c3cb37e-2a00-8061-b5a1-dc7103f30fe3","name":"In-Situ Resource Utilization","description":"Everything from mining the moon and using those resources to build things in space to fuel etc etc"}]},{"id":"1c1cb37e-2a00-8085-bd56-cdf566ac7e42","slug":"lack-of-a-dedicated-field-for-planetary-terraforming","name":"Lack of a Dedicated Field for Planetary Terraforming","description":"There is currently no established field for systematically studying and applying planetary terraforming methods, leaving key challenges in transport, energy supply, and civil engineering largely unaddressed.","field":"Space Engineering","outcome":"Terraforming gets a research community and a shared problem list, so its transport, energy and civil engineering questions get worked on deliberately.","outcome_rationale":"Their description states the gap as the absence of an established field, not the absence of a technique.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Researchers, funded programmes and publications are countable as inputs, but whether the field is producing useful knowledge is not measurable from those counts.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Coordination and institutions","maturity":"Speculative","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 Speculative / v2 Working now; mechanical rule had given Working now). Convening a terraforming workshop is available today and does not move this gap. What the gap names as unaddressed — transport, energy supply and civil engineering at planetary scale — has no clear path from anything that exists. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"Speculative","primary_rationale":"Maturity re-adjudicated on the merits (v1 Speculative / v2 Working now; mechanical rule had given Working now). Convening a terraforming workshop is available today and does not move this gap. What the gap names as unaddressed — transport, energy supply and civil engineering at planetary scale — has no clear path from anything that exists. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-8024-94cd-eb132988955f","name":"Directed Work on Biological Approaches","description":"Initiate dedicated research into biological terraforming strategies that utilize living organisms to modify planetary environments."},{"id":"1c3cb37e-2a00-808c-8ef2-eee8df4ec9e5","name":"Directed Work on Non-Biological Approaches","description":"Invest in research exploring non-biological methods for planetary engineering, such as atmospheric modification and energy-efficient infrastructure."}]},{"id":"1c1cb37e-2a00-8027-9ab6-d431ec43d57f","slug":"microbes-quickly-out-evolve-our-defenses","name":"Microbes Quickly Out-Evolve Our Defenses","description":"Pathogenic microbes evolve quickly, and bad actors may exploit biotechnology for harmful purposes. Our current defenses struggle to keep pace with these evolving threats.","field":"Biosecurity","outcome":"Countermeasures outpace resistance. Right now resistance sets the shelf life of the entire antimicrobial arsenal.","outcome_rationale":"Their framing is a race; the unlock is winning it, which means development time falling below evolution time.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Time from pathogen sequence to validated candidate, and resistance prevalence over time, are both tracked surveillance quantities.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":0,"rationale":"Closing the design-test loop for candidate validation is the near-term extension of working self-driving chemistry.","confidence":"confident"},{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-803b-9766-c484a72e1d85","name":"Improved Antibiotics Discovery","description":"Develop next-generation antibiotics and novel antimicrobial compounds using advanced discovery platforms to stay ahead of evolving pathogens."},{"id":"1c3cb37e-2a00-804f-90d5-fd6cefb1bd42","name":"Improved Broad-Spectrum Antiviral Discovery","description":"Develop new broad-spectrum antivirals that can be used to treat or prevent infection from evolving viruses."},{"id":"1c3cb37e-2a00-8092-9907-ea3ab9c2cfc7","name":"Novel Vaccine Technologies","description":"Develop innovative vaccine platforms that can adapt to or be robust to rapidly mutating viruses (e.g., influenza, HIV, coronaviruses) using methods such as mosaic nanoparticles and mRNA cocktails.\n\nSafety and security considerations: https://www.sciencedirect.com/science/article/pii/S0264410X21001717 "},{"id":"1c3cb37e-2a00-80fa-b511-e48564667c48","name":"Decentralized, Low-Resource Vaccine Production","description":"Establish scalable, decentralized vaccine production systems to rapidly deploy immunizations during outbreaks, reducing reliance on centralized facilities."},{"id":"1c3cb37e-2a00-80ed-afe8-d2d72d41d0aa","name":"Prototype Medical Countermeasures","description":"Develop prototype vaccines or therapeutics for viruses in each viral family that infects humans, to be rapidly adapted for the next pandemic."},{"id":"1c3cb37e-2a00-8063-a343-f0c191624910","name":"Novel Methods for Rapid Antibody or Antibody-Like Molecule Discovery","description":"Use technologies such as rapid B-cell sorting and computational design to quickly characterize and discover antibodies and antibody-like molecules to combat emerging biosecurity threats."},{"id":"1c3cb37e-2a00-8005-8bbe-fe3e4de307ea","name":"Self-Administrable Nasal Sprays","description":"Develop nasal sprays that could be applied daily for broad-spectrum protection against respiratory pathogens"}]},{"id":"1c1cb37e-2a00-80ed-a10f-e4099f056390","slug":"we-lack-basic-capabilities-that-are-necessary-for-travel-far-beyond-earth","name":"We Lack Basic Capabilities that Are Necessary for Travel Far Beyond Earth","description":"We lack the many technologies to support exploration and survival off earth. Our exploration efforts remain confined to our solar system, limiting our potential to explore beyond and understand the broader cosmos.","field":"Space Engineering","outcome":"Missions beyond the solar system turn survivable and reachable, pushing direct observation past the boundary that currently confines it.","outcome_rationale":"Their description names confinement to the solar system as the limit on understanding the broader cosmos.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Achievable delta-v, supported mission duration and life-support closure fraction are direct engineering quantities even where the capability is far off.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"Speculative","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Physical build","primary_maturity":"Speculative","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-8063-8590-dfcf7456e758","name":"Interstellar Probes","description":"Invest in the development of interstellar probe technologies that can traverse vast distances, opening the door to exploration beyond our solar system."},{"id":"1c3cb37e-2a00-8058-9ccb-c6a8ec7fbf72","name":"Cryosleep","description":"Assuming no FTL the only way to get outside of the solar system in a single lifetime is with cryosleep."},{"id":"1c3cb37e-2a00-80b1-9ac6-f1b8237d9521","name":"Pressurizing Technology or Habitation Domes","description":"Methods to pressurize large surface areas or positive pressure habitation domes will be necessary for living most places off earth."},{"id":"1c3cb37e-2a00-80d9-9aaa-fadc0ce5db08","name":"Air Breathing Fusion for Single-Stage-to-Orbit Vehicle","description":"Air breathing fusion to propel single-stage-to-orbit (SSTO) vehicles for a single-stage, reusable spaceplane."}]},{"id":"1c1cb37e-2a00-802e-bd40-dd7361f31518","slug":"inadequate-blockers-of-transmission","name":"Inadequate Blockers of Transmission","description":"Our ability to block the transmission of pathogens is limited. Without effective strategies, airborne and surface-based transmission continues to spread diseases. A meta roadmap is here.","field":"Biosecurity","outcome":"Buildings and materials interrupt transmission on their own, so outbreak control no longer depends on everybody complying.","outcome_rationale":"Every listed capability acts on the physical environment rather than on behaviour, which is where the leverage sits.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Filtration efficiency, air changes per hour and chamber-trial transmission reduction are direct, standardised quantities.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":0,"rationale":"Surface and coating screening is exactly the bench-scale formulation loop self-driving labs already run in adjacent domains.","confidence":"confident"},{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-80e3-887b-ea30b2e7b27d","name":"Built Environment Protection","description":"Enhance indoor environmental controls to reduce pathogen transmission through advanced sensor networks, UV/opto-acoustic disinfection, and improved HVAC systems."},{"id":"1c3cb37e-2a00-80f7-97b9-d80fafddabfd","name":"Better PPE","description":"Develop next-generation, affordable, high-performance personal protective equipment to reduce transmission risks."},{"id":"1c3cb37e-2a00-80ee-a5be-fe6e7252ca2a","name":"Transmission Reduction Through Surfaces and Textiles","description":"Innovate new materials and coatings for surfaces and textiles that actively reduce pathogen viability and transmission."},{"id":"1c3cb37e-2a00-8099-9a67-f69a6a0da355","name":"Engineering the Microbiome to Improve Immune Response ","description":"Understand the role of the microbiome in immunity against infection and develop the ability to engineer or transplant microbiome communities that protect against infection."}]},{"id":"1c1cb37e-2a00-80be-82f0-c671a80de28e","slug":"insufficient-surveillance-of-bio-threats","name":"Insufficient Surveillance of Bio-Threats","description":"Rapid detection of emerging bio-threats is critical for effective intervention and containment. However, many pathogens are detected only after widespread transmission has occurred. In addition, attributing the source of these threats remains challenging.","field":"Biosecurity","outcome":"A novel pathogen gets caught while transmission is still containable, and its source can be attributed afterwards.","outcome_rationale":"Their description names both halves, detection latency and attribution, and both are needed for intervention to be possible.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Time from emergence to first detection and estimated case count at detection are both reconstructable for past outbreaks.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":0,"rationale":"Sample sharing and cross-border reporting agreements gate how much of the working detection capability actually gets used.","confidence":"confident"},{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Working now","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-80e4-a29e-f4a415ee535c","name":"Early Detection via Sequencing","description":"Deploy comprehensive sequencing-based early warning systems to rapidly detect and attribute emerging bio-threats. It is important to distinguish detection of known pathogens (easy to find in sequencing data) vs. previously unknown ones (hard)."},{"id":"1c3cb37e-2a00-805c-a65f-dcfa630b1b7e","name":"Volatilomics for Early Detection","description":"Utilize the analysis of volatile organic compounds (VOCs) as an early detection tool to identify pathogenic outbreaks."},{"id":"1c3cb37e-2a00-8014-b206-c973af2009be","name":"Enhanced Epidemiology and Epidemic Tracking","description":"Improve epidemic surveillance and modeling by integrating data from multiple sources for timely, actionable insights."},{"id":"1c3cb37e-2a00-80dd-926c-ec8f8c925cf2","name":"Fine-Grained Economic Modeling of Policy Options","description":"Develop detailed economic models to better assess the costs and benefits of different bio-threat interventions and inform policy decisions."},{"id":"1c3cb37e-2a00-8008-a821-e606c69ab239","name":"Universal Detection and Modulation of the Host Response to Infections","description":"Technologies to detect and modulate the host immune response across a broad range of infections to improve patient outcomes and treatment strategies."},{"id":"1c3cb37e-2a00-806b-97a6-ddf1ecee3ed6","name":"Emerging Bio-Threat Forensics","description":"Capabilities to determine the geographical source of bio-threats and whether they stem from natural sources or human activity. "},{"id":"1c3cb37e-2a00-802d-ad17-d21dc66f5e67","name":"Rapidly Adaptable Rapid Test Platforms ","description":"Develop rapid test platforms that can be reconfigured within days or weeks for emerging pathogens, enabling quick self-testing during a pandemic. Address the limitations of traditional lateral flow tests—which rely on antibodies and take months to develop—by significantly accelerating the deployment of diagnostic tools."}]},{"id":"1c1cb37e-2a00-8084-a877-c83d9b9c3ece","slug":"risks-of-malicious-bioengineering","name":"Risks of Malicious Bioengineering","description":"Advances in synthetic biology have unlocked unprecedented innovations, but also raise concerns about the potential for harmful bioengineering. Preventing misuse requires robust screening and control measures around DNA synthesis. Implementation must be coordinated and universal to effectively minimize the risk of malicious actors.","field":"Biosecurity","outcome":"Synthesis screening becomes universal and hard to evade, closing the cheapest route from a dangerous sequence to physical material.","outcome_rationale":"Their description states the requirement itself: implementation must be coordinated and universal to reduce risk at all.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Share of global synthesis capacity under screening is measurable; misuse prevented is not, since the counterfactual attacks are unobserved.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"Working now","is_primary":0,"rationale":"Sequence-level hazard screening and watermarking are working techniques; coverage, not capability, is the shortfall.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Screening technology exists and works — SecureDNA, IBBIS, the IGSC protocol. The gap is universal adoption, which the description itself names, and no synthesis provider outside the voluntary consortium screens today. Type unchanged from the mechanical adjudication.","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Screening technology exists and works — SecureDNA, IBBIS, the IGSC protocol. The gap is universal adoption, which the description itself names, and no synthesis provider outside the voluntary consortium screens today. Type unchanged from the mechanical adjudication.","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-80cd-8236-fc854c493289","name":"Trustless DNA Synthesis Screening","description":"Implement advanced, potentially cryptographic (where appropriate), and/or adversarial AI-based DNA synthesis screening methods to prevent misuse and ensure biosecurity."},{"id":"1c3cb37e-2a00-8031-93a9-f1fa9f84e0a0","name":"Hardware Lock DNA Synthesizer","description":"Integrate hardware-based locks into DNA synthesizers to ensure secure operation and prevent unauthorized use. Enhance systems for detecting and reporting flagged orders—enabling cross-verification with other intelligence data—and secure AI-driven biodesign tools to mitigate potential biosecurity risks."},{"id":"1c3cb37e-2a00-80c2-b5c6-c2ef9f128a0c","name":"Watermarking AI Generated Protein Sequences","description":"Develop embedded watermarks in generative protein models to promote traceability of engineered sequences."}]},{"id":"1c1cb37e-2a00-80b5-8594-d53ea33db290","slug":"fragile-supply-chains-and-lack-of-backup-for-critical-infrastructure","name":"Fragile Supply Chains and Lack of Backup for Critical Infrastructure","description":"Many critical supply chains and infrastructure systems are fragile and lack robust backup mechanisms, leaving society vulnerable.","field":"Biosecurity","outcome":"A shock to one node stops cascading into lost industrial capability, because the failure points are mapped and the fallbacks are already in place.","outcome_rationale":"Their framing is fragility plus absent backup; the unlock is that a single-node failure stops being systemic.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Single-source dependence and inventory depth are measurable inputs, but resilience itself only reveals under a shock that has not happened.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"Working now","is_primary":0,"rationale":"Network cascade modelling to locate single points of failure works now and is not what is missing.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-80d4-b84d-d16c67dced35","name":"Optimize Circular and Robust Supply Chains","description":"Adopt circular economy principles and advanced design strategies to build more resilient, robust, and sustainable supply chains."},{"id":"1c3cb37e-2a00-802e-9833-ee4a4f6ef10b","name":"Civilization Reboot Toolkit","description":"Develop a comprehensive toolkit to reboot essential infrastructure and restore societal functions in the event of widespread failure."}]},{"id":"1c1cb37e-2a00-80e2-a028-ebaa685df3da","slug":"limited-tools-for-improving-individual-social-and-societal-epistemics-in-the-face-of-misinformation","name":"Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation ","description":"In an era of relentless information overload and pervasive misinformation—fueled by algorithms that prioritize fleeting engagement over meaningful value—we have the opportunity to reshape our digital spaces. By leveraging AI and more intentional design of social media spaces for epistemic improvement, we can empower users to curate, evaluate, and contextualize content more effectively to create a healthier digital world.","field":"Social Science","outcome":"Platform operators get an objective other than engagement that they can actually implement: information environments designed to reward being right.","outcome_rationale":"Their description contrasts algorithms optimising for fleeting engagement with design for epistemic improvement.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Belief accuracy, exposure diversity and correction uptake are measurable, but collective epistemic health is inferred from these rather than observed.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Reading and synthesis vs Coordination and institutions; maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Reading and synthesis vs Coordination and institutions; maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-803f-b732-c20815089b6b","name":"Crowd-Sourced Fact-Checking ","description":"Build on initiatives like Community Notes–a crowdsourced fact-checking system that attaches contextual annotations to tweets–to build robust community-driven evaluation tools for social media platforms. This allows users to create and vote on annotations, while an open source algorithm determines which context notes to attach."},{"id":"1c3cb37e-2a00-805b-823e-e81416816f79","name":"Contextualization Engines","description":"Develop systems that provide context to content, helping users understand the broader background and counteract misinformation without censorship."},{"id":"1c3cb37e-2a00-8045-9957-ed3956e3dd70","name":"Detection of Coordinated Misinformation and Scams","description":"Develop digital tools to shield individuals from online scams and fraudulent schemes, and that detects and filters or flags coordinated misinformation campaigns."},{"id":"1c3cb37e-2a00-80ed-a19e-d545fb316961","name":"Content Curation Tools","description":"Develop tools that empower users to both create personalized information streams and collaboratively curate content. \nUsers can set up custom feeds—filtered by semantic content, social network data, and engagement signals—to tailor the information they receive. "},{"id":"1c3cb37e-2a00-80c0-86ba-fd683a74a625","name":"Incentives for Verification and Accountability","description":"Create market-based mechanisms that incentivize accurate fact-checking and hold sources accountable."},{"id":"1c3cb37e-2a00-8089-9a06-f03ed07b6cf3","name":"“System 2” Recommenders in Social Media Spaces","description":"Develop recommender systems that prioritize content we’ll appreciate upon reflection, rather than content that only captures our immediate attention. "},{"id":"1c3cb37e-2a00-8087-9222-cb8fb1dbc758","name":"Tools to Promote Constructive Dynamics in Social Media Spaces","description":"Develop “systems to  increase mutual understanding and trust across divides, creating space for productive conflict, deliberation, or cooperation”"},{"id":"1c3cb37e-2a00-80f5-b71c-f54a5fb5504a","name":"Improved Observatory of the Information Environment","description":"Build centralized platforms to monitor and analyze the information ecosystem, enabling better identification of misinformation trends."},{"id":"1c3cb37e-2a00-80df-a75c-eb2a4f073153","name":"Structured Transparency","description":"Implement frameworks that allow actors to reduce collaboration risks and costs by defining and enforcing precise flows of information. This allows a larger negotiating space for strategic actors dealing with advanced technologies. "},{"id":"1c3cb37e-2a00-8048-99ed-d6c0bc89d890","name":"Projects to Crowdsource Ground Truth Information","description":"Global distributed crowdsourced public knowledge and knowledge graphs"},{"id":"1c3cb37e-2a00-80bd-abd8-c924eccbff2f","name":"Privacy Preserving Identity Verification","description":"Robust identity verification systems that safeguard personal privacy."}]},{"id":"1c1cb37e-2a00-8055-b34e-fe7956440682","slug":"labor-replacing-ai-could-lead-to-human-disempowerment","name":"Labor-Replacing AI Could Lead to Human Disempowerment","description":"As AI systems become the cornerstone of competitive advantage, they can inadvertently marginalize human roles and decision-making. The drive for efficiency and cost reduction may lead organizations to rely predominantly on AI, sidelining human judgment, creativity, and accountability. This dynamic risks creating environments where economic and social inequities widen, and the intrinsic value of human input is systematically undermined (see examples). The gradual disempowerment of individuals under such competitive pressures poses significant challenges for societal well-being and democratic governance. \n\nSee: https://gradual-disempowerment.ai/","field":"Social Science","outcome":"Erosion of human influence over consequential decisions becomes visible while it is still reversible.","outcome_rationale":"Their capability list includes metrics to track human influence and early-warning signs, which makes visibility the near-term unlock.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Candidate metrics for human influence are named in their own capability list and could be computed, but no agreement exists that any of them would establish disempowerment rather than ordinary automation. Sits close to counterfactual-required, since the comparison is a world that did not occur, hence the guess.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"2-5 years","is_primary":0,"rationale":"Their proposed AI agents advocating for human interests and deliberation tools are language-model applications.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"Speculative","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"Speculative","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-80b6-95ba-e9105ec21743","name":"Metrics to Track Human Influence","description":"We need comprehensive methods to detect and quantify human disempowerment including economic, cultural, and political metrics as well as research and education."},{"id":"1c3cb37e-2a00-806c-a3cf-f293481902b8","name":"Tools for Deliberative Democracy","description":"Develop AI-driven platforms that facilitate deliberative democracy by aggregating diverse perspectives, advancing community moderation, and guiding public decision-making processes to move beyond corporate and partisan pressures."},{"id":"1c3cb37e-2a00-80e3-bfbd-f6722332930c","name":"Improved Voting and Auditing Protocols","description":"Develop next-generation voting systems and auditing protocols that are secure, transparent, and capable of supporting robust collective decision-making."},{"id":"1c3cb37e-2a00-8023-b6d2-c5bbb41599d3","name":"Early Warning Signs for Human Disempowerment","description":"Tools to forecast and monitor key thresholds or tipping points beyond which human influence becomes critically compromised, and the ability to measure effectiveness of intervention strategies."},{"id":"1c3cb37e-2a00-80a3-9617-dc436a21a54a","name":"Interventions for Maintaining Human Oversight","description":"Develop direct interventions for preventing accumulation of excessive AI influence: \n• Regulatory frameworks mandating human oversight for critical decisions, limiting AI autonomy in specific domains, and restricting AI ownership of assets or participation in markets\n• Progressive taxation of AI-generated revenues both to redistribute resources to humans and to subsidize human participation in key sectors\n• Cultural norms supporting human agency and influence, and opposing AI that is overly autonomous or insufficiently accountable"},{"id":"1c3cb37e-2a00-803c-aa2a-d32f9b6d74da","name":"AI Agents to Advocate for Human Interests","description":"Develop AI delegates who can advocate for people's interest with high fidelity, while also being better able to keep up with the competitive dynamics that are causing the human replacement."},{"id":"1cccb37e-2a00-80d0-90ce-f08e3adefc9c","name":"Technology to Augment Humans","description":"Augmentation technology, from better interfaces to AI tools to brain computer interfaces (BCIs), can amplify human strengths, democratize high-skill work, enable greater oversight, and help make humans more economically capable.\nSee BCI-related Foundational Capabilities:\n• Minimally Invasive Ultrasound–Based Whole Brain Computer Interface\n• Fully Noninvasive Read–Write Technologies\n• Micro- to Nano-Scale Minimally Invasive BCI Transducers"}]},{"id":"1c1cb37e-2a00-801b-bbfd-c904de35f2b4","slug":"our-platforms-for-civic-engagement-and-democratic-decision-making-dont-take-advantage-of-21st-century-scalable-technology","name":"Our Platforms for Civic Engagement and Democratic Decision-Making Don’t Take Advantage of 21st Century Scalable Technology","description":"Our current systems for democratic participation are hindered by outdated platforms and tools that fail to scale with modern needs. Limited survey infrastructure, insecure voting methods, and under-informative deliberative tools restrict our capacity for informed, collective decision-making. By harnessing AI to facilitate clearer expression of public opinion and leveraging innovative technologies for secure, scalable engagement, we can transform civic participation into a more robust, effective, and inclusive process.","field":"Social Science","outcome":"Deliberation stops being limited to the number of people who fit in a room. Public preferences get elicited and aggregated at population scale, verifiably.","outcome_rationale":"Their description names survey infrastructure, voting security and deliberative tools as the three limits.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Participation rates and demographic representativeness are measurable, but whether collective decisions improved has no agreed observable.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). LLM-mediated deliberation works in pilots — vTaiwan, the Habermas Machine — but the gap names the platforms that actually carry civic decisions, and no national system has adopted any of it. Distinct from clinical trials, where the institutional innovation did reach scale. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). LLM-mediated deliberation works in pilots — vTaiwan, the Habermas Machine — but the gap names the platforms that actually carry civic decisions, and no national system has adopted any of it. Distinct from clinical trials, where the institutional innovation did reach scale. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-80e3-bfbd-f6722332930c","name":"Improved Voting and Auditing Protocols","description":"Develop next-generation voting systems and auditing protocols that are secure, transparent, and capable of supporting robust collective decision-making."},{"id":"1c3cb37e-2a00-806c-a3cf-f293481902b8","name":"Tools for Deliberative Democracy","description":"Develop AI-driven platforms that facilitate deliberative democracy by aggregating diverse perspectives, advancing community moderation, and guiding public decision-making processes to move beyond corporate and partisan pressures."}]},{"id":"1c1cb37e-2a00-8013-a845-c84eb5cde804","slug":"policy-creation-and-evaluation-is-manual-and-suffers-from-low-efficiency-and-accountability","name":"Policy Creation and Evaluation is Manual and Suffers from Low Efficiency and Accountability","description":"Policy development and evaluation processes today rely heavily on manual human review to ensure accountability. However, as AI systems increasingly support or automate these processes, this human-centered accountability becomes challenging. Human reviewers risk becoming a critical bottleneck, slowing policy implementation. New tools are needed to streamline policy creation and evaluation, and to ensure consistency and compliance before deployment.","field":"Social Science","outcome":"Policy gets written formally enough to check automatically for consistency and compliance, and human review stops being the throughput limit.","outcome_rationale":"Their description identifies human reviewers becoming a bottleneck as the specific failure.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Time from draft to cleared policy, and detected inconsistencies per document, are direct process metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-8089-b88f-e1232ffdec29","name":"Formalization of Policy","description":"Build tools to enable policy-makers to write mechanized (i.e. runnable software) versions, moving more of the subjective evaluation ahead of the action (rather than interpreting more things post-facto). This would leverage formal logic to streamline policy development and governance."}]},{"id":"1c1cb37e-2a00-800f-9e2a-cf61f6212c7e","slug":"underdevelopment-of-modern-tools-in-the-social-sciences","name":"Underdevelopment of Modern Tools in the Social Sciences","description":"The social sciences need new tools to help researchers identify and prioritize important questions that will have an impact, and better infrastructure to collect qualitative data. Qualitative methods are powerful for understanding the how and why behind social outcomes, yet even the most comprehensive surveys don’t capture all the factors that contribute to social outcomes. AI-enabled qualitative methods could super-charge the social sciences, but there is much work to be done.\n\nSimilarly, many archaeological methods remain manual and lack the technological revolution seen in other fields, limiting discovery and analysis.","field":"Social Science","outcome":"The how and why of social outcomes stops being the underpowered half of the field. Qualitative and survey evidence scales the way quantitative data already does.","outcome_rationale":"Their description names qualitative methods as powerful but unscaled, and archaeology as similarly un-modernised.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Interviews coded per researcher-hour and inter-coder agreement are measurable, but whether the field is asking better questions has no observable.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":0,"rationale":"Satellite and machine-learning-enabled archaeology is a detection problem already producing results.","confidence":"confident"},{"ai_type":"Reading and synthesis","maturity":"Working now","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"Working now","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-8036-bd26-fff62f3c54f2","name":"Infrastructure for Problem-Driven Research","description":"A framework to help researchers identify and pursue meaningful questions. This includes support structures—such as fellowship programs to train students in problem-driven research workflows, and innovative funding mechanisms to empower researchers to focus on questions that truly matter."},{"id":"1c3cb37e-2a00-80b5-8c03-f4525550cb12","name":"Modernize Survey Infrastructure","description":"Develop modern, scalable survey platforms that can efficiently capture and analyze public opinion data and other qualitative data for social sciences. New tools are needed to collect, analyze, curate, and model large-scale qualitative data."},{"id":"1c3cb37e-2a00-80b4-bea0-c9314c934760","name":"Satellite and Machine Learning-Enabled Archaeology","description":"Apply satellite imagery and ML techniques to modernize archaeological surveys and analysis, enabling more rapid and systematic discoveries."}]},{"id":"1c1cb37e-2a00-8097-9eca-e48345e90acc","slug":"education-modalities-suffer-from-scaling-limitations","name":"Education Modalities Suffer From Scaling Limitations","description":"Current education systems face structural inefficiencies such as excessive administrative workloads on educators, overcrowded classrooms, and inequitable resource distribution. Innovative technologies have the potential to significantly reduce these burdens by providing tools that assist teachers with scheduling, grading, and creating personalized, adaptive lesson plans. Digital platforms could dynamically tailor learning experiences to individual student progress, complementing classroom teaching. Additionally, technology-driven improvements in administrative efficiency could free valuable resources, enhancing educational equity and overall student experiences.\n\n“US K-12 teachers are 30% more likely to face burnout than U.S. soldiers, whose lives are defined by relentless duty, perpetual war and low wages.” - Adrienne Williams \n\n“Given recent improvements in the quality, affordability, and usability of technologies like AI, computer vision, and AR/VR, we can reimagine a more personalized, research-driven K-12 experience—one better able to meet diverse learning needs and set students up for success both in school and in life. From chatbots for individual coaching to immersive mixed reality solutions to adaptive technologies supporting culturally responsive pedagogy, we have an opportunity to leverage new technologies to tackle pressing needs across literacy education, STEM instruction, and preparing students for the workforce.” - Kumar Garg ","field":"Social Science","outcome":"One-to-one teaching at the price of one-to-many. Teaching quality stops being rationed by teacher hours per student.","outcome_rationale":"Their description names administrative burden and class size as the structural limits that technology would lift.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Learning gains against control, teacher hours reclaimed and cost per student are measurable and routinely trialled.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"Working now","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"Working now","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-80c4-992a-d9b1e7e9748d","name":"Digital Tutor Technologies","description":"Develop sophisticated digital tutor systems that offer personalized learning assistance by adapting to individual learning styles using AI."},{"id":"1c3cb37e-2a00-80de-9d21-f9456cf622c4","name":"Digital Tools for Educators and School Districts","description":"Tools for lesson planning and administrative tasks. Tools for administration of school systems, including managing resource distribution."},{"id":"1c3cb37e-2a00-80da-83db-c653aed8ea20","name":"New Architectures for Learning Assistance","description":"Explore novel architectures and settings for AI assisted learning"}]},{"id":"1c1cb37e-2a00-80c2-a1bb-c735bf51fdd9","slug":"underdevelopment-of-deep-tooling-for-economic-modeling-and-future-forecasting","name":"Underdevelopment of Deep Tooling for Economic Modeling and Future Forecasting","description":"Current economic models are often too simplistic to capture the intricate dynamics of our global economy, limiting effective policy-making and forecasting. Experimentation with innovative economic models—such as those incorporating universal basic income or alternative market systems—is rare, leaving us unprepared for emerging trends. The inherent complexity of global systems further complicates accurate forecasting, underscoring the urgent need for more sophisticated, adaptive tools that can better predict and navigate the economic landscape of tomorrow.","field":"Social Science","outcome":"Economic models represent structural transformation rather than marginal change, so policy for unprecedented conditions can be tested before it is lived through.","outcome_rationale":"Their description names modelling radical transformations and unpreparedness for emerging trends as the loss.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Forecast accuracy is measurable at short horizons, but models of radical transformation cannot be validated against an economy that has not undergone one.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-8023-a74b-d0f38b2b559e","name":"Real-World SimCity with AI Agents","description":"Create complex simulation models using AI agents that mimic real-world economic dynamics, providing a richer framework for policy analysis."},{"id":"1c3cb37e-2a00-80bf-af77-e2b1e8c8c2e0","name":"Complex Systems Simulation Techniques for Economics","description":"Utilize techniques from chaos and complex systems theory to develop more realistic economic models that account for non-linear dynamics."},{"id":"1c3cb37e-2a00-806e-ad6e-cab9270d32f9","name":"Model Radical Economic Transformations","description":"Develop simulation models to explore scenarios of profound economic change to help plan for disruptive shifts. For example: a future where extended youth replaces traditional retirement and end-of-life expenses are significantly reduced, assessing the economic impact of a major AI transition."},{"id":"1c3cb37e-2a00-80ec-b33e-f94f3ae319c0","name":"Prediction Markets and AI-Based Forecasting","description":"Develop and refine prediction markets augmented by AI to improve future forecasting accuracy and better inform policy decisions."}]},{"id":"1c1cb37e-2a00-80e6-a011-db27cb962cd8","slug":"translational-gaps-in-development-economics","name":"Translational Gaps in Development Economics","description":"There exists a disconnect between academic research and the practical implementation of development economics, hampering the conversion of theoretical insights into effective real-world interventions. Current mechanisms for delivering public goods and fostering collective cooperation are inefficient, limiting our capacity to coordinate resources and drive meaningful change. Innovative approaches are needed to bridge this gap, ensuring that cutting-edge economic theories can be transformed into actionable policies and scalable interventions that truly improve development outcomes.","field":"Social Science","outcome":"Development economics findings get implemented at the scale they were validated for. Evidence currently accumulates faster than anyone uses it.","outcome_rationale":"Their description names the disconnect between academic research and practical implementation as the gap itself.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Adoption counts and programme reach are measurable, but the welfare gain forgone by non-translation is not observed.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":0,"rationale":"Cost-effectiveness modelling and large-scale economic experiments in their capability list are prediction and simulation work.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Algorithmic targeting has worked where a government adopted it — Togo's Novissi transfers are a real instance. The blocker the gap names is that adoption, which depends on state capacity and political will, not on the technique. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Algorithmic targeting has worked where a government adopted it — Togo's Novissi transfers are a real instance. The blocker the gap names is that adoption, which depends on state capacity and political will, not on the technique. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-80f7-b6e9-e06114fdf9e6","name":"Enhance Operational Approaches in Development Economics","description":"Bridge the translational research gap by adopting more operational, nonacademic approaches to development economics, moving beyond traditional models like JPAL."},{"id":"1c3cb37e-2a00-8062-819f-d3e6f4453f9e","name":"Long-Run Followups on Major Economic Studies","description":"Most major economic studies include data collection for 1-2 years, but the outcomes (e.g. education, child health) could plausibly impact large sections of one’s life.  There is little funding for followups on existing studies."},{"id":"1c3cb37e-2a00-8054-8535-ff41613f656c","name":"Experiments in Alternative Economic Structures","description":"Run controlled experiments—such as basic income trials and alternative market models for carbon pricing—to test and refine innovative economic structures."},{"id":"1c3cb37e-2a00-8020-9022-fbf332f10d35","name":"Large-Scale Economic Experiments in Multiplayer Games","description":"Utilize multiplayer online games as testbeds for large-scale economic experiments, enabling controlled studies of human behavior in dynamic environments."},{"id":"1c3cb37e-2a00-808f-9622-e31df00ada7b","name":"Invent New Decentralized Coordination and Contract Mechanisms","description":"Design and implement novel, decentralized models for public goods allocation and cooperation that overcome existing systemic inefficiencies."},{"id":"1c7cb37e-2a00-8056-8855-f053debb8e30","name":"Cost Effectiveness and Techno-Economics Calculations","description":"Make it easier to calculate cost effectiveness of interventions without a large staff through AI assisted methods"}]},{"id":"1c1cb37e-2a00-801b-bd01-c3b284abeb44","slug":"ephemeral-societal-data-on-proprietary-platforms","name":"Ephemeral Societal Data on Proprietary Platforms","description":"Much critical data is stored on proprietary platforms and is at risk of disappearing, hindering long-term research and reproducibility.","field":"Social Science","outcome":"Socially significant data outlives the platform that held it, and longitudinal research stops depending on corporate retention policy.","outcome_rationale":"Their description names long-term research and reproducibility as what disappearing data forecloses.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Datasets successfully archived is countable, but the quantity of interest is what has already been lost, which by construction left no record.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"Maturity decided by David, 2026-08-23, overriding a label both independent passes agreed on. Web-scale archiving has been technically solved for years and the data is still disappearing, which is the proof that the blocker is institutional rather than technical. Working now would say the problem is scaling a working solution; it is not, it is permission and who funds preservation.","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"Maturity decided by David, 2026-08-23, overriding a label both independent passes agreed on. Web-scale archiving has been technically solved for years and the data is still disappearing, which is the proof that the blocker is institutional rather than technical. Working now would say the problem is scaling a working solution; it is not, it is permission and who funds preservation.","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c7cb37e-2a00-8003-9a14-d43d9eafa0f0","name":"Create an Internet Archive for Critical Data","description":"Establish initiatives to systematically archive and preserve essential datasets from proprietary platforms, ensuring long-term accessibility and reproducibility."}]},{"id":"1c1cb37e-2a00-8007-8b2a-c8c51e4c39d0","slug":"fraud-in-the-scientific-literature","name":"Fraud in the Scientific Literature","description":"Scientific literature is plagued by fraudulent publications, undermining trust and slowing progress.","field":"Metascience","outcome":"Fabricated work gets caught before other work is built on top of it.","outcome_rationale":"Their description names trust and slowed progress as the two losses, and both come from undetected fraud propagating.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Retractions and detected fraud are countable, but they measure detection capability rather than fraud prevalence. The denominator, undetected fraud, is by definition unobserved.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"Working now","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"Working now","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[{"id":6,"quantity":"Retractions per 10,000 published articles, averaged over 2000-2024","current_value":"11.82","unit":"retractions per 10,000 articles","as_of":"2024-12-31","target_value":null,"target_basis":null,"source_title":"Zhou, Lou, Shen & Li, 'Mapping Academic Integrity: Global Retraction Trends Explored through a Topic Lens', arXiv:2511.21176v2 (6 July 2026). Preprint, not peer reviewed.","source_url":"https://arxiv.org/abs/2511.21176","source_doi":null,"source_checked":"verified","reads_as":"Averaged across 2000-2024, about 12 papers in every 10,000 were eventually retracted. For papers published in 2022 alone the rate reached 42.58 per 10,000.","direction":"ambiguous","context":"Web of Science (SCI, SSCI, ESCI), 2000-2024: nearly 42 million articles, of which over 49,000 were retracted. Retractions compounded at 22.29% a year against 6.24% for publications. By discipline the rate runs from 3.49 per 10,000 in Physics to 33.97 in Electrical Engineering & Computer Science. Earlier studies put the rate between 2 and 8 per 10,000, so estimates vary several-fold with the database and the denominator.","caveat":"This counts detection, not fraud. It rises when policing improves and when fraud increases, and the number alone cannot tell you which — that is why this gap is tiered Proxy only. It is also an average over twenty-five years and must not be read as a current rate: the source puts 2022 nearly four times higher. Recent years are additionally depressed by the lag between publication and investigation.","is_null_result":0,"rationale":"REVISED after Gate A, which graded the previous version wrong-quantity and blocking, correctly. The row had recorded 11.04 per 10,000 with as_of 2025-08-31 and read it as the present-day rate. It is not a present-day rate. The source states it as an average over 2000-2024, and says in its own words that 'the retraction rate in 2022 reached 42.58 per 10,000, markedly exceeding the average rate of 11.82 observed from 2000 to 2024'. A long-run average presented as current understated the recent figure roughly fourfold.\n\nRe-fetching the source also showed the number itself had moved. v2 (6 July 2026) reports over 49,000 retractions and an overall rate of 11.82 per 10,000; the 11.04 figure, and the 55,000-paper corpus and 22.09%/6.25% growth rates in the previous rationale, are all from v1 and are superseded. The as_of is now the close of the source's own measurement window rather than a date attached to a different sentence.\n\nOne part of the gate's finding did not hold and is recorded so it is not re-fixed: it reported the discipline range 0.035%-0.34% as unsupported, against 3.25 and 31.97 per 10,000 in the body. v2's abstract gives exactly 0.035% in Physics to 0.34% in Computer Science, and its body gives 3.49 and 33.97 per 10,000 — the same figures rounded. The range was read off the source correctly.\n\nNo target is recorded, and that is deliberate: a lower number is not unambiguously better, so writing one down would assert the opposite. Still marked a guess — the source remains a preprint.","confidence":"guess"}],"capabilities":[{"id":"1c7cb37e-2a00-8057-8c61-c5453214d1a0","name":"Automated Scientific Fraud Detection","description":"Develop AI and multimodal LLM systems to automatically detect fraudulent research, flag suspicious publications, and improve overall scientific integrity.\n\nEspecially as rapid advancements in AI models make it feasible to generate inaccurate scientific content at scale to disrupt.\n"}]},{"id":"1c1cb37e-2a00-8012-ae64-f8ff6f67ee0b","slug":"inability-to-comprehend-and-synthesize-the-entire-scientific-literature-at-scale","name":"Inability to Comprehend and Synthesize the Entire Scientific Literature at Scale","description":"The volume of scientific publications is overwhelming, making it difficult for humans to read, comprehend, and synthesize the entire body of literature. How can AI-generated knowledge become cumulative? What should a machine-human shared Wikipedia look like? We should collect and synthesize all the world’s knowledge, accelerate its development, and make it universally available in a compelling form.","field":"Metascience","outcome":"The literature can be queried as one body of knowledge rather than searched as a pile of documents, and what is already known stops being rediscovered.","outcome_rationale":"Their description asks how AI-generated knowledge becomes cumulative, which is the unlock beyond mere retrieval.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Synthesis quality on benchmark questions is measurable, but their actual target, whether machine-generated knowledge becomes cumulative, has no agreed observable.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"Working now","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Reading and synthesis","primary_maturity":"Working now","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c7cb37e-2a00-80b5-a6b6-d72a0e6d6596","name":"AI Scientific Literature Agents","description":"Develop AI agents that can read, summarize, and integrate scientific literature, providing researchers and policymakers with synthesized insights."},{"id":"1c7cb37e-2a00-80e8-8c20-c5615927d5e4","name":"Knowledge Synthesis Infrastructure","description":"Build robust infrastructure to integrate and synthesize diverse scientific findings into cohesive, accessible formats for researchers."},{"id":"1c7cb37e-2a00-8084-a621-f55ac259c6b8","name":"Intelligent Databases","description":"Intelligent databases and automatic probabilistic integration across multiple databases"},{"id":"1c7cb37e-2a00-80d5-8d2c-d403864922db","name":"Large Knowledge Models","description":"Knowledge models that can facilitate reasoning by synthesizing and clarifying relevant information transparently from multiple domains.\n\n“Provide a semantic medium that is both more expressive and more computationally tractable than natural language, a medium able to support formal and informal reasoning, human and inter-agent communication, and the development of scalable quasilinguistic corpora with characteristics of both [scientific] literatures and associative memory”.\n\nLLMs have powerful capabilities but their knowledge is opaque,  not cumulative, and not easily updatable and comparable. Human society has not only individual brains with memory but a cumulative scholarship to grow knowledge, compare alternative views, etc.\n\nSuch AI could be used as a collective strategic assistant. "},{"id":"1c7cb37e-2a00-8038-9f8a-e18059923bf7","name":"Infrastructure for Research Curation","description":"Infrastructure to support and incentivize the diverse curation of research through science social media and/or dedicated spaces. This infrastructure would support rapid dissemination of research and encourage broader exploration of the research landscape, mitigating the risk of homogenous research focus, and maladaptive collective attention patterns in science."}]},{"id":"1c1cb37e-2a00-8063-9b9f-f3c2aa784e50","slug":"doing-and-publishing-research-is-expensive-and-subject-to-structural-roadblocks","name":"Doing and publishing research is expensive and subject to structural roadblocks","description":"Traditional structures dominate in how research is conducted and how its outputs are disseminated. Expensive publishing practices restrict and slow the spread of knowledge. We should replace outdated publishing practices and complement research practices with new approaches that leverage frugal innovation, community-led platforms, and open access. We imagine a future where scientific discovery is more inclusive and dynamic.","field":"Metascience","outcome":"Getting a result into the accepted, verified scientific record is cheap and quick.","outcome_rationale":"Axis chosen is cost; their sentence bundles cost, speed and inclusiveness, and the other two are separate chains.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Article processing charges, total publishing spend, time from submission to publication and reviewer invitations per accepted review are all directly reported.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"Working now","is_primary":0,"rationale":"Drafting, screening and reviewer-matching assistance work today, which is precisely why they are not the binding link.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Launching a cheap open platform is available this afternoon; shifting the prestige and incentive economy that keeps publishing expensive is not. Open access has moved real ground over a decade, so the path exists, but nothing applied today closes it. Type unchanged from the mechanical adjudication.","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). Launching a cheap open platform is available this afternoon; shifting the prestige and incentive economy that keeps publishing expensive is not. Open access has moved real ground over a decade, so the path exists, but nothing applied today closes it. Type unchanged from the mechanical adjudication.","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c7cb37e-2a00-80ba-9a56-e74ba0d95dfb","name":"Post-Publication Peer Review Layer","description":"Establish a community-driven, post-publication peer review system that supplements traditional publishing, enhancing transparency and accountability."},{"id":"1c7cb37e-2a00-803b-b89f-c5094b5dfd83","name":"Disrupt Traditional Publishing Models","description":"Support modern, open, and community-driven publishing platforms that challenge and replace outdated models."},{"id":"1c7cb37e-2a00-803a-a33e-d146c8783c98","name":"Frugal Science Initiatives","description":"Promote and support frugal science projects that empower communities worldwide to participate in scientific research, especially in developing countries."},{"id":"1c7cb37e-2a00-8048-96c1-e7cb85103c80","name":"New Protocols for Knowledge Production and Verification","description":"Combination of new tools and norms that are situated outside of the traditional publishing pipeline: protocols  for empowering and recognizing citizen science on social media. New practices around micro/nanopublishing to lower the barrier to publishing"}]},{"id":"1c1cb37e-2a00-80df-ae9c-cc561c437a83","slug":"clinical-trials-are-poorly-optimized-for-evidence-gathering","name":"Clinical Trials Are Poorly Optimized for Evidence Gathering","description":"Current clinical trial designs are not sufficiently optimized for gathering robust evidence, leading to inefficiencies and suboptimal outcomes.","field":"Metascience","outcome":"More evidence per patient and per dollar. The number of clinical questions answerable in a decade rises without a matching rise in cost.","outcome_rationale":"Their description frames the loss as inefficiency in evidence gathering rather than as a shortage of trials.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Cost per trial, time to result, statistical power achieved and the share of trials ending inconclusive are all recorded in trial registries.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":0,"rationale":"Digital twins and synthetic control arms are prediction problems, and both appear in their capability list.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"Maturity decided by David, 2026-08-23. Bayesian adaptive designs have existed since 2010 and are in live use, and practitioners running them report they do not speed things up much — so the mechanism being deployed somewhere is not the same as it moving this gap. The residual blocker is institutional, in the same way as scientific publishing and civic deliberation, and gets the same label.","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"Maturity decided by David, 2026-08-23. Bayesian adaptive designs have existed since 2010 and are in live use, and practitioners running them report they do not speed things up much — so the mechanism being deployed somewhere is not the same as it moving this gap. The residual blocker is institutional, in the same way as scientific publishing and civic deliberation, and gets the same label.","primary_confidence":"confident","indicators":[{"id":4,"quantity":"Median estimated cost of a pivotal clinical trial supporting a new FDA approval","current_value":"19.0","unit":"million USD","as_of":"2016-12-31","target_value":null,"target_basis":null,"source_title":"Estimated Costs of Pivotal Trials for Novel Therapeutic Agents Approved by the US Food and Drug Administration, 2015-2016","source_url":"https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2702287","source_doi":"10.1001/jamainternmed.2018.3931","source_checked":"verified","reads_as":"Half the pivotal trials behind new FDA approvals cost less than 19 million dollars, half cost more.","direction":"lower is better","context":"Trials measuring clinical outcomes averaged 64.7 million against 24.0 million for surrogate endpoints. Evidence that answers the question costs roughly 2.7 times evidence that does not.","caveat":"2015 to 2016 data, and the publisher blocks automated link checks even though the DOI verifies.","is_null_result":0,"rationale":"138 pivotal trials across 59 novel therapeutic agents approved in 2015–2016, median estimated cost 19.0 million USD, IQR 12.2–33.1 million. The same paper shows why cost is the right axis for a gap about evidence-gathering efficiency rather than a proxy for it: trials measuring clinical outcomes averaged 64.7 million against 24.0 million for surrogate endpoints, so the price of the evidence that actually answers the question is roughly 2.7 times the price of the evidence that does not. No target is recorded. A tempting one exists — a separate dataset of 524 US-funded phase 3 trials gives a median near 2.5 million — but that population is publicly funded trials rather than industry pivotal trials, and pretending the two are the same quantity would manufacture a 7-fold improvement target out of a sampling difference. Marked confident on the value and deliberately empty on the target. One sourcing note, recorded rather than hidden: the JAMA article page was fetched and read for these figures, but the publisher returns HTTP 403 to non-browser clients, so an automated link check flags it as unreachable while a human clicking it sees the paper. The DOI verifies against Crossref, which is why source_checked reads 'verified' and the URL check does not. Gate A re-fetched this source and confirmed every value, including the $64.7M and $24.0M context figures. It also found that the claim recorded here about JAMA returning 403 to automated clients does not reproduce — the full body fetched cleanly — so source_checked is 'verified' on a fetch the reviewer repeated, not on the original assertion.","confidence":"confident"}],"capabilities":[{"id":"1c7cb37e-2a00-8061-a74c-f427a8541c44","name":"Increase Predictive Validation of Early Studies","description":"Enhance early clinical studies by integrating predictive validation techniques to forecast outcomes and optimize trial design."},{"id":"1c7cb37e-2a00-809e-8052-f9515d089115","name":"Quasi-Experimental Causality in Biomedical Research","description":"Apply quasi-experimental designs to strengthen causal inference in clinical studies, improving evidence quality, and do this for wider and more unified datasets (e.g., existing medical records) and across many conditions, e.g., combined with federated data approaches.\n\nHistorically it has taken 5-10 years for advanced methods to percolate into relevant areas of omics / biotechnology x clinical area. It is also changing a culture of thinking — that there exists a different kind of validation that is neither 'do a perfect experiment' nor 'I tested on an external hold-out' but a third thing."},{"id":"1c7cb37e-2a00-801b-b0df-db9928460a72","name":"Federated Data Approaches","description":"Implement federated data and differential privacy systems to aggregate clinical trial data from multiple sites while ensuring patient privacy."},{"id":"1c7cb37e-2a00-805d-8fcc-e75f3d0c3201","name":"Out-of-Clinic Studies","description":"Promote studies conducted outside traditional clinical settings to gather real-world evidence more effectively."},{"id":"1f0cb37e-2a00-80db-940f-de96a93d37f6","name":"Infrastructure for Decentralized Patient-Reported Clinical Studies","description":"Decentralized patient reporting for clinical studies would amplify and diversify patient participation at decreased cost by reducing logistical and geographic barriers. It could compound and incorporate subjective patient experiences at large scale and would improve public accessibility of trial data."}]},{"id":"1c1cb37e-2a00-80b5-8cf4-c8195273fa45","slug":"a-limited-set-of-rigid-organizational-structures-for-organizing-and-funding-research-constrains-the-forms-of-rd-that-get-done","name":"A Limited Set of Rigid Organizational Structures for Organizing and Funding Research Constrains the Forms of R&D That Get Done","description":"This is more of a meta-bottleneck. But scientists are spending a lot of time not doing science, and the institutional structures in which they work are often set up with incentive structures that hinder certain kinds of outcomes, like more coordinated research.","field":"Metascience","outcome":"Long-horizon, heavily coordinated research that no current institutional form can house becomes fundable and organisable.","outcome_rationale":"Their description names incentive structures that hinder certain kinds of outcomes, coordinated research specifically.","outcome_confidence":"confident","tier":"Counterfactual required","tier_rationale":"The quantity of interest is research that never happened because no institution could house it. There is no observation set for unproposed work, and proposals never written leave no record.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Coordination and institutions","maturity":"Speculative","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 Speculative / v2 Working now; mechanical rule had given Working now). Redesigning how research is funded and organised has no demonstrated mechanism at all, AI-mediated or otherwise. This is the map's counterfactual-required gap: nothing applied today moves it, and no clear path runs from what exists to moving it. Type unchanged from the mechanical adjudication.","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"Speculative","primary_rationale":"Maturity re-adjudicated on the merits (v1 Speculative / v2 Working now; mechanical rule had given Working now). Redesigning how research is funded and organised has no demonstrated mechanism at all, AI-mediated or otherwise. This is the map's counterfactual-required gap: nothing applied today moves it, and no clear path runs from what exists to moving it. Type unchanged from the mechanical adjudication.","primary_confidence":"confident","indicators":[{"id":8,"quantity":"Progress toward less rigid organizational and funding structures for research","current_value":null,"unit":null,"as_of":null,"target_value":null,"target_basis":null,"source_title":null,"source_url":null,"source_doi":null,"source_checked":null,"reads_as":null,"direction":null,"context":null,"caveat":null,"is_null_result":1,"rationale":"NULL CORROBORATED by Gate A, and this is now the strongest row in the sample. A second searcher, working from the gap description alone and without sight of the queries below, ran eight independent searches of its own wording and also found no maintained series measuring which forms of R&D get done. Two searchers failing with different wording is a materially different claim from one searcher failing, and it is the reason to trust this row.\n\nThe original six searches (recorded in research-log/searches/phase-3.json, phase 3, this gap — the search_log database table is empty, so that file is the provenance): researcher time spent on administration; metrics for diversity of research funding mechanisms; counterfactual measurement of funding structure against discoveries; outcome metrics for new organizational forms including FROs; indicators of institutional innovation in science funding; and randomised experiments in grant allocation.\n\nNear-misses, all failing for stateable reasons. (1) Administrative burden: US federally funded principal investigators report spending roughly 42% of their research time on administration. The gate objected that the gap's own sentence opens 'scientists are spending a lot of time not doing science', so this measures the first clause and was rejected against the second. The objection is fair and the rejection still stands, but on narrower grounds than before: the 42% is a real measure of one symptom, and cutting it to 20% would be a genuine gain that left the gap's actual claim — which forms of R&D get done — untouched. It is the wrong quantity, not a trivial one, and the FDP survey behind it last ran in 2018. (2) NIH's median age at first R01-equivalent award, an official series updated annually, has held near 42 across FY21-25 (mean 43-44). It is maintained, current and genuinely about rigidity, which makes it a stronger near-miss than either of the two originally listed — and it still measures who gets funded rather than which forms of research become possible. (3) The randomised evidence on funding allocation, including grant lotteries, is the right instrument but is a set of individual studies, not a series, and its own conclusion is that counterfactual assessment of funding structures is usually impossible because researchers have alternative sources.\n\nThis is what Counterfactual required means: the quantity of interest is the research a different structure would have produced and this one did not, and no observation of the world we are in contains it. The null is not a search failure. It is the tier being correct.","confidence":"confident"}],"capabilities":[{"id":"1c7cb37e-2a00-803f-beeb-c7c0b19f4cd7","name":"Novel and Diverse Research Organization Structures ","description":"We should design research institutions to support the particular kinds of research they need to house, not fit every square peg into the same round hole."},{"id":"1c7cb37e-2a00-801e-96ef-c03df8d69833","name":"Decentralized Science Funding","description":"Diversify who can be a funder of science by launching new fast grants-style programs. "},{"id":"1c7cb37e-2a00-80ca-abe3-ea2550fdc7f0","name":"New Models to Reduce Fundraising Burden of Scientists","description":"There are various solutions that could help scientists spend less time fundraising and more time doing science. This could include new structures for research institutes, where researchers don’t need to apply for external grants, and new tools for decentralized science funding."}]},{"id":"1c1cb37e-2a00-8089-a828-f143cf6f5c86","slug":"under-provisioning-of-antibiotics-vaccines-and-other-interventions-for-major-global-health-challenges","name":"Under-Provisioning of Antibiotics, Vaccines and Other Interventions for Major Global Health Challenges","description":"Many of the world’s most deadly diseases—such as tuberculosis, Group A Streptococcus, hepatitis C, hepatitis B, and syphilis—lack effective vaccines or cures. Additionally, the pace of developing effective, low-cost, therapeutics for emerging pathogens in low resource settings is too slow to meet global health needs. Malnutrition exacerbates susceptibility to disease and impedes recovery; food security is important especially for early child development. Understanding the basic science of malnutrition during development is important for the development of more effective interventions.","field":"Global Health","outcome":"Tuberculosis and Group A Streptococcus have resisted vaccines for decades. They get effective countermeasures, and malnutrition's developmental mechanism turns into a design target.","outcome_rationale":"Their description names the specific pathogens and adds the basic science of malnutrition as a prerequisite for better intervention.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Vaccine efficacy in trial, disease incidence and cost per course are direct and long-tracked quantities.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Coordination and institutions","maturity":"Working now","is_primary":0,"rationale":"These are under-provisioned rather than undiscoverable; the market failure is well understood and pull mechanisms already exist.","confidence":"confident"},{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c7cb37e-2a00-809f-b187-c169c5c789f9","name":"Effective Vaccines for Global Health Challenges","description":"Accelerate clinical trials and support research into novel vaccine technologies that can be distributed in low-resource settings, especially those that do not require cold chain."},{"id":"1c7cb37e-2a00-80a2-8933-dedaa4b8a3da","name":"Novel Therapeutic Approaches for Chronic Infections","description":"New treatments for chronic infections (e.g., achieving a functional cure for hepatitis B) and novel monoclonal antibodies for diseases such as malaria.\n\nNew formulations (e.g., one time dose time release delivery) could improve delivery and compliance in low resource settings."},{"id":"1c7cb37e-2a00-80ef-9bc7-c3b83e978b6c","name":"Improved Interventions for Malnutrition","description":"We need to better understand the physiology of malnutrition and develop improved interventions."}]},{"id":"1c1cb37e-2a00-8045-ad83-cb1081a2eaf1","slug":"limited-diagnostic-tools-optimized-for-low-resource-settings","name":"Limited Diagnostic Tools Optimized for Low-Resource Settings","description":"Current diagnostic tests are costly or often offer only limited information, failing to reveal the cause of disease and delaying or preventing administration of available treatments. Moreover, early detection systems for emerging pathogens are fragmented, delaying critical public health interventions.","field":"Global Health","outcome":"A clinician learns the cause rather than the syndrome, at the point of care, and treatments that already exist reach the patients they would help.","outcome_rationale":"Their description names the specific loss: tests that do not reveal cause delay or prevent administration of treatments that already exist.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Sensitivity, specificity, cost per test and time to result are the standard diagnostic metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c7cb37e-2a00-8013-82fa-efbd8086d874","name":"Multiplexed Molecular Diagnostic Platforms","description":"Develop rapid, integrated diagnostic tests for many diseases at once that are low cost.\n \nFrom Jacob Trefethen’s essay:\n“I would count success in the diagnostic row as: a multiplex diagnostic for at least 3 pathogens (i.e. flu + COVID does not count), available over the counter for use at home. Either a respiratory panel (e.g. flu + COVID + strep throat) or a fever panel (e.g. malaria + dengue + typhoid) would count. An at-home multiplex STI panel would be great (e.g. chlamydia + gonorrhea + syphilis)”."},{"id":"1c3cb37e-2a00-805c-a65f-dcfa630b1b7e","name":"Volatilomics for Early Detection","description":"Utilize the analysis of volatile organic compounds (VOCs) as an early detection tool to identify pathogenic outbreaks."},{"id":"1c7cb37e-2a00-8056-9873-ea134135ee94","name":"Point of Care Diagnostics for Lead Testing","description":"Lead exposure remains a critical but under-addressed public health challenge. Approximately 1 in 3 children globally have toxic levels of lead in their bloodstream, leading to developmental delays, cognitive impairment, and increased risk of chronic diseases. However, as it stands most countries do not do comprehensive, routine surveillance of blood lead levels. This is because the primary method for blood lead surveillance is costly and inconvenient."}]},{"id":"1c1cb37e-2a00-80c0-aae9-fc48f6ef2af6","slug":"lack-of-infrastructure-technologies-and-strategies-optimized-for-low-resource-settings","name":"Lack of Infrastructure Technologies and Strategies Optimized for Low-Resource Settings","description":"Global health outcomes are compromised by insufficient health systems and infrastructure that limit our ability to prevent and control infectious diseases. Key deficiencies include the lack of cost-effective antimicrobial materials to block pathogen transmission, underdeveloped intervention models for effective public health strategies, and outdated sanitation solutions that fail to meet the needs of vulnerable populations.","field":"Global Health","outcome":"Infection control no longer assumes infrastructure that low-resource settings do not have, so prevention decouples from a country's health budget.","outcome_rationale":"Their description frames every deficiency as an infrastructure mismatch rather than a missing scientific result.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Cost per unit, coverage achieved and infection rates in served populations are all routinely reported by health systems.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Design search vs Coordination and institutions; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Design search vs Coordination and institutions; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c3cb37e-2a00-803b-9766-c484a72e1d85","name":"Improved Antibiotics Discovery","description":"Develop next-generation antibiotics and novel antimicrobial compounds using advanced discovery platforms to stay ahead of evolving pathogens."},{"id":"1c3cb37e-2a00-804f-90d5-fd6cefb1bd42","name":"Improved Broad-Spectrum Antiviral Discovery","description":"Develop new broad-spectrum antivirals that can be used to treat or prevent infection from evolving viruses."},{"id":"1c3cb37e-2a00-80fa-b511-e48564667c48","name":"Decentralized, Low-Resource Vaccine Production","description":"Establish scalable, decentralized vaccine production systems to rapidly deploy immunizations during outbreaks, reducing reliance on centralized facilities."},{"id":"1c7cb37e-2a00-801a-bec0-d048c8d745af","name":"Improved Modeled Intervention Strategies","description":"Advance public health by creating better models and strategies for intervention, integrating data-driven insights to inform policy and practice."},{"id":"1c3cb37e-2a00-80ee-a5be-fe6e7252ca2a","name":"Transmission Reduction Through Surfaces and Textiles","description":"Innovate new materials and coatings for surfaces and textiles that actively reduce pathogen viability and transmission."},{"id":"1c3cb37e-2a00-80f7-97b9-d80fafddabfd","name":"Better PPE","description":"Develop next-generation, affordable, high-performance personal protective equipment to reduce transmission risks."},{"id":"1c7cb37e-2a00-8056-a4df-f9f5c21235c5","name":"Enhanced Sanitation Solutions","description":"Innovate and scale sanitation technologies to improve water, sanitation, and hygiene, to reduce disease transmission."},{"id":"1c7cb37e-2a00-80f6-a55e-e559745cf6de","name":"Improved Vaccine Distribu","description":"Vaccine distribution is also a significant challenge in low-resource settings."}]},{"id":"1c1cb37e-2a00-80fb-9bbd-cbbe06f1776e","slug":"in-silico-molecular-simulation-is-slow-and-kludgy","name":"In-Silico Molecular Simulation Is Slow and Kludgy","description":"In-silico molecular simulation has not received the necessary push, despite the promise of machine learning-based surrogate models. Moreover, advancements in quantum chemistry—both AI accelerated and quantum/ASIC-enabled—remain underexploited.","field":"Chemistry","outcome":"Quantum-chemical accuracy at molecular-dynamics cost. Reaction energetics get screened computationally instead of measured one system at a time.","outcome_rationale":"Their description names surrogate models and AI-accelerated quantum chemistry as the underexploited route; the unlock is the accuracy-cost tradeoff collapsing.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Error against high-level reference calculations at a stated compute budget is the field's standard benchmark.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"Working now","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"Working now","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-8092-9680-ead8a5b3ce64","name":"AI-Enabled Molecular Dynamics","description":"Use neural network potentials and force fields to enhance molecular dynamics simulations, making them more efficient and accurate."},{"id":"1c1cb37e-2a00-802f-b0e3-f52a0e70218a","name":"AI-Accelerated Quantum Chemistry","description":"Leverage AI techniques to accelerate quantum chemistry calculations, improving the speed and accuracy of electronic structure predictions."},{"id":"1c1cb37e-2a00-80d2-a150-f1cae464aab5","name":"Quantum Computing-Enabled Quantum Chemistry","description":"Utilize hybrid quantum algorithms to perform quantum chemistry simulations, capitalizing on recent progress in quantum computing."},{"id":"1c1cb37e-2a00-80c8-bdee-d5550911fdf8","name":"ASIC-Enabled Quantum Chemistry","description":"Develop application-specific integrated circuits (ASICs) tailored for quantum chemistry calculations, aiming to combine efficiency with high computational power."},{"id":"1c1cb37e-2a00-80d5-9692-eae26b80e333","name":"High-Quality Experimental Chemistry Benchmarks","description":"Produce a high-quality experimental dataset to validate molecular simulation techniques and transition to frictionless reproducibility."},{"id":"1c1cb37e-2a00-802e-90c9-dbf91f2f6746","name":"Ab-Initio Calculation of Heavier Elements","description":"Create accurate, thoroughly benchmarked calculations for elements beyond the second row of the periodic table, for clusters of 3 or more atoms. Could complement 29.5."},{"id":"1c1cb37e-2a00-80bb-811b-c75611b77bba","name":"High-Quality Open Reaction and Structure Datasets","description":"Support open alternatives to Reaxys and the Cambridge Structural Dataset to the point where they are equal or superior in quality"},{"id":"1c2cb37e-2a00-8055-adeb-f33cf28d09e7","name":"Machine Learning Force Fields for Electrochemistry","description":"Extend work on ML force fields to charge transfer problems in external potentials, enabling in-silico discoveries in batteries, electrolysis, carbon capture, biochemistry and the origins of life"}]},{"id":"1c1cb37e-2a00-8026-a0fa-ff9d396d425a","slug":"limited-ability-to-identify-molecular-structures-through-spectroscopy","name":"Limited ability to identify molecular structures through spectroscopy","description":"Most molecular structure determination methods lose critical information, and solving the inverse problem remains challenging. This limits our ability to accurately reconstruct molecular structures from spectral data.","field":"Chemistry","outcome":"Read the structure straight off the spectrum. Crystallisation and purification drop out of the critical path for structure determination.","outcome_rationale":"Their description frames it as an inverse problem where information is lost; recovering it removes the slow experimental step.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Structure elucidation accuracy on held-out spectra with known answers is a direct, already-used benchmark.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"Working now","is_primary":0,"rationale":"Forward spectrum prediction from structure works now and supplies the training signal for the inverse direction.","confidence":"confident"},{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-80e7-8884-fc5e33b4f454","name":"Microwave Spectroscopy for 1:1 Mapping","description":"Employ microwave spectroscopy to establish a direct one-to-one mapping between molecular structure and spectrum, preserving detailed structural information."},{"id":"1c1cb37e-2a00-80e5-9a5d-d849f053228b","name":"Universal Spectroscopic Databases","description":"Develop comprehensive spectroscopic databases to train AI models for both forward and inverse predictions, enabling more accurate structure determination."}]},{"id":"1c1cb37e-2a00-804f-bf69-f232177933ce","slug":"understanding-life-as-a-far-from-equilibrium-physical-phenomenon","name":"Understanding Life as a Far-From-Equilibrium Physical Phenomenon","description":"Our ability to analyze organisms holistically as systems which emerge from fundamental physics is limited by our lack of formal frameworks for distinguishing living and nonliving systems which are precise enough to be useful for practical scientific problems","field":"Biophysics","outcome":"Origin-of-life research and biosignature detection get a testable target instead of a definitional argument, because there is an operational line between living and nonliving matter.","outcome_rationale":"Their description asks precisely for a framework precise enough to be practically useful, which is what biosignature work currently lacks.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Candidate measures exist and are computed on real molecules, assembly theory and free-energy dissipation among them, but the field actively disputes whether any of them settles what life is. A measurement can be taken; the verdict is what is contested.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"Speculative","is_primary":0,"rationale":"Formal framework construction is theory work, where AI contribution remains unproven.","confidence":"guess"},{"ai_type":"Prediction and modeling","maturity":"Speculative","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 Speculative / v2 2-5 years; mechanical rule had given 2-5 years). The gap asks for a formal framework precise enough to be useful, not for a faster calculation. Surrogates need a target function to approximate and this gap is the absence of one; there is no clear path from an emulator to a definition of life. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"Speculative","primary_rationale":"Maturity re-adjudicated on the merits (v1 Speculative / v2 2-5 years; mechanical rule had given 2-5 years). The gap asks for a formal framework precise enough to be useful, not for a faster calculation. Surrogates need a target function to approximate and this gap is the absence of one; there is no clear path from an emulator to a definition of life. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-8099-9364-c076edd8af7e","name":"Simulating Nonequilibrium Systems at Scale","description":"Develop software for simulating stochastic thermodynamics at scale with modern hardware accelerators (GPUs, etc), going beyond Gillespie algorithm."}]},{"id":"1c1cb37e-2a00-808d-9e6c-d7c767eee358","slug":"live-cell-imaging-at-deep-nanoscale-resolution-is-destructive","name":"Live Cell Imaging at Deep Nanoscale Resolution is Destructive","description":"Techniques that achieve deep nanoscale resolution in live cell imaging often destroy the sample, limiting the ability to conduct longitudinal studies on the same specimen.","field":"Biophysics","outcome":"Image the same living cell at nanoscale again and again. Subcellular dynamics get watched directly rather than reconstructed from fixed snapshots of different cells.","outcome_rationale":"Their description names the loss precisely: destruction forecloses longitudinal study of one specimen.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Resolution achieved against dose delivered against post-imaging viability is a three-way tradeoff that is routinely quantified.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"Speculative","is_primary":0,"rationale":"Quantum electron microscopy and ghost imaging require instruments that do not yet exist outside proof-of-principle.","confidence":"confident"},{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-804a-b3a1-c7d6e61ba942","name":"Quantum Non-Demolition X-Ray (Ghost) Imaging","description":"Explore quantum non-demolition x-ray imaging (or ghost imaging) methods that use quantum correlations to image samples with minimal perturbation, preserving sample integrity for longitudinal studies."},{"id":"1c1cb37e-2a00-80af-b662-db8edddc1ae0","name":"Quantum Electron Microscopy","description":"Develop a quantum electron microscope that leverages quantum principles to achieve high-resolution imaging while minimizing sample damage, enabling repeated measurements on the same specimen."}]},{"id":"1c1cb37e-2a00-8031-b79b-c13cca53b9c9","slug":"difficulty-delivering-physical-probes-for-imaging-into-living-cells","name":"Difficulty Delivering Physical Probes for Imaging into Living Cells","description":"Delivering physical probes for imaging into living cells is challenging due to barriers in cell membranes and potential perturbation of cellular function. New approaches are required to enable high-dimensional biosensing without invasive probes.","field":"Biophysics","outcome":"Watch biochemistry happen inside a living cell without disturbing the process you are watching.","outcome_rationale":"Their framing is that probes both fail to enter and perturb what they measure; removing both makes in-situ dynamics readable.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Species resolved per cell, perturbation magnitude and time to phototoxicity are standard reported instrument metrics.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Design search vs Measurement and sensing; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Design search vs Measurement and sensing; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-8077-8755-ce584ba364c3","name":"Label-Free High-Dimensional Biosensing","description":"Develop label-free biosensing techniques that leverage vibrational signatures to capture complex cellular information without the need for physical probe delivery."},{"id":"1c1cb37e-2a00-8043-9729-ef46e2c27e5a","name":"Metabolically Incorporated Labels","description":"Engineer cells to incorporate imaging labels metabolically, thereby eliminating the need for external probe delivery and enabling noninvasive imaging."}]},{"id":"1c1cb37e-2a00-8044-9b62-ef8af2ed2fa6","slug":"light-scattering-in-living-tissue-prevents-optical-access-to-deeper-regions","name":"Light Scattering in Living Tissue Prevents Optical Access to Deeper Regions","description":"Living tissue exhibits strong light scattering, which hampers deep-tissue imaging and limits resolution. Overcoming this barrier is critical for mapping neural activity and enabling noninvasive diagnostic imaging.","field":"Biophysics","outcome":"Optical imaging reaches depths that currently require cutting the tissue out, so neural activity and diagnostic imaging happen in intact organisms.","outcome_rationale":"Their description names deep-tissue neural mapping and noninvasive diagnostics as the two downstream unlocks.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Imaging depth in millimetres at a stated resolution and wavelength is the field's standard benchmark.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). The depth gains of the last decade came from three-photon excitation and adaptive optics hardware, not from algorithms. Learned descattering and computational wavefront correction are real but research-stage, and the ballistic-photon limit is physics. Type unchanged from the mechanical adjudication.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). The depth gains of the last decade came from three-photon excitation and adaptive optics hardware, not from algorithms. Learned descattering and computational wavefront correction are real but research-stage, and the ballistic-photon limit is physics. Type unchanged from the mechanical adjudication.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-808f-9e59-d5474f0360d2","name":"Magnetic-Based Imaging","description":"Develop magnetic imaging (magneto-genetics) approaches as an alternative modality that circumvents the limitations of optical scattering, potentially offering noninvasive imaging of deep tissue structures."},{"id":"1c1cb37e-2a00-8092-b82e-f7c67eeb8b2a","name":"Anti-Scattering Optical and Opto-Acoustic Methods","description":"Develop novel optical and opto-acoustic techniques that reduce scattering, enabling deep-tissue imaging without expensive MRI. These methods aim to improve resolution and enable whole-brain activity mapping as well as cost-effective “body scanners” for diagnostics."}]},{"id":"1c1cb37e-2a00-8099-b97a-ce85f02ab5a5","slug":"lack-of-direct-measurement-of-quantum-effects-in-biological-systems","name":"Lack of Direct Measurement of Quantum Effects in Biological Systems","description":"Despite theoretical predictions, quantum effects in biological systems remain largely unmeasured. Direct experimental evidence is needed to explore how quantum phenomena influence biomolecular interactions.","field":"Biophysics","outcome":"Whether quantum coherence does anything functional in biology gets settled experimentally rather than argued theoretically.","outcome_rationale":"Their description is explicit that predictions exist and direct evidence does not; the unlock is adjudication.","outcome_confidence":"confident","tier":"Verification contested","tier_rationale":"Candidate observables exist and have been measured, notably long-lived coherences in photosynthetic complexes, but the field disputes whether they are functional or artifacts of the spectroscopy, so measurement does not settle the question.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-8063-aa4b-fd082be08bff","name":"Quantum Biology Measurements","description":"Develop the instrumentation and controls required to directly measure quantum effects in biological systems."}]},{"id":"1c1cb37e-2a00-804b-8411-ee3977ce38e1","slug":"lack-of-structure-prediction-for-highly-dynamic-proteins","name":"Lack of Structure Prediction for Highly Dynamic Proteins","description":"Current structure prediction tools like AlphaFold excel for stable proteins but struggle with highly dynamic proteins whose structures fluctuate continuously, leaving a gap in our understanding of intrinsically disordered proteins and protein allostery.","field":"Biophysics","outcome":"Structure-based design extends to the disordered and allosteric proteins that AlphaFold-class tools currently cannot represent, which is a large share of the proteome.","outcome_rationale":"Their description names the boundary precisely: stable proteins solved, dynamic ones not.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Ensemble prediction is scorable against NMR and SAXS observables, and CASP-style blind assessment already exists for the static case.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":0,"rationale":"NMR ensemble mapping supplies the training and validation labels the prediction problem lacks.","confidence":"confident"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-8063-87bf-c59f48040713","name":"NMR Mapping and Modeling of Intrinsically Disordered Proteins","description":"Utilize nuclear magnetic resonance (NMR) techniques to capture the dynamic ensembles of intrinsically disordered proteins, enabling accurate modeling of their fluctuating structures."},{"id":"1c1cb37e-2a00-8081-bc05-e07c372ba0fe","name":"Protein Allostery Prediction Across the Proteome","description":"Develop methods to  measure and predict allosteric regulation mechanisms across the proteome, capturing dynamic conformational changes that impact protein function."}]},{"id":"1c1cb37e-2a00-80cd-aef2-fe6e9d360328","slug":"some-proteins-are-still-recalcitrant-to-experimental-structure-analysis","name":"Some Proteins Are Still Recalcitrant to Experimental Structure Analysis","description":"Membrane proteins are notoriously difficult to analyze experimentally and to incorporate into technological applications due to their inherent insolubility in aqueous environments. Their recalcitrance limits our capacity to study their structure and function in detail. Other challenges include the difficulty of studying small proteins with cryo-EM.","field":"Biophysics","outcome":"Membrane proteins are the biggest drug-target class and small proteins are the hardest cryo-EM case. Both turn structurally tractable.","outcome_rationale":"Their description names both sub-cases; membrane proteins are the majority of drug targets, which is where the downstream unlock sits.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Structures deposited and resolution achieved below 50 kDa are directly countable in the PDB.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":0,"rationale":"Orientation assignment for small particles in cryo-EM is a learned inference problem already being attacked.","confidence":"confident"},{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-802d-b9de-d07b4dfd9970","name":"Recode Membrane Proteins for Solubility","description":"Modify the genetic code or structure of membrane proteins so they become soluble in water, thereby enabling more robust experimental analysis and practical applications."},{"id":"1c1cb37e-2a00-8056-bab4-eec1398a6048","name":"Make it Easier to Map Orientations of Small Proteins in Cryo-Em","description":"Barcode the orientations of small proteins in the cryo-EM"}]},{"id":"1c1cb37e-2a00-8022-8d55-d645dc19dcf6","slug":"when-we-put-a-molecule-in-the-human-body-we-cant-predict-what-it-will-do","name":"When We Put a Molecule in the Human Body, We Can’t Predict What It Will Do","description":"Drug development is often hampered by failures related to absorption, distribution, metabolism, excretion, and toxicity (ADME/Tox). Improved predictive models for molecular interactions are essential for designing safer, more effective drugs, as well as evaluating the impact of environmental chemicals.\n\nAdditionally, there is a significant gap in our knowledge of what exactly is present in foods and how these components affect human biology. A comprehensive mapping of the “foodome” and studies on food component functionality are needed to advance nutrition science and personalized dietary interventions.","field":"Physiology and Medicine","outcome":"ADME and toxicity failures show up before dosing rather than in clinical trials.","outcome_rationale":"Their description names ADME/Tox failures as the dominant cause of drug development failure.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Prediction accuracy against measured pharmacokinetics and observed clinical toxicity is directly scorable on held-out compounds.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c8cb37e-2a00-806e-a3b2-d656247d1f19","name":"Pharm-ome Mapping","description":"Systematically map drug–target interactions (the “pharm-ome”) to better predict drug efficacy, side effects, and repurposing"},{"id":"1c8cb37e-2a00-8090-a5e4-cebc0eb7f707","name":"Predictive Drug Safety and Efficacy Models","description":"Develop predictive models for ADME-Tox to lower drug candidate failure rates and increase clinical safety and efficacy."},{"id":"1c8cb37e-2a00-8068-a40a-fc08a69a477f","name":"Microplastics Characterization and Degradation","description":"Characterize microplastics in food and water and understand human exposure and impacts. Develop new technologies to degrade PET, polystyrene, and microplastics (e.g., microbe/ enzyme systems)."},{"id":"1c8cb37e-2a00-80ef-825d-f46ba209a0a3","name":"Toxin Mapping and Neutralization","description":"Adapt pharm-ome mapping approaches to environmental toxins to predict their biological impacts and improve safety assessments. \n\nFor example, scalable solutions for rapid detection and neutralization of mycotoxins in food systems, which contaminate 25% of agricultural products and post significant health risks (e.g., portable sensors and enzyme/ microbial/ RNA-based detoxification systems)."},{"id":"1c8cb37e-2a00-8008-933d-fe972adef17b","name":"Scalable Tool Compounds","description":"Develop large libraries of tool compounds to systematically probe molecular interactions, aiding both drug discovery and toxicity prediction."},{"id":"1c8cb37e-2a00-806c-954e-c64cc93cc7e5","name":"Immunogenicity Prediction for Biologics","description":"Create computational models to predict and mitigate immune responses to biologic drugs, improving safety profiles."},{"id":"1c8cb37e-2a00-804c-b1a9-c0333ba552c4","name":"Immunodominance Prediction for New Pathogens","description":"Develop models to forecast which epitopes will dominate immune responses upon exposure to new antigens, guiding vaccine and therapeutic development."},{"id":"1c8cb37e-2a00-808d-8a72-d5b96ca43396","name":"Surveillance of Microbes and Fungi","description":"Surveillance networks of genetic mutations in bacteria and fungi–both foodborne pathogens and antimicrobial resistance trends."},{"id":"1c8cb37e-2a00-801f-8653-dd1242e1c317","name":"Comprehensive Foodome Mapping","description":"Systematically catalog the chemical and biological components of foods and study their interactions with the human body at multiple scales—from receptors to whole organisms."},{"id":"1c8cb37e-2a00-80f3-a427-d4128d753435","name":"Functional Food Component Analysis","description":"Measure the biological effects of food components at the receptor, cellular, organ, and organism levels to understand their impact on human health."}]},{"id":"1c1cb37e-2a00-80ac-842e-e5f6ea4906cb","slug":"our-measurements-and-tests-arent-revealing-what-is-actually-causing-many-diseases","name":"Our Measurements and Tests Aren’t Revealing What Is Actually Causing Many Diseases","description":"Our understanding of human physiology and disease remains incomplete. In the last century, we have developed cures for many diseases with well-defined root causes (polio, smallbox, cholera, SMA, cervical cancer, etc.). However, a wide array of conditions still eludes cures and treatments. We have yet to fully decipher the dynamic interplay between brain and peripheral systems, the bioenergetic processes underlying chronic conditions, and the multifactorial pathways that drive aging. The biological mechanisms driving complex diseases and the aging process are multifactorial, involving multiple interacting pathways. \n\nAlthough we understand some individual aging mechanisms, we do not yet have line of sight to comprehensively rejuvenating mammals or extending lifespan. To overcome these challenges, we need combinatorial approaches that can modulate multiple mechanisms simultaneously, allowing us to measure multi-system impacts and develop effective interventions.","field":"Physiology and Medicine","outcome":"Multifactorial disease and ageing get identified causal mechanisms rather than correlates, which makes them targetable the way single-cause diseases already are.","outcome_rationale":"Their description contrasts diseases with well-defined root causes, which were cured, against multifactorial ones, which were not.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Biomarkers and intervention effects are measurable, but the underlying causal structure of a multifactorial disease is inferred from those proxies rather than observed.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":0,"rationale":"Better systems measurements and resolved brain-body mapping are detection capabilities their list names directly.","confidence":"confident"},{"ai_type":"Running experiments","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Running experiments","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-80f3-a615-de70d7aa1e22","name":"Epidemiological Tools to Identify the Root Causes of Disease","description":"Integrated analytic tools to study clinical, molecular, and environmental datasets to identify patterns and infer the underlying causes of disease."},{"id":"1c8cb37e-2a00-8075-b9fd-e4e70d8882d8","name":"Chemical and Cell-Type Resolved Mapping of Brain Activation","description":"Create detailed, functional maps of hypothalamus/brainstem activation by specific peptides and hormones to link neural activity with physiological outcomes and decipher how the brain orchestrates complex physiological responses."},{"id":"1c8cb37e-2a00-80a9-9338-e41e0b54230e","name":"Whole Body Connectomes","description":"Map the connectivity not only within the brain but also between the brain and peripheral organs to reveal integrated regulatory networks."},{"id":"1c8cb37e-2a00-8073-8eb1-eeb020a12336","name":"Engineering Endosymbionts","description":"Study and engineer endosymbiotic relationships (e.g., mitochondria) to better understand and manipulate cellular energy production, potentially offering new avenues to treat bioenergetic disorders."},{"id":"1c8cb37e-2a00-800f-9695-c0558dc96e19","name":"Combinatorial Aging Interventions Screening","description":"Develop aging-relevant in vitro models and screen combinations of interventions (e.g., small molecules, gene therapies) using multi-omic and functional readouts to identify synergistic treatments that extend lifespan or promote regeneration."},{"id":"1c8cb37e-2a00-807b-b5e3-ddbd1ab17bec","name":"In-Vivo/ In-Situ Pooled Screening","description":"Implement pooled screening techniques directly in living organisms to test multiple intervention combinations concurrently in aged context, accelerating discovery in complex disease and aging research."},{"id":"1c8cb37e-2a00-8031-adbc-d93ca42a77e3","name":"Developing Immortal Biological Models","description":"Develop immortal model organisms beyond cell lines to enable the study of longevity and underlying mechanisms of aging. "},{"id":"1c8cb37e-2a00-80ca-94d3-e1c71507787b","name":"Long-Term Multimorbidity Endpoints","description":"Human long term multimorbidity endpoint  trials of known to be safe compounds"},{"id":"1c8cb37e-2a00-8091-aac6-ee716164c482","name":"Better Systems Measurements","description":"To study interactions in complex systems we need to measure multiple agents simultaneously, ideally with timecourse data. Doing this in live aged organisms would require new tools."},{"id":"1c8cb37e-2a00-801d-af58-da384dc30288","name":"Understudied Biological Systems","description":"There are many other biological systems that are understudied. We need field building to drive greater study of important biological dark matter, e.g., extracellular matrix biology, pregnancy, thymic involution, chronic infections driving chronic disease, menopause biology, and many others"}]},{"id":"1c1cb37e-2a00-80fd-87e6-c92b83132c48","slug":"limited-longitudinal-data-in-humans","name":"Limited Longitudinal Data in Humans","description":"A comprehensive understanding of human health over time is hindered by the lack of longitudinal data from cohorts that are diverse and globally representative. Such datasets are essential to track developmental, nutritional, and environmental influences on long-term health outcomes.","field":"Physiology and Medicine","outcome":"Follow health trajectories across decades and diverse populations. Developmental and environmental causes stop being inferred from cross-sections.","outcome_rationale":"Their description names developmental, nutritional and environmental influences on long-term outcomes as the target.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Cohort size, years of follow-up, retention and demographic representativeness are direct and reported.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":0,"rationale":"Easy longitudinal sampling depends on low-burden assays, which is a working detection capability.","confidence":"confident"},{"ai_type":"Coordination and institutions","maturity":"2-5 years","is_primary":1,"rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). The mechanism is proven — UK Biobank and All of Us are coordination capability applied and running. What the gap asks for is diverse and globally representative cohorts, and standing those up in low- and middle-income settings is not possible at that scale today. Type unchanged from the mechanical adjudication.","confidence":"confident"}],"primary_ai_type":"Coordination and institutions","primary_maturity":"2-5 years","primary_rationale":"Maturity re-adjudicated on the merits (v1 2-5 years / v2 Working now; mechanical rule had given Working now). The mechanism is proven — UK Biobank and All of Us are coordination capability applied and running. What the gap asks for is diverse and globally representative cohorts, and standing those up in low- and middle-income settings is not possible at that scale today. Type unchanged from the mechanical adjudication.","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c8cb37e-2a00-8071-a638-cbb3eccff38c","name":"Accessing Human Cohorts","description":"Develop platforms (e.g., multi-channel data collection systems) to continuously monitor health parameters in diverse human cohorts, e.g., existing or planned clinical trials. "},{"id":"1c1cb37e-2a00-8018-9b59-e5623edbdd0a","name":"Mapping Developmental Malnutrition","description":"Systematically profile cellular mechanisms by which early-life malnutrition affects physiology across the lifespan, providing insights into long-term developmental impacts."},{"id":"1c1cb37e-2a00-80d5-a586-df58692bd270","name":"Breast Milk-ome Profiling","description":"Conduct comprehensive “-omics” studies of breast milk to capture its molecular and microbial composition and its role in infant development."},{"id":"1c8cb37e-2a00-8006-b1bf-dc940ba4bc84","name":"Easy Longitudinal Sampling","description":"Implement noninvasive, low-cost sampling methods (e.g., breath analysis and point-of-care nucleic acid sequencing) to collect repeated, high-resolution physiological data. This would augment the depth of data from initiatives such as the NIH All of Us research program."},{"id":"1c8cb37e-2a00-80dd-a39a-e71bc705d42a","name":"Expanded Biobank Initiatives","description":"Broaden the scope and diversity of existing biobanks (e.g., the UK BioBank) to include more global populations and additional longitudinal health measures."}]},{"id":"1c1cb37e-2a00-8086-9443-f719d9caef36","slug":"we-cant-safely-and-controllably-deliver-complex-molecular-payloads-to-the-targets-we-want-in-the-body","name":"We Can’t Safely and Controllably Deliver Complex Molecular Payloads to the Targets We Want in the Body","description":"Current in-vivo delivery systems (viral vectors, nanoparticles, microchips) face challenges such as off-target accumulation and inefficiency, particularly in delivering therapies to the brain. Novel delivery approaches are needed to improve targeting and performance.","field":"Physiology and Medicine","outcome":"Payloads reach the tissue they are aimed at, the brain included, so what can be delivered stops limiting what can be treated.","outcome_rationale":"Their description names off-target accumulation and brain delivery as the specific failures.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Percent of injected dose reaching target tissue versus liver and other off-target organs is a direct, standard biodistribution readout.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Design search","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-8062-9340-fb9a8e8324ee","name":"Map of Cell-Type or Tissue Specific Targets","description":"Systematically investigate cell-type and tissue specific targets equivalent to ASGPR for the liver. "},{"id":"1c1cb37e-2a00-8059-a445-fd11c443fd95","name":"Delivery Systems that “De-Target” the Liver and Macrophages","description":"New delivery mechanisms are needed to target the desired tissues and avoid uptake by hepatocytes and macrophages. Off-target uptake necessitates higher dosing, increasing toxicity risks."},{"id":"1c8cb37e-2a00-805d-ac85-d2e3598dda45","name":"Microelectronics for Tissue-Targeted Delivery","description":"Utilize ultra-small microelectronic devices delivered via endovascular routes to target specific tissues (including the brain), bypassing the limitations of viral vectors and bulky implants."},{"id":"1c8cb37e-2a00-805d-a86f-dccdda6d37a3","name":"Improved & New Gene Therapy Delivery Vehicles","description":"Develop improved gene therapy vectors that can be manufactured at scale, as well as completely novel therapeutic gene therapy delivery vehicles. "},{"id":"1c8cb37e-2a00-801f-87d2-fe4cc430af32","name":"Map the Determinants of Cellular Migration in the Body","description":"Map the determinants of cellular migration in the body to block metastasis and control cellular delivery."}]},{"id":"1c1cb37e-2a00-80fd-a046-f81466f201e7","slug":"inadequate-models-of-human-physiology","name":"Inadequate Models of Human Physiology","description":"Current preclinical models of human physiology, including animals and organoids, do not fully capture the complexity of human physiology, limiting the predicting power of preclinical experiments and explaining, in part, the costly failures of drug development in clinical trials. This is especially true for complex disorders including those of aging, neurological disorders, and female reproductive biology. More systematic and representative models—including ex vivo human organ systems or even whole bodies and novel animal species—are needed to improve the predictive power of biomedical research. These technologies also have applications in addressing organ shortages, improving neonatal care, and other unmet medical needs.","field":"Physiology and Medicine","outcome":"Drugs fail cheaply on the bench instead of expensively in late-stage trials, because preclinical results start predicting human outcomes.","outcome_rationale":"Their description ties model inadequacy directly to costly drug-development failures in clinical trials.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"The share of preclinical successes that survive to clinical success is a direct, historically tracked translation rate.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Design search vs Prediction and modeling; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Design search vs Prediction and modeling; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1c1cb37e-2a00-80b2-ac6c-e4151b58e835","name":"3D Organoid Technology","description":"Lab-grown 3D organ tissues have become an established additional model system to recapitulate aspects of human biology. This technology could also enable the development of functional organs for transplant.\n\nIt is also important to make tissue models that recapitulate the effects of aging, or a form of “accelerated lifetime testing”."},{"id":"1c1cb37e-2a00-8058-9990-c236f77a985c","name":"Grow and Maintain Human Organs Ex Vivo or in Animals","description":"Grow human organs in animals to study disease and drug response more accurately. Use advanced stem cell technologies to grow patient-specific tissues and organs in animals for transplantation. \n\nHuman organs could also be maintained ex vivo (“in a vat”) for research purposes. Perfused organ systems (including cadaver-based models) that maintain the structure and function of human tissues ex vivo would also be enabling."},{"id":"1c1cb37e-2a00-800a-a7de-f1e61729b3ef","name":"Improved and Diverse Research Models","description":"A greater variety of small animal models (along with corresponding suites of tools such as species-specific antibodies, annotated genomes, transgenics, etc.) would enable novel biological insights and could be used to develop models of complex human diseases. Additional rodent models, as well as those beyond mouse and rat would be highly enabling.\n\nMore realistic models are also critical for aging research–many diseases of aging are studied in young animals.\n\nAnalytical tools are also important to make it easier for researchers to understand a) limitations of their research models, b) be aware of superior but less commonly used models. For example, a “Maniatis” style handbook detailing which human pathophysiology is mirrored in different species."},{"id":"1c1cb37e-2a00-80a9-b398-e93dcf51775e","name":"More Information Gleaned from Single Animals","description":"High-throughput testing in a single animal would enable entire studies to be run in rare/exceptional animals, e.g. with spontaneous disease mimicking humans or species not suitable for research labs."},{"id":"1c1cb37e-2a00-80b2-b6aa-dfc3b8210092","name":"Ectogenesis (Artificial Wombs)","description":"Artificial wombs could revolutionize neonatal care and reduce preterm birth complications. They are an early stage research area with various positive biomedical externalities. "},{"id":"1c1cb37e-2a00-8003-8656-e203281a11a5","name":"Cryopreservation and Rewarming to Extend the Viability of Biological Tissues","description":"Develop engineering methods for cryopreserving and safely rewarming large organs or whole bodies, which could revolutionize transplantation and long-term tissue storage."},{"id":"1c1cb37e-2a00-8093-9040-d4f7948c446e","name":"Living Human Bodies Without Brains","description":"Living human bodies created from stem cells without neural components could be transformative for medical research and drug development. There are many open questions–for example, the long time it takes for maturation, whether a body would function without neural components, etc."}]},{"id":"1b4cb37e-2a00-80af-8e6b-ff088bc25396","slug":"clinical-trials-are-inefficient-slow-and-scarce","name":"Clinical Trials are Inefficient, Slow and Scarce","description":"Clinical trial designs are often inefficient, resulting in high costs, lengthy timelines, and suboptimal patient outcomes. Innovative trial designs and decision-support tools are required to streamline the clinical evaluation process and accelerate therapeutic development.","field":"Physiology and Medicine","outcome":"More therapies evaluated per year, without a matching rise in cost or patient burden. Promising candidates stop dying in the queue.","outcome_rationale":"Their description names cost, timeline and scarcity together, and their capability list is dominated by trial-design and recruitment tools.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Cost per trial, enrolment time and trials completed per year are recorded in registries.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":0,"rationale":"Digital twins, synthetic control arms and proxy biomarkers are prediction problems, and all appear in their capability list.","confidence":"confident"},{"ai_type":"Design search","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Design search","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1b4cb37e-2a00-8096-a2ca-c60028df9030","name":"Real-Time Dynamics of Patient Molecular Biology","description":"The ability to study the real-time dynamics of molecular pathways in living humans remains extremely limited in time-resolution (frequency and duration) and scalability (beyond lab testing). Continuous, minimally invasive technologies are needed to improve our understanding of human physiology, accurately diagnose patients, and assess the impact of clinical interventions. Molecular engineering, particularly protein engineering, and integrated circuit design can be combined to develop new, miniaturized devices."},{"id":"1c1cb37e-2a00-805e-956b-dfa2baef87f1","name":"Improved Diagnostics","description":"Diagnostic tools for the top 10 hard to diagnose diseases could have a significant impact on patients and the healthcare ecosystem. Diagnostics as an industry is in a state of market failure. Many disorders remain challenging to diagnose, impacting patient outcomes and burdening the healthcare system.\n\nMonitoring the epigenome can enable the identification of exposures to infectious disease and reveal exposure to threat agents."},{"id":"1c1cb37e-2a00-8070-b295-e2c1cdfc466f","name":"Advanced Noninvasive Measurements of Patient Biology","description":"Noninvasive monitoring technologies can enable high-resolution, point-of-care data collection."},{"id":"1c1cb37e-2a00-80a5-a157-cca91fc429b4","name":"Faster Proxy Biomarkers","description":"Identify and validate robust biomarkers and surrogate endpoints to serve as effective proxies in clinical trials, enabling faster and more informative evaluations. This is especially important for aging (e.g., as proposed by the Norn Group)."},{"id":"1c1cb37e-2a00-8023-9cd1-c986c8de5af8","name":"Digital Twins / Synthetic Clinical Trial Models","description":"Create high-fidelity computational models (digital twins) that accurately simulate human physiology, enabling synthetic clinical trials and faster hypothesis testing."},{"id":"1c1cb37e-2a00-80c0-8b4e-d1ccc14b4367","name":"New Clinical Trial Designs and Tools","description":"Clinical trial designs such as challenge trials, adaptive trial design, and group testing can improve efficiency. \n\nThese can be complemented with analytics and decision-support systems to optimize design, patient enrollment, and outcome interpretation."},{"id":"1c1cb37e-2a00-80be-a815-de00ff02359b","name":"Enhanced Participant Recruitment","description":"Recruitment is a bottleneck that could be addressed by socio-technical programs. Increased interoperability and unity of data access and patient recruitment across centers and disease states would help de-silo recruitment and improve efficiency."},{"id":"1c1cb37e-2a00-806a-87b8-cfa0f619cd67","name":"Prioritization of Clinical Trials and Practice Updates","description":"Mine biomedical knowledge to decide which neglected assets to run trials on for which conditions"},{"id":"1c1cb37e-2a00-80bb-86e9-f2464ccc2edd","name":"Multi-Disease Efficacy Testing with Standardized Biomarkers","description":"Infrastructure and coalition to collect extra blood samples during trials and apply biomarkers for different diseases to get more value from each trial."},{"id":"1f0cb37e-2a00-80db-940f-de96a93d37f6","name":"Infrastructure for Decentralized Patient-Reported Clinical Studies","description":"Decentralized patient reporting for clinical studies would amplify and diversify patient participation at decreased cost by reducing logistical and geographic barriers. It could compound and incorporate subjective patient experiences at large scale and would improve public accessibility of trial data."}]},{"id":"1b4cb37e-2a00-80b8-a4f8-ee697b40331b","slug":"cellular-and-biomolecular-states-are-highly-multimodal-and-complex","name":"Cellular and Biomolecular States Are\n  Highly Multimodal and Complex","description":"Cellular state is a multifaceted and complex phenomenon, involving multiple overlapping omics layers that vary in time and space. Capturing and representing this multimodal complexity is essential for predictive modeling of cell behavior and for advancing our understanding of cellular function. ","field":"Cellular and Molecular Biology","outcome":"One model of cell state that predicts the response to a perturbation it has never seen, replacing a stack of per-assay descriptions.","outcome_rationale":"Their description names predictive modelling of cell behaviour as the goal; the unlock is generalisation to unseen perturbations.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Held-out perturbation prediction accuracy is a direct benchmark and is already the field's standard evaluation.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Reading and synthesis","maturity":"2-5 years","is_primary":0,"rationale":"Their capability list includes automated mechanistic interpretability of those models, which is a reasoning task over learned representations.","confidence":"confident"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (maturity: Working now vs 2-5 years). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1b4cb37e-2a00-808f-8016-ed0e34bb66e3","name":"Universal Latent Variable Model of\n  Cellular State","description":"Train predictive models that can probabilistically infer any missing\n  omics data from any available measurements by learning a broadly useful\n  latent space. Such a universal representation of cellular state could become\n  the foundation for many predictive and diagnostic applications."},{"id":"1b4cb37e-2a00-80b0-8054-d56e1fa698d0","name":"Automated Mechanistic Interpretability of Generative AI Models Trained on Biological Data","description":"Develop automated approaches for mechanistic interpretability of virtual\n  neural network models trained on biological state data. This would enable the\n  extraction of mechanistic insights from predictive models, potentially\n  informing both basic science and therapeutic design."},{"id":"1b4cb37e-2a00-800a-a99a-c6b3bf8a960a","name":"Models for Dynamic Cellular and Subcellular Data","description":"Models that represent cells as self-organizing and adaptive are critical for enabling the simulation of multi-cellular systems. \r\n\r\nThese should incorporate subcellular data from experimental techniques tracking individual molecules within cells in real time, as well as capture the fundamental interaction between external and internal molecular environment."}]},{"id":"1b4cb37e-2a00-80e6-8b25-fd9e770fbb3f","slug":"limited-ability-to-image-molecules-in-their-native-contexts","name":"Limited Ability to Image Molecules in Their Native Contexts","description":"We are currently limited in our ability to image molecules in their native contexts—for example, within live 3D tissues. Achieving scalable, high-resolution imaging of biomolecules in situ would de-risk many areas of biomedical science by enabling integrative, comprehensive molecular mapping within intact specimens.","field":"Cellular and Molecular Biology","outcome":"Molecular identity and position get read inside intact three-dimensional tissue, so spatial maps describe real organs rather than dissociated cells.","outcome_rationale":"Their description names the loss caused by dissociation and the value of measuring within intact specimens.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Species multiplexed, spatial resolution and imaging depth are direct instrument metrics with published benchmarks.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"Working now","is_primary":0,"rationale":"Their listed 'binders for every epitope' capability is de novo binder design, which works now at increasing hit rates.","confidence":"confident"},{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1b4cb37e-2a00-803c-9ae5-f521d7801fbc","name":"Freezing and Unfreezing Time for Living\n  Cells","description":"Develop methods to “freeze” and subsequently “unfreeze” living cells,\n  effectively capturing dynamic states for later analysis and then resuming\n  cellular function."},{"id":"1b4cb37e-2a00-80be-946e-c837881d5ef5","name":"Molecular 3D Scanning","description":"Create whole-cell, nanometer-resolved, in-situ multi-omic imaging\n  approaches that can map hundreds to thousands of molecules (e.g., proteins,\n  RNAs) within intact specimens at nanoscale resolution. The core chemistries\n  and imaging technologies for this exist, but they need to be integrated and\n  brought to scale. We could, for example, comprehensively measure the many\n  aspects of the “Hallmarks of Aging” within a single tissue sample."},{"id":"1b4cb37e-2a00-80ca-b324-cdcdac9a2f59","name":"Live Cell Subcellular Imaging &\n  Foundation Models","description":"Develop live cell subcellular imaging techniques in tissues combined with\n  computational foundation models to interpret the resulting data, enabling\n  detailed mapping of subcellular structures and their dynamics  in native environments."},{"id":"1b4cb37e-2a00-80e3-a775-ffb918e3064c","name":"Binders for Every Epitope","description":"Highly specific mAb/nanobody type binders for every target epitope,\n  including for specific post-translational modifications"},{"id":"1f0cb37e-2a00-8038-85f1-c83d88b15505","name":"Multiplexed Live-Cell Dynamics Recordings","description":"Develop techniques for spatial multiplexing of dynamic signals in live cells, capturing real-time changes and molecular ticker-tapes that record cellular events over time."}]},{"id":"1b4cb37e-2a00-8060-a0b6-cc4e1c0e782b","slug":"fundamental-biomolecular-actors-in-cells-remain-largely-invisible","name":"Fundamental Biomolecular Actors in Cells Remain Largely Invisible","description":"Many of the fundamental actors in cells—proteins, lipids, and metabolites—are still mostly invisible to us, especially when considering their extensive multiplexity, diversity, cell-to-cell heterogeneity, and temporal variation. Without scalable, cost-effective technologies to capture these molecular details, our comprehensive analysis of complex biological systems remains limited.","field":"Cellular and Molecular Biology","outcome":"Transcriptomics has deep coverage; the molecules that do the work do not. Proteins, lipids and metabolites get measured per cell and at scale.","outcome_rationale":"Their description names the asymmetry directly: the fundamental actors remain invisible while RNA is routinely profiled.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Molecules quantified per cell, cells processed per run and cost per cell are the field's routine reported figures.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Design search","maturity":"Working now","is_primary":0,"rationale":"Affinity reagent design for broader molecular coverage is a working protein-design capability.","confidence":"confident"},{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1b4cb37e-2a00-8048-9641-e10b875bb35b","name":"Metabolomics","description":"Establish foundational datasets and machine learning models for inverse\n  mass spectrometry of small molecules, enabling interpretation of metabolomic\n  data at the single-cell level."},{"id":"1b4cb37e-2a00-805b-bbfb-e4702bbbb27c","name":"High-Throughput Single-Cell Proteomics","description":"Develop new technologies to improve the cost-performance of single-cell\n  proteomics (>100x), enabling proteome-wide analysis at scale. This would\n  allow proteins—the functional output of genetics—to be analyzed\n  comprehensively across complex biological systems. Some of these technologies\n  are single-molecule, some are not. "},{"id":"1b4cb37e-2a00-809f-99e9-f255d62e2815","name":"Mapping Other Key Biomolecules","description":"Develop technology suites for single-cell glycomics and lipidomics using\n  methods such as in-situ multi-cycle imaging or spatially resolved\n  nanopore/mass spectrometry. This would scale up the mapping of key molecules\n  beyond proteins and metabolites."},{"id":"1b4cb37e-2a00-80f5-aa25-e712c235768f","name":"Deciphering the DNA Regulatory Code","description":"Decipher the “DNA regulatory code” that governs gene expression.\n  Understanding it would enable the prediction of how perturbations to cell\n  state affect transcription in development and disease."}]},{"id":"1b4cb37e-2a00-80b5-a4f1-f702c75399f8","slug":"our-immune-system-can-uniquely-recognize-nearly-any-molecule-but-we-dont-know-the-recognition-code","name":"Our Immune System Can Uniquely Recognize Nearly Any Molecule but We Don’t Know the Recognition Code","description":"A better understanding of how the immune system interacts at the molecular level with threats and triggers is critical. This knowledge would enable the development of predictive tools and technologies to augment immune responses—improving interventions against infections, cancers, and autoimmune disorders.","field":"Immunology","outcome":"Predict which receptor binds which antigen, and immune responses get designed rather than found by screening.","outcome_rationale":"Their description names predictive tools to augment immune responses across infection, cancer and autoimmunity as the unlock.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Binding prediction accuracy on held-out receptor-antigen pairs is directly scorable against experimental assays.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1b4cb37e-2a00-8038-a083-c4bda0f7538b","name":"T-cell and B-cell Receptor Antigen Mapping","description":"Construct a comprehensive map linking trillions of T-cell receptors (TCRs) and B-cell receptors (BCRs) to millions of disease antigens. This map will facilitate antibody–antigen binding prediction, allowing the identification of target antigens based solely on immune cell DNA sequences."}]},{"id":"1b4cb37e-2a00-80d3-b382-c445e2718659","slug":"our-immune-memory-contains-a-detailed-history-of-exposures-but-we-cant-read-it","name":"Our Immune Memory Contains a Detailed History of Exposures but We Can’t Read It","description":"Immunological diseases often have nonobvious, complex etiologies and pathophysiologies that are difficult to identify.","field":"Immunology","outcome":"The body already keeps a record of every exposure. Reading it turns immune repertoire into a diagnostic for diseases with obscure causes.","outcome_rationale":"Their gap title states it exactly: the history is there and unreadable, and their capability list targets links such as infection to neurodegeneration.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Predicted exposure history can be scored against known infection records in existing longitudinal cohorts.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":0,"rationale":"Repertoire sequencing and multi-omic profiling are working assays; the interpretation layer is what is missing.","confidence":"confident"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1b4cb37e-2a00-800c-b9d2-df04a8402a16","name":"Large-Scale Profiling of Infection to Expose Links to Neurodegeneration","description":"Conduct a large-scale, comprehensive study to establish causal links between persistent microbial/viral infections (such as herpes simplex 1) and neurodegenerative diseases like Alzheimer’s Disease. Such a study would illuminate the role of infections in increasing disease risk and progression.\n\nThis effort could serve as a sequel to the Genome-Wide Association Studies (GWAS) that have been performed since the completion of the Human Genome Project. Many diseases failed to show obvious genetic etiologies from those GWAS efforts, suggesting a role for the environment in disease causation. Large-scale infection profiling could therefore unearth etiologies that were not possible to detect by GWAS."},{"id":"1b4cb37e-2a00-8013-ab45-d8eab2f7a3fb","name":"Repertoire- and Cellular Subset-Level Adaptive Immune Tools","description":"Adapt and build experimental and computational methods to read out distributed and sparse immune memory signatures from adaptive immune cells. Beyond individual antibody-antigen binding, this approach focuses on signatures distributed across cellular subpopulations and the repertoire. Decoding these signatures could identify hidden causes of and cures for disease, enabling more accurate diagnosis, treatment, and prevention of chronic conditions."},{"id":"1b4cb37e-2a00-8034-ba73-cefc0f554570","name":"Longitudinal Multi-Omic Immune Profiling\n  Across Populations","description":"Generate longitudinal, multi-omic immunological data from a diverse\n  cohort of individuals. This dataset would be critical for enabling immune-ome\n  modeling and prediction (for example, in forecasting vaccine responses)."},{"id":"1b4cb37e-2a00-8058-8429-cfd3c79b7f53","name":"Universal Immune-Computer Interface (ICI)","description":"Develop a universal immune-computer interface to enhance the immune\n  system’s targeting of pathogens and cancers, while reducing issues such as\n  autoimmunity and transplant rejection. The ICI would involve two-way\n  coupling—where the immune system and computer mutually optimize their\n  matching processes—and real-time feedback loops. An example could be a\n  wearable device that integrates mRNA manufacturing with single-cell\n  sequencing."},{"id":"1b4cb37e-2a00-80f3-9824-e12d0bfe0f19","name":"Immune Tolerance Induction to Enable In-Human Synthetic Biology with Foreign Proteins","description":"We need a way to programmably induce immune tolerance to a user-defined foreign protein, in order to enable many new forms of gene and cell therapy (not to mention help with autoimmune diseases). This is especially needed for brain computer interface as many of the most powerful concepts for BCI would involve adding foreign protein such as transducer proteins for optical or acoustic signals. As Hannu Rajaniemi wrote, “a flexible ability to induce immune tolerance to opsins is a prerequisite of a two-way mind meld with computers”."}]},{"id":"1b4cb37e-2a00-80da-94b7-d6c79e96cf17","slug":"we-cant-take-high-resolution-movies-of-or-intervene-in-brain-computation-at-the-single-neuron-level","name":"We Can’t Take High-Resolution Movies of or Intervene in Brain Computation at the Single Neuron Level","description":"Capturing the dynamics of large brain networks at single-neuron resolution in vivo is extremely challenging. Advanced imaging methods that record fast, high-resolution activity without destructive intervention are required to unravel the complex interplay of neuronal circuits in real time.","field":"Neuroscience","outcome":"Watch large neural populations at single-neuron resolution and behavioural timescales in a living brain, instead of reconstructing circuits from averages.","outcome_rationale":"Their description names fast, high-resolution, non-destructive recording as the requirement for unravelling circuit interplay in real time.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Neurons recorded simultaneously, volume rate and imaging depth are direct instrument metrics with published records.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":1,"rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Working now","primary_rationale":"Independent passes disagreed (maturity: 2-5 years vs Working now). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1b4cb37e-2a00-8043-a6ad-e037b8aad421","name":"Novel Fast-Scanning Microscopy of Brain Cortex","description":"Develop innovative microscopy techniques that enable rapid, high-resolution imaging of neuronal networks in vivo at single–neuron resolution (e.g., Light Beads Microscopy)"},{"id":"1b4cb37e-2a00-806f-8585-e74b56bf89f4","name":"In-Vivo Connectomics","description":"Create methods to map neuronal connectivity in living brains, capturing the dynamic interactions between neurons at a single-cell level."},{"id":"1b4cb37e-2a00-8082-a289-e7c4d1fe7f6e","name":"Low Energy Multiphoton Brain Imaging","description":"Develop a novel physical pathway for multiphoton fluorescence generation that enables optical sectioning and excitation at red wavelengths with simple systems like continuous wave lasers."},{"id":"1b4cb37e-2a00-808e-8f5e-f2f13b51343d","name":"In-Vivo Optical Transparency of Brain Tissue","description":"Develop techniques that render brain tissue optically transparent in vivo, allowing deeper and higher-resolution imaging of neural networks without invasive sectioning."}]},{"id":"1b4cb37e-2a00-807c-88c6-d543174bdbd2","slug":"current-model-systems-for-brain-function-are-not-representative-of-the-real-human-brain","name":"Current “Model Systems” for Brain Function are Not Representative of the Real Human Brain","description":"Current in vivo and in vitro models often fail to capture human brain function. Innovative model systems—including digital reconstructions, embodied simulations, and new biological models—are needed.","field":"Neuroscience","outcome":"Results from model systems survive the jump to the human brain.","outcome_rationale":"Their description frames the loss as models failing to capture human brain function, which is a transfer failure.","outcome_confidence":"confident","tier":"Proxy only","tier_rationale":"Correspondence on specific named measures can be scored, but representativeness of the human brain overall has no single observable and is argued rather than measured.","tier_confidence":"guess","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"2-5 years","is_primary":0,"rationale":"Ex vivo human brains and in vitro cortex models are biological engineering capabilities rather than computational ones.","confidence":"guess"},{"ai_type":"Prediction and modeling","maturity":"2-5 years","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Prediction and modeling","primary_maturity":"2-5 years","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[],"capabilities":[{"id":"1b4cb37e-2a00-8017-9c3a-faca36da22fb","name":"Digital Replicas of the Whole Brain of a Model Organism","description":"Construct a digital replica of a model organism’s brain (e.g., C. elegans) that accurately recapitulates neuronal activity and behavior. "},{"id":"1b4cb37e-2a00-803f-a319-d774aea566fd","name":"Datasets that Quantitatively Capture the Full Space of Behavior","description":"Use rich data collection and machine learning to correlate natural behavior with neural activity in animals and humans and decipher the “grammar” and subcomponents of bodily movement."},{"id":"1b4cb37e-2a00-8040-b9d2-f221d4502f91","name":"Mapping the Hypothalamus and Brainstem","description":"Systematically map how specific brain regions like the hypothalamus and brainstem (arguably the “steering subsystem” of the brain) drive innate behaviors and learning signals, and understand their role in obesity, chronic pain, fertility, inflammation, and other disorders."},{"id":"1b4cb37e-2a00-80c8-8277-dc36cd05a858","name":"Data–Driven AI Models of Brain Systems","description":"Use machine learning to construct functional digital emulations of human\n  and primate brain systems. These models, built at varying levels of fidelity,\n  support automated interpretability and data-driven discovery."},{"id":"1b4cb37e-2a00-80dd-bd1d-c9863e9b5fae","name":"Embodied Testbeds for Neuro–Cognitive Models","description":"Develop virtual models (e.g., a virtual fly, rodent) to simulate neuro–cognitive processes in a controlled, embodied environment."},{"id":"1b4cb37e-2a00-80f6-8f74-c5228ba7e8d8","name":"In Vitro Models of the Human Cortex","description":"Develop an in-dish model that mimics the structure and function of the human cortex, providing a controllable platform for studying cortical development, function, and disease."},{"id":"1b4cb37e-2a00-80f7-9b26-d31b726d6b88","name":"Ex Vivo Human Brains","description":"Utilize live human brains maintained ex vivo (“in a vat”) to study\n  disease and drug responses more accurately."}]},{"id":"1b4cb37e-2a00-80e4-bb14-ecc39aa04c12","slug":"most-of-the-human-brain-remains-inaccessible","name":"Most of the Human Brain Remains Inaccessible","description":"Large portions of the living human brain are difficult to observe and modulate with current technologies. Safer, noninvasive, or minimally invasive methods are needed to capture real-time brain state information. \n\r\nOne funding program dedicated to making advancements in this space is that of ARIA (UK science R&D agency), which launched the Scalable Neural Interfaces opportunity space to support a new suite of tools to interface with the human brain at scale.","field":"Neuroscience","outcome":"Read and modify brain state across the whole living human brain without opening the skull, which takes neural interfaces from a few thousand patients to a general population.","outcome_rationale":"Their description names safer, noninvasive or minimally invasive real-time access as the specific requirement.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Channels recorded, spatial and temporal resolution, depth reached and invasiveness are direct device specifications.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Measurement and sensing","maturity":"2-5 years","is_primary":1,"rationale":"Independent passes disagreed (type: Physical build vs Measurement and sensing; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","confidence":"guess"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"2-5 years","primary_rationale":"Independent passes disagreed (type: Physical build vs Measurement and sensing; ). v2 taken because it had the complete eight-category taxonomy available; flagged guess because two independent labelers reading the same text did not converge.","primary_confidence":"guess","indicators":[],"capabilities":[{"id":"1b4cb37e-2a00-802d-bceb-e985e3133c26","name":"Noninvasive Blood–Based Measurement of Brain Biomarkers","description":"Use peripheral sampling methods to indirectly monitor brain molecular\n  biomarkers. One approach involves using ultrasound to transiently open the\n  blood–brain barrier, releasing engineered protein markers into the\n  bloodstream for detection."},{"id":"1b4cb37e-2a00-804d-8639-eaa7b5153815","name":"Minimally Invasive Ultrasound–Based Whole\n  Brain Computer Interface","description":"Develop a minimally invasive ultrasound-based platform that can interface\n  programmably with the whole human brain. This approach leverages ultrasound’s\n  ability to penetrate deep tissues, offering scalable imaging and modulation\n  with minimal invasiveness."},{"id":"1b4cb37e-2a00-80af-8c60-cca83e5a230c","name":"Less Invasive Access to the Brain Through Very Tiny Skull Holes","description":"Access brain fluids through much tinier holes than currently possible to facilitate less invasive delivery of drugs or devices to intracranial or intraventricular spaces."},{"id":"1b4cb37e-2a00-80bd-bee4-fd98fc7e54d8","name":"Fully Noninvasive Neural Read–Write Technologies","description":"Use noninvasive modalities—such as ambient field magnetoencephalography\n  (MEG) with quantum gradiometers, sono–magnetic tomography, optical\n  interference methods, and ultrasound modulated optical tomography—to record\n  and modulate brain activity without surgery."},{"id":"1b4cb37e-2a00-80f0-965d-e326c333b76e","name":"Micro- to Nano-Scale Minimally Invasive\n  BCI Transducers","description":"Harmless nanoscale transducers to record or modulate brain activity that\n  can be delivered minimally invasively."},{"id":"1c0cb37e-2a00-8089-8cfe-eb2ad9182896","name":"Precision Immunology for Indirect Read/ Write Access to the Central Nervous System","description":"Use peripheral immune cells as both reporters and interventions to decode and influence the brain. This approach leverages the adaptive immune system’s inherent function as sentinel and archivist for etiological and pathological processes in other end organs (including the brain)."},{"id":"1c0cb37e-2a00-80e0-9a6e-e255870837e2","name":"Gene Expression Control of Cellular Transplants in the Brain","description":"Technologies to control gene expression in single neurons post-transplantation. Light-based or acoustic methods could offer precision for neuro-activation to enable axonal guidance and integration and enable cellular transplantation for neuroregeneration, circuitry reconstruction etc."}]},{"id":"1b3cb37e-2a00-80cd-9213-c8eb788cc657","slug":"most-brain-circuitry-is-still-invisible","name":"Most Brain Circuitry is Still Invisible","description":"Understanding the complete wiring of the brain at single–cell resolution, along with detailed molecular annotations, is critical for revealing how neural circuits support learning, memory, and behavior. Current technologies are prohibitively expensive and lack scalability, limiting our ability to link molecular composition with circuit connectivity and to understand the alterations present in brain disorders. This gap fundamentally makes diagnosis, treatment, and prevention of many brain disorders more difficult. Beyond the biomedical applications, maps of brain circuitry could play a fundamental role in grounding principles of safety for brain-like AI systems.\n\nInitiatives like the NIH BRAIN Initiative’s transformative projects (the BRAIN Initiative Cell Atlas Network (BICAN), the BRAIN Initiative Connectivity Across Scales (BRAIN CONNECTS) Network, and the Armamentarium for Precision Brain Cell Access) represent important efforts to illuminate foundational principles governing the circuit basis of behavior and to inform new approaches to treating human brain disorders by radically enhancing our understanding of brain cell types and the tools needed to access them (The BRAIN Initiative® 2.0: From Cells to Circuits, Toward Cures).","field":"Neuroscience","outcome":"Complete wiring diagrams with molecular annotation, affordable at brain scale. Circuit-level explanations of behaviour and disorder turn testable rather than assumed.","outcome_rationale":"Their description ties the gap directly to diagnosis, treatment and grounding safety principles for brain-like AI.","outcome_confidence":"confident","tier":"Directly measurable","tier_rationale":"Cubic millimetres reconstructed per dollar, synapses annotated and error rate per reconstructed neuron are direct and published.","tier_confidence":"confident","is_new":0,"ai_types":[{"ai_type":"Physical build","maturity":"2-5 years","is_primary":0,"rationale":"Their stated blocker is that the imaging hardware is prohibitively expensive and unscalable, which is instrument throughput.","confidence":"confident"},{"ai_type":"Measurement and sensing","maturity":"Working now","is_primary":1,"rationale":"both independent passes agreed","confidence":"confident"}],"primary_ai_type":"Measurement and sensing","primary_maturity":"Working now","primary_rationale":"both independent passes agreed","primary_confidence":"confident","indicators":[{"id":3,"quantity":"Volume of brain tissue reconstructed at synaptic resolution in a single dataset","current_value":"1.0","unit":"mm³","as_of":"2025-04-09","target_value":"≈500 (one whole mouse brain)","target_basis":"NIH BRAIN CONNECTS (RFA-NS-22-049) names the transformative project as 'a complete nanometer-level reconstruction of an entire mouse brain'. The programme states the goal in whole-brain terms rather than in mm³; the ≈500 mm³ conversion is ours and is approximate.","source_title":"Functional connectomics spanning multiple areas of mouse visual cortex","source_url":"https://www.microns-explorer.org/cortical-mm3","source_doi":"10.1038/s41586-025-08790-w","source_checked":"verified","reads_as":"The largest volume ever reconstructed at synaptic resolution is about one cubic millimetre of mouse visual cortex: 200,000 cells and 523 million synapses.","direction":"higher is better","context":"A whole mouse brain is roughly 500 mm³, and NIH BRAIN CONNECTS aims at exactly that. The gap is a factor of about 500.","caveat":"Cost per mm³ would be the better number and is not published in a form comparable across projects.","is_null_result":0,"rationale":"The MICrONS cortical mm³ dataset is 1.4 × 0.87 × 0.84 mm of mouse primary visual cortex, containing more than 200,000 cells, 120,000 neurons and 523 million automatically detected synapses; the figures are read off the MICrONS Explorer dataset page and the flagship paper is Nature 640(8058):435–47. Volume is the right axis because it is the one Convergent's phrase 'prohibitively expensive and lack scalability' is about, and because the same quantity is comparable across the human H01 fragment (also roughly one cubic millimetre, 1.4 petabytes) and the completed FlyWire fly brain. The gap between the current 1 mm³ and the programme's own stated whole-mouse-brain goal is about a factor of 500, and that ratio is the finding. Cost per mm³ would be the better indicator and is not published in a form that can be compared across projects, which is why volume is used instead.","confidence":"confident"}],"capabilities":[{"id":"1b4cb37e-2a00-8036-aa4f-f8930ffac5ee","name":"Faster Electron Microscopy for Connectomics","description":"Develop new physical detection methods for electron microscopy that improves scalability of visualization of brain circuitry.\n    \n    "},{"id":"1b4cb37e-2a00-80ae-bb31-f2cc484cbf8f","name":"Scalable, Anatomically and Molecularly Dense Brain Mapping Technology","description":"Develop a scalable technology that can map the brain’s wiring at the single–cell level and link molecular and circuit properties."},{"id":"1b4cb37e-2a00-80cc-8e5f-f4dc2d15c252","name":"Rapid High–Resolution Neural Circuit Visualization","description":"Combine advanced imaging\n  methods—such as synchrotron X–ray microscopy and expansion microscopy—to\n  rapidly and scalably image neural circuits in both small and large brain\n  regions. This method promises high spatial resolution with faster throughput.\n    \n    "}]}],"new_gaps":[{"id":"new-coating-thermal-noise","name":"Coating Thermal Noise Sets the Sensitivity Floor for Precision Optical Instruments","slug":"coating-thermal-noise-sets-the-sensitivity-floor-for-precision-optical-instruments","description":"The mechanical loss of dielectric mirror coatings sets a Brownian noise floor that limits gravitational wave detectors, optical clocks and cavity-stabilized lasers alike. We fund coating work as a subsystem of each instrument, so the improvement that would serve all three is nobody's program. Lower-loss coatings deposited at meter scale are needed.","field_id":"1c1cb37e-2a00-8018-9437-eaafd782ddbb","outcome":"A whole class of precision instruments gains sensitivity at once rather than one at a time, and third-generation gravitational wave detectors reach their design sensitivity in the band where they are most useful.","ai_type":"Design search","maturity":"2-5 years","tier":"Directly measurable","tension_test":"Wide agreement that the goal is transformative: the Einstein Telescope and Cosmic Explorer design cases both state that coating thermal noise limits sensitivity in the most sensitive frequency band, and ET's coatings programme describes minimal-noise coatings as essential to reaching expected sensitivity. Genuine debate about near-term feasibility: amorphous coatings have resisted large loss-angle reductions for two decades, crystalline AlGaAs and mixed TiO2:GeO2 candidates each have unresolved problems at metre-scale deposition, and there is no consensus on whether the required loss angle is reachable at all in an amorphous material. That is the tension: agreement on the prize, no agreement on whether it is available.","unlock_test":"The dominoes are unusually well aligned because the quantity is shared. A lower loss angle improves gravitational wave detector reach directly; it also improves optical cavity stability, which sets the short-term stability of the best optical clocks, which in turn sets the sensitivity of clock-based tests of fundamental physics and of relativistic geodesy. Convergent's own Astrophysics gap on gravitational wave frequency coverage sits downstream of coating performance without naming it. Gate B corrected an overreach here: the formation flying gap was previously listed as downstream too, and it is not — sub-micron spacecraft station-keeping is a controls problem that a better coating does not touch.","dedup_check":"Grepped all 103 gaps and all 369 capabilities on /coating|thermal noise|mechanical loss|gravitational wave|interferomet|precision measurement|noise floor/. Two gaps matched, both Astrophysics: 'Limited Detection of Gravitational Waves Across the Frequency Spectrum', which is about extending detection to bands outside the audio band and whose three capabilities are space-based, decihertz and high-frequency detection; and 'Higher-Resolution Views of the Universe Are Roadblocked by Formation Flying Technology', which is about sub-micron spacecraft station-keeping. Five capabilities matched, none about coatings or mirror materials. The four Materials Science gaps cover discovery search, crystallization, scalable macroscale synthesis and element acquisition. Closest match overall is the gravitational wave frequency gap; it is distinct because frequency coverage and in-band sensitivity are different axes and a perfect coating does nothing for the first.","nearest":"Closest is your Astrophysics gap “Limited Detection of Gravitational Waves Across the Frequency Spectrum”. It is about extending detection into new frequency bands; this is about sensitivity inside the band you already cover, which a perfect coating improves and a new band does not.","funding_check":"Not clear of funding, and the honest version matters more than a clean answer. The NSF has funded the LSC Center for Coatings Research since 2017 (award 1707868, renewed as 2429369), whose stated premise is that 'the A+ upgrade and all 3rd generation detector designs depend on the development of mirrors with low coating thermal noise'. Italy's ETIC project funds coatings work for the Einstein Telescope through INFN, INAF and the universities of Padua and Bologna, testing mixed titanium and silicon dioxide films. Both are funded as detector subsystems inside one instrument programme. What is not funded anywhere found is the cross-instrument materials problem: no programme treats low-mechanical-loss optical coatings as a materials target in its own right, with the result that optical clock and cavity groups re-derive the same materials work. The gap proposed here is that framing, not the existence of coatings research, and it is stated that way deliberately.","rationale":"Chosen as the safe case: unambiguous metric (coating loss angle, and coating thermal noise in the detector band), a concrete mid-scale build, and a large downstream unlock. It is tier 1 on their own roadmapping criterion, which is why it is offered first. REVISED after Gate B, which confirmed novelty against both the 103 gaps and the 369 capabilities with its own search terms, and against Convergent's twelve-FRO portfolio, but found two real problems. The description previously asserted that coating improvements 'transfer poorly between instrument classes', which the field's flagship result refutes: Cole, Zhang, Martin, Ye and Aspelmeyer, 'Tenfold reduction of Brownian noise in optical interferometry', Nature Photonics 2013 (arXiv:1302.6489), reports a tenfold reduction in mechanical damping from crystalline coatings and frames it for 'the next generation of ultra-sensitive interferometers, as well as for new levels of laser stability' — both instrument classes at once, with a clock physicist and an optomechanics physicist among the authors. The claim has been narrowed to what survives: transfer happens, but it is not what the funding is organized to produce. The formation flying claim in unlock_test was simply wrong and has been removed.","confidence":"confident","created_at":"2026-09-06 16:38:41","field":"Materials Science"},{"id":"new-record-of-what-did-not-work","name":"There Is No Machine-Readable Record of What Did Not Work","slug":"there-is-no-machine-readable-record-of-what-did-not-work","description":"Negative results, null findings and abandoned approaches are largely missing from the scientific record, and where they are captured the formats are neither standardized nor machine-readable. Systems that could learn from failure have little to learn from. The obstacle is credit, not storage.","field_id":"1c7cb37e-2a00-806b-a9e9-fbec288c4ed2","outcome":"The failed half of the scientific record becomes available to both people and models, so an approach that has already been tried and abandoned can be recognised as such before it is funded again.","ai_type":"Coordination and institutions","maturity":"2-5 years","tier":"Proxy only","tension_test":"Wide agreement that the goal is transformative: publication bias is one of the oldest and best-documented pathologies in metascience, and 2025 saw a values-based framework for surfacing null results published in PLOS Biology. Genuine debate about feasibility: every previous attempt at a negative-results journal or database has struggled, because the cost of depositing falls on the author and the benefit accrues to everyone else. The debate is not about whether the record would be useful, it is about whether any credit arrangement can make depositing rational for the person who has to do it.","unlock_test":"Downstream of this sit three of Convergent's own metascience gaps. Literature synthesis at scale is synthesising a censored corpus. Fraud detection is easier when the distribution of real null results is known, because implausible cleanliness is what detection methods look for. Clinical trial optimisation depends on knowing which designs have already failed. It also acts on the materials and chemistry gaps: their own 'Open Synthesis Database' capability asks for procedure logs including failed experiments, which is this gap restricted to one field.","dedup_check":"Grepped all 103 gaps and all 369 capabilities on /negative result|null result|failed experiment|unpublished|rejected proposal|reproducib|replicat|preregist|machine-readable|data sharing/. Six gaps matched, none about the missing record: they are about synthetic olfaction, manual chemical synthesis, manual bioengineering, evolved computation, superconductivity noise, and ephemeral platform data. Convergent's five metascience gaps cover fraud, synthesis at scale, publishing cost, clinical trial design and organisational rigidity, and none covers this. The closest match in the whole export is a capability, not a gap: 'Open Synthesis Database', attached to the Chemistry gap 'Limited Understanding of the Chemical Reaction Space', which asks for 'procedure logs and outcomes including failed experiments'. That is the same idea scoped to materials synthesis. The proposal here is that it generalises, and that generalising it is the gap.","nearest":"Closest is not a gap but a capability: “Open Synthesis Database”, attached to your Chemistry gap “Limited Understanding of the Chemical Reaction Space”, which asks for procedure logs including failed experiments. That is this idea scoped to materials synthesis. The proposal is that it generalises, and that generalising it is the gap.","funding_check":"Initiatives exist and none is infrastructure. The Null Hypothesis Initiative and its Null Compass tool are running; ICPC and other venues run replication-and-negative-results tracks; Open Grants (ogrants.org) collects funded and unfunded proposals voluntarily; NegResCo built a COVID-specific negative results collection on ELIXIR platforms. Reporting on these efforts is consistent that they lack coordination and that data formats are not standardised. On the proposal side, the case for releasing unfunded applications is argued in Issues in Science and Technology and an OpenProposal platform was proposed in a 2025 preprint; no funder has implemented it. Nothing found is a funded, cross-field, machine-readable record under construction. Gate B found one omission worth naming: Octopus (https://www.octopus.ac/), which received GBP 650,000 from Research England's emerging priorities fund in August 2021 and is run with Jisc and the UK Reproducibility Network. Octopus breaks a paper into eight independently publishable, linked units — problem, hypothesis, method, results, analysis, interpretation, real-world application and peer review — which is a structured, machine-readable primary record and is closer to this gap than anything previously listed here. It is not the same thing, and the distinction is narrow enough to state carefully: the eight types are stages of the research process, not outcome categories, so Octopus does not create a negative-results record as such. What it does is decouple depositing a result from writing a narrative paper around it, which removes one of the two barriers this gap names while leaving the credit arrangement untouched.","rationale":"Proposed at tier 3 in effect and labelled Proxy only rather than dressed up: deposited records are countable, but the quantity that matters is the share of failures captured, and that denominator is unobservable by construction. Their own roadmapping criterion asks whether success is unambiguously measurable, and for this gap the honest answer is no. The audit in Phase 2 found Proxy only to be the least reliable tier in the taxonomy, which is a caveat on this label rather than on the gap. REVISED after Gate B, which confirmed novelty by reading in full every capability that could plausibly cover this — including 'Open Synthesis Database', 'Materials Property Bank', 'Intelligent Databases', 'Infrastructure for Research Curation' and 'Create an Internet Archive for Critical Data' — and reached the same conclusion by its own route. It also verified the productive tension with named opposition on both sides, Fanelli against the 50-author PLOS Biology consensus. Octopus was added to the funding check as the nearest funded thing.","confidence":"confident","created_at":"2026-09-06 16:38:41","field":"Metascience"}],"critical_paths":[{"id":"path-publishing-cost","gap_id":"1c1cb37e-2a00-8063-9b9f-f3c2aa784e50","title":"From draft to credited contribution: what sets the cost of publishing research","axis":"Cost, measured as reviewer and editor labour per published paper","axes_excluded":"Speed, meaning elapsed time from submission to publication, and inclusiveness, meaning who can afford to publish and who can read the result. Their sentence bundles all three. Speed in particular behaves differently by venue: journal latency is reviewer-supply-driven while conference latency is set by a fixed programme committee calendar, so a chain mixing them would produce a cost concentration that is an artifact of the venue mix.\n\nAlso excluded: the 'doing' half of their gap statement. 'Doing and publishing research is expensive' covers the cost of the research itself, which is most of what the other 102 gaps in the map are about. Instruments, reagents, compute and staff are priced elsewhere. This chain runs from a finished result to a credited contribution.\n\nAlso outside the axis, and worth stating because it is the biggest number in the room: author time. The axis counts reviewer and editor labour, so the weeks a researcher spends writing the paper are not in it. That is very likely the largest single cost in getting a result published, and this chain does not measure it.","expectation":"This chain must produce a different answer from chain 1 or the pair demonstrates nothing. AI does act on several links here: drafting and screening work now, and reviewer-to-proposal matching is tractable, though there is published evidence that traditional statistical representations outperform generative AI at identifying expert reviewers. So the expectation is not that AI does nothing. It is that AI acts on the links that were never rate-limiting, while reviewer recruitment, judgment consistency and the credit and legitimacy step remain untouched. Stated explicitly so it cannot be read the other way: this is not a claim that publishing is a cognitive bottleneck. It is not, and asserting it would be wrong. Recorded before the link analysis.","finding":"Drafting was expensive, and AI has made it much cheaper. That saving is real and large. It has not arrived as cheaper publishing: submissions rose 42% after ChatGPT's release relative to the prior two-year window, in the one corpus where a journal has published full figures, and the labour that the saving displaced landed further down the chain rather than disappearing.\n\nWhere exactly it landed is a correction Gate C forced, and it matters because it was this chain's marquee result. I wrote that the labour moved to reviewer recruitment. The source does not say that. Pierce, Gartenberg, Hasan and Murray put the displaced load on volunteer editors at desk screening — among manuscripts with 70%+ AI scores nearly 70% are desk-rejected, against 44% for low-AI submissions — and give no invitation or recruitment figures at all. So the arrow from cheaper drafting runs to screening, which is a step this chain calls tractable and AI-reached. Recruitment strain is real and separately evidenced, by Silverchair's 4.5 invitations per accepted review and the fall in acceptance from 43% to 22%; it is simply not what the 42% surge is shown to have caused. The general claim survives and the specific one does not: relieving a step upstream of a concentration moves cost downstream, and it moved to the nearest downstream step, not the furthest.\n\nThe prediction holds, in the shape the pair needs. AI is not absent here: it acts on the first four of seven steps, which is more than it touches anywhere in chain 1. It can match a reviewer to a paper. It cannot make that reviewer say yes, agree with the other reviewer, or persuade a hiring committee to count the work.\n\nRecruitment is where the axis choice earns itself. Splitting it into matching and willingness separates a well-posed prediction problem from a labour-supply problem. On reviewer identification specifically, published evidence finds traditional statistical representations outperform generative AI. Gate C's caveat is recorded with it: ESO names matching difficulty as a reason for distributed peer review too, so the split is this analysis's, not ESO's.\n\nOne observation about your capability set, and it runs the other way from chain 1. Two of the four capabilities you attach to this gap, a post-publication peer review layer and new protocols for knowledge production and verification, act directly on credit and legitimacy, which carries cost. For the telescope gap none of the three touches a decision step; here half of them do.\n\nTo be plain about scope: this chain covers getting a finished result into the record and credited. The cost of doing the research is the other half of their gap statement and is most of what the rest of the map is about.","duration_basis":null,"programmes_json":"[]","axis_kind":"cost","reviewed":"human","created_at":"2026-09-06 16:38:41","programmes":[],"gap_name":"Doing and publishing research is expensive and subject to structural roadblocks","gap_field":"Metascience","links":[{"path_id":"path-publishing-cost","seq":1,"link":"Production and drafting","blocker":"Author time. Almost certainly the largest single cost in getting a paper published, and the one this chain's axis does not count.","ai_type":"Reading and synthesis","maturity":"Working now","is_binding":0,"evidence":"Organization Science AI Task Force, 'More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review', Organization Science editorial, 2026, doi 10.1287/orsc.2026.ed.v37.n3. Roughly 7,000 manuscripts with full-submission access. Submissions up 42% since ChatGPT's November 2022 release against the prior two-year window; for scale, the COVID-19 period produced a 20% bump. The rise is almost entirely manuscripts with substantial AI-generated text, and submissions scoring low for AI have declined over the same period.","rationale":"Scope first, because it decides how to read this step. The axis is reviewer and editor labour per published paper, so author time sits outside it by construction. Writing the paper is very probably the most expensive thing that happens in this pipeline, and this chain does not measure it. What it can say is that AI has taken a large share out of the per-paper drafting cost, that the number of papers went up, and that the labour this displaced arrived downstream: submissions rose 42% after ChatGPT's release against the prior two-year window, in the one corpus with published full-submission figures, and the load landed on volunteer editors at desk screening. Whether total author cost across all papers fell is unknown and this chain has no way to find out. Two corrections from Gate C are folded in above: the 42% is growth since ChatGPT within a five-year corpus, and the source places the displaced load at screening.","duration_years":null,"duration_span":null,"duration_note":null,"figure":"author time, not counted on this axis; submissions +42% since ChatGPT's release in November 2022 against the prior two-year window, roughly 7,000 manuscripts (Organization Science, doi 10.1287/orsc.2026.ed.v37.n3)","duration_days":null,"duration_span_note":"No published figure. Author time sits outside this chain's axis and no one reports a median for it.","duration_covers_json":null,"ai_acts":1,"capabilities_json":"[]","capabilities":[],"duration_covers":null,"capability_links":[]},{"path_id":"path-publishing-cost","seq":2,"link":"Submission and desk screening","blocker":"Editor triage time.","ai_type":"Reading and synthesis","maturity":"Working now","is_binding":0,"evidence":"ESO reports that classical triage 'has significantly degraded the quality of feedback for the triaged proposals' — screening under load trades feedback quality for throughput.","rationale":"Not binding on the cost axis. Screening is genuinely tractable for current models, which is why it is already being done and why it does not set the cost.","duration_years":null,"duration_span":null,"duration_note":null,"figure":"triage under load 'significantly degraded the quality of feedback' (ESO)","duration_days":119,"duration_span_note":"Submission to acceptance, which brackets steps 2 to 5 together rather than any one of them. Median 119 days across ophthalmology journals in 2020, IQR 83-168. Carried on step 2 because that is where the span opens, and it is not a measurement of desk screening alone.","duration_covers_json":"[2,3,4,5]","ai_acts":1,"capabilities_json":"[]","capabilities":[],"duration_covers":[2,3,4,5],"capability_links":[]},{"path_id":"path-publishing-cost","seq":3,"link":"Reviewer recruitment and matching","blocker":"Willingness to serve, and matching. The chain's argument turns on the first, and Gate C established that the source cannot be used to dismiss the second: ESO's own Phase 1 page lists 'it has become progressively more difficult to find optimal proposal-referee matches' as a distinct reason for deploying distributed peer review, beside the supply bullet the artifact quotes. Both are real. The claim that survives is narrower — AI reaches the matching half and not the willingness half — and it no longer rests on ESO having said matching is solved.","ai_type":"Prediction and modeling","maturity":"2-5 years","is_binding":1,"evidence":"4.5 invitations per accepted review, nearly double the 2018 rate, and reviewer acceptance down from 43% in 2018 to 22% in 2024 (Silverchair, Future of Peer Review 2026, from eight years of ScholarOne Manuscripts activity plus 2,000+ survey responses). Per 100 invitations sent, editors wait a combined 407 days on reviewers who ultimately decline or never answer. A survey of 139 editors of Australian journals, largely local and society titles, with 27 interviews, finds 55% rate finding reviewers a significant or very significant challenge, and some 'described having to send out 30 or more invitations to secure just two reviewers' (Jamali et al., Learned Publishing 2026, doi 10.1002/leap.2034, and Luca et al., doi 10.1002/leap.2041; popularised in The Conversation, 15 February 2026). The sample is national and small-journal weighted, which is where recruitment is hardest, so the bias runs in this chain's favour and is noted for that reason. Meyerson documents 21 years of declining reviewer acceptance, 2002-2024. ESO's Phase 1 page on distributed peer review states that 'it has become progressively harder to find scientists willing to serve in the panels and in the OPC'. And on the part AI could do: traditional statistical representations outperform generative AI at identifying expert peer reviewers, tested on NASA/ADS records and observatory DPR data across 379 researchers.","rationale":"Binding, and it is the link that most rewards being split in two. Matching is a well-posed prediction problem where AI is applicable and, on current evidence, not even the best method. Willingness is a labour-supply problem that no matching system addresses. The binding half is the half AI does not touch, which is the whole finding of this chain in one link. REVISED after Gate C. The previous version asserted that matching is not the problem and cited ESO for it; ESO says both. The split between a well-posed prediction problem and a labour-supply problem still does the work this chain needs, but it is now stated as this analysis's distinction rather than the source's.","duration_years":null,"duration_span":null,"duration_note":null,"figure":"4.5 invitations per accepted review, nearly double 2018 (Silverchair, Future of Peer Review 2026); reviewer acceptance fell from 43% in 2018 to 22% in 2024; 55% of 139 editors of Australian journals call recruitment a significant or very significant challenge; 21 years of declining acceptance","duration_days":null,"duration_span_note":"Measured inside the 119-day submission-to-acceptance span, which brackets steps 2 to 5 together. No figure separates this step from the others in that span.","duration_covers_json":null,"ai_acts":1,"capabilities_json":"[]","capabilities":[],"duration_covers":null,"capability_links":[]},{"path_id":"path-publishing-cost","seq":4,"link":"Review judgment","blocker":"Irreducible subjectivity in ranking work that is above the bar.","ai_type":"Reading and synthesis","maturity":"2-5 years","is_binding":1,"evidence":"NeurIPS 2021 consistency experiment: two independent committees disagree on 23% of papers, and approximately half the accepted list would change on a random rerun — consistent with 26% in 2014. Cortes and Lawrence, revisiting the 2014 experiment, find no correlation between review scores and later citation impact for accepted papers, and a correlation for rejected ones: 'the reviewing process for the 2014 conference was good for identifying poor papers, but poor for identifying good papers.'","rationale":"Binding, and unusually well quantified for an institutional link. The Cortes and Lawrence result also sets the bar any automated reviewer has to clear, and it is a low bar in one direction and an unreachable one in the other: filtering bad work is where review already works, ranking good work is where it does not, and it is not established that the second is a solvable problem for anyone.","duration_years":null,"duration_span":null,"duration_note":null,"figure":"23% committee disagreement; ~half the accept list changes on a rerun; scores predict impact for rejected papers only","duration_days":null,"duration_span_note":"Measured inside the 119-day submission-to-acceptance span, which brackets steps 2 to 5 together. No figure separates this step from the others in that span.","duration_covers_json":null,"ai_acts":1,"capabilities_json":"[\"Post-Publication Peer Review Layer\"]","capabilities":["Post-Publication Peer Review Layer"],"duration_covers":null,"capability_links":[{"name":"Post-Publication Peer Review Layer","url":"https://www.gap-map.org/capabilities/post-publication-peer-review-layer/","initiatives":[{"title":"APPRAISE (A Post-Publication Review and Assessment In Science Experiment)","url":"https://asapbio.org/eisen-appraise"},{"title":"PREreview","url":"https://prereview.org/"}]}]},{"path_id":"path-publishing-cost","seq":5,"link":"Editorial decision","blocker":"Editor labour, and accountability for the decision.","ai_type":"Coordination and institutions","maturity":"2-5 years","is_binding":0,"evidence":"Follows directly from the recruitment and judgment links; editors carry the residual cost when reviews do not arrive.","rationale":"Not binding on the cost axis. Its cost is largely inherited from link 3.","duration_years":null,"duration_span":null,"duration_note":null,"figure":"cost inherited from link 3","duration_days":null,"duration_span_note":"Measured inside the 119-day submission-to-acceptance span, which brackets steps 2 to 5 together. No figure separates this step from the others in that span.","duration_covers_json":null,"ai_acts":0,"capabilities_json":"[]","capabilities":[],"duration_covers":null,"capability_links":[]},{"path_id":"path-publishing-cost","seq":6,"link":"Dissemination","blocker":"Article processing charges and platform cost.","ai_type":"Coordination and institutions","maturity":"Working now","is_binding":0,"evidence":"Three of Convergent's four capabilities for this gap act here: disrupting traditional publishing models, frugal science initiatives, and the micro/nanopublishing half of new protocols for knowledge production and verification. That third one is counted at link 7 as well, because its description covers both lowering the barrier to publishing and new norms for recognising work; it is one capability acting on two steps, not two capabilities.","rationale":"Not binding on the labour-cost axis, and this deserves care because it is the link their gap sentence centres on. On the inclusiveness axis — explicitly excluded from this chain — it may well be binding. That is why the axes were separated.","duration_years":null,"duration_span":null,"duration_note":null,"figure":"article processing charges and platform cost","duration_days":30,"duration_span_note":"Acceptance to first online release. Median 30 days, IQR 10-71, same 2020 ophthalmology cohort.","duration_covers_json":"[6]","ai_acts":0,"capabilities_json":"[\"Disrupt Traditional Publishing Models\",\"Frugal Science Initiatives\",\"New Protocols for Knowledge Production and Verification\"]","capabilities":["Disrupt Traditional Publishing Models","Frugal Science Initiatives","New Protocols for Knowledge Production and Verification"],"duration_covers":[6],"capability_links":[{"name":"Disrupt Traditional Publishing Models","url":"https://www.gap-map.org/capabilities/disrupt-traditional-publishing-models/","initiatives":[{"title":"Arcadia Science publications","url":"https://research.arcadiascience.com/"},{"title":"ResearchHub","url":"https://www.researchhub.com/about"},{"title":"Unjournal","url":"https://www.unjournal.org/"}]},{"name":"Frugal Science Initiatives","url":"https://www.gap-map.org/capabilities/frugal-science-initiatives/","initiatives":[{"title":"Frugal Science","url":"https://www.frugalscience.org/"}]},{"name":"New Protocols for Knowledge Production and Verification","url":"https://www.gap-map.org/capabilities/new-protocols-for-knowledge-production-and-verification/","initiatives":[{"title":"Cosmik","url":"https://www.cosmik.network/"},{"title":"Nanopublications","url":"https://nanopub.net/"}]}]},{"path_id":"path-publishing-cost","seq":7,"link":"Credit and legitimacy","blocker":"Hiring, tenure and funding committees decide what counts. Nothing in the pipeline can make them count something new.","ai_type":"Coordination and institutions","maturity":"Speculative","is_binding":1,"evidence":"A model can review a paper today; it cannot make a hiring committee count that review. The same asymmetry explains why reviewer supply falls: the labour is unpriced and uncredited, and pricing it is an institutional act.","rationale":"Binding, and the only link in either chain where the blocker is purely a matter of what institutions agree to recognise. Two of Convergent's four capabilities for this gap — a post-publication peer review layer, and new protocols for knowledge production and verification — act here.","duration_years":null,"duration_span":null,"duration_note":null,"figure":"no published quantity, the blocker is what committees agree to count","duration_days":null,"duration_span_note":"No published figure, and probably not measurable. The quantity is how long institutions take to count a new kind of work.","duration_covers_json":null,"ai_acts":0,"capabilities_json":"[\"Post-Publication Peer Review Layer\",\"New Protocols for Knowledge Production and Verification\"]","capabilities":["Post-Publication Peer Review Layer","New Protocols for Knowledge Production and Verification"],"duration_covers":null,"capability_links":[{"name":"Post-Publication Peer Review Layer","url":"https://www.gap-map.org/capabilities/post-publication-peer-review-layer/","initiatives":[{"title":"APPRAISE (A Post-Publication Review and Assessment In Science Experiment)","url":"https://asapbio.org/eisen-appraise"},{"title":"PREreview","url":"https://prereview.org/"}]},{"name":"New Protocols for Knowledge Production and Verification","url":"https://www.gap-map.org/capabilities/new-protocols-for-knowledge-production-and-verification/","initiatives":[{"title":"Cosmik","url":"https://www.cosmik.network/"},{"title":"Nanopublications","url":"https://nanopub.net/"}]}]}]},{"id":"path-telescope-elapsed-time","gap_id":"1c1cb37e-2a00-8008-9ca1-cf81d4b44116","title":"From science case to first light: what sets the elapsed time of a frontier telescope","axis":"Elapsed time from first concept study to first light","axes_excluded":"Cost per unit of collecting area, and cost per unit of science return. Their gap statement bundles cost with schedule — 'cost-prohibitive and slow' — and the two run over a partly different set of links. Each is a separate chain, and a chain that mixes them produces a binding link that is an artifact of the mixing rather than a property of the world.","expectation":"Design and optimisation search helps at the concept and optics stages, and the binding links will fall on the physical build and on the funding decision, neither of which any current AI capability touches. If that holds, the demonstration is that closing every cognitive link in the chain changes the total duration very little. A second prediction, stated separately so it can fail separately: Convergent's three capabilities attached to this gap — modular assembly with reduced launch costs, leveraging commercial component advances, and a space telescope factory — all act on the fabrication and assembly links, and none acts on decision, approval or funding. Both predictions are recorded before any programme history was assembled, and both are testable against published JWST, Rubin and ELT milestone dates.","finding":"AI acts on three of the eight steps. How many years those three are is the part this chain can no longer put a single number on, and saying so is the honest result.\n\nAs labelled, the three are 9.5 of the 32.5 years, leaving 23. Two independent reviews then found corrections pointing in opposite directions. Gate C found that links 2 and 3 count eight years against AI on blockers JWST's own record does not show in those windows, and that design maturation demonstrably ran four years past the interval it is given — all of which move years into the AI column and pull the residual toward the mid-teens. Gate D found the opposite at link 1: it is labelled Working now while its own blocker says the constraint is community consensus rather than analysis capacity, and correcting that moves seven years out of the AI column and pushes the residual toward thirty.\n\nSo the residual is somewhere between about fifteen and about thirty of the 32.5 years, depending on judgments the decomposition does not settle. That is a weaker claim than 23 and a more defensible one, and it is what a chain of eight hand-assigned intervals can actually support. What neither reading disturbs is the shape: on every version of the arithmetic, most of the elapsed time of a frontier telescope sits in steps no current AI capability reaches.\n\nThe prediction about your capability set holds exactly. Modular assembly with reduced launch costs, a space telescope factory, and leveraging commercial component advances all act on fabrication, integration and launch. None acts on ranking or funding. Stated as an observation and not a deficiency: those three are aimed at the part of the chain where the years actually are.\n\nOne honest correction to how I first wrote this up. I originally called four of the eight steps binding, which was circular: a chain of steps that run strictly one after another has no non-binding steps, because removing any of them shortens the total. The set I had marked was really just the set AI does not touch.\n\nWhat was not predicted is that the split moves between programmes. JWST is build-dominated, with seventeen of its thirty-two years after construction started. Rubin is decision-dominated, twenty-two years to a construction award against eleven building. The ELT sits between, with about two and a half years attributable to nothing but a funding condition on a design already approved.","duration_basis":"Per-step durations are JWST's, because its milestone record is the most completely published of the three. Real phases overlap, design maturation and fabrication in particular, so each figure is the interval between the published milestones that best bracket that step, and the eight sum to 32.5 against STScI's stated 32 from conception to launch. Those two totals are not the same measurement — 32.5 runs to first images in July 2022 and STScI's 32 runs to launch in December 2021 — so their agreement is a coincidence of rounding and is not corroboration. Gate C's audit of the intervals is the reason the finding below is stated as a bound rather than a point: two links carry titles that describe only part of what their intervals contain, and the overlap between design maturation and fabrication is real and undated.","programmes_json":"[{\"programme\":\"JWST\",\"concept\":\"1989 NGST workshop\",\"to_build\":\"2004 construction start\",\"build\":\"2004 → Dec 2021\",\"years_to_build\":15,\"years_build\":17,\"total\":32,\"note\":\"Build-dominated. Seventeen of the thirty-two years fall after construction started.\"},{\"programme\":\"Rubin Observatory\",\"concept\":\"early-1990s planning\",\"to_build\":\"Aug 2014 NSF construction award\",\"build\":\"Aug 2014 → Jun 2025\",\"years_to_build\":22,\"years_build\":11,\"total\":33,\"note\":\"Decision-dominated. Thirteen years from the 2001 National Academies recommendation alone.\"},{\"programme\":\"ESO ELT\",\"concept\":\"1998 OWL studies\",\"to_build\":\"Dec 2014 green light\",\"build\":\"Dec 2014 → Mar 2029\",\"years_to_build\":16,\"years_build\":14,\"total\":31,\"note\":\"Roughly two and a half of those years are a funding condition on an already-approved design.\"}]","axis_kind":"time","reviewed":"ai-only","created_at":"2026-09-06 16:38:41","programmes":[{"programme":"JWST","concept":"1989 NGST workshop","to_build":"2004 construction start","build":"2004 → Dec 2021","years_to_build":15,"years_build":17,"total":32,"note":"Build-dominated. Seventeen of the thirty-two years fall after construction started."},{"programme":"Rubin Observatory","concept":"early-1990s planning","to_build":"Aug 2014 NSF construction award","build":"Aug 2014 → Jun 2025","years_to_build":22,"years_build":11,"total":33,"note":"Decision-dominated. Thirteen years from the 2001 National Academies recommendation alone."},{"programme":"ESO ELT","concept":"1998 OWL studies","to_build":"Dec 2014 green light","build":"Dec 2014 → Mar 2029","years_to_build":16,"years_build":14,"total":31,"note":"Roughly two and a half of those years are a funding condition on an already-approved design."}],"gap_name":"Frontier Telescopes Are Expensive and Take Decades to Build","gap_field":"Astrophysics","links":[{"path_id":"path-telescope-elapsed-time","seq":1,"link":"Science case definition","blocker":"Community consensus on what to build, not analysis capacity. The step ends when a field agrees on one instrument, and agreement is produced by workshops and committee reports.","ai_type":"Reading and synthesis","maturity":"Working now","is_binding":0,"evidence":"JWST: NGST workshop at STScI September 1989; an STScI committee recommends a larger infrared telescope in 1995-1996. Seven years. ESO: OWL concept pursued from 1998, OWL Blue Book published end of 2005. Seven years.","rationale":"Not binding, and the reason matters for the whole chain. Synthesis and trade-study tools are mature and could compress the analytical content of these seven years substantially. They cannot compress the part that actually consumes them, which is a community converging. Treating the full seven years as AI-addressable is the most generous accounting available and is used below as an upper bound.\n\nCONTESTED after Gate D. This link is labelled maturity 'Working now' with ai_acts = 1, and its own blocker field says the constraint is 'Community consensus on what to build, not analysis capacity'. Under the efficacy reading the maturity repair established — would applying this capability move this step — those two statements cannot both be right: a capability that addresses analysis capacity does not move a step whose stated blocker is agreement. The critical_path_links.maturity column was never in the repair's scope and has never had a second labeler, which is why the label is left as it stands and flagged here rather than flipped. It carries 7.0 of the chain's 9.5 AI-acted years, so flipping it is not a detail: it would take the chain's only quantitative claim from 9.5 of 32.5 to about 2.5 of 32.5. Restating a headline on one reader's unaudited judgment is the thing the gates exist to prevent.","duration_years":7,"duration_span":"1989 → 1996","duration_note":"NGST workshop at STScI to the STScI committee recommendation","figure":null,"duration_days":null,"duration_span_note":null,"duration_covers_json":"[1]","ai_acts":1,"capabilities_json":"[]","capabilities":[],"duration_covers":[1],"capability_links":[]},{"path_id":"path-telescope-elapsed-time","seq":2,"link":"Concept studies, strategic ranking and design competition","blocker":"Two things at once, and the interval does not separate them. A fixed external cadence — a project that misses a decadal survey waits for the next one, and the wait is ten years regardless of readiness — running alongside the Phase A design competition that occupies the same years.","ai_type":"Coordination and institutions","maturity":"Speculative","is_binding":0,"evidence":"Rubin: planning from the early 1990s, recommended by a National Academies report in 2001, top-ranked large ground-based project only in the 2010 decadal survey, NSF construction award August 2014. Thirteen years from first national recommendation to construction start. ELT: ESO Council names ELTs its highest strategic priority in December 2004; construction green light December 2014. Ten years. JWST's own window, 1996 to 2002, contains the industry study teams (4 June 1997), the Yardstick Design report (6 August 1998), the Lockheed Martin and TRW Phase A selection (7 July 1999) and the TRW/Ball selection (14 August 2002), per STScI's mission timeline; the strategic ranking in it is the 2001 McKee-Taylor decadal survey, which STScI's timeline does not list. The Rubin and ELT figures above are the ones that isolate ranking cleanly; JWST's does not, and that is stated here rather than left for a reader to notice.","rationale":"Six years, and they are not six years of one thing. Gate C found that the duration_note attached to this link describes a design competition while the link title describes strategic ranking, and both happened in the window. The 2001 decadal survey ranked NGST top; the Phase A competition ran either side of it. ai_acts is recorded as 0 because the cadence is what sets the length — no amount of readiness moves a project forward inside a decadal cycle, so the marginal value of accelerating everything upstream is zero until a whole cycle is saved — but design and optimization search does act on the competition inside these years. This link is therefore the largest single source of uncertainty in the chain's arithmetic, and the finding now says so.","duration_years":6,"duration_span":"1996 → 2002","duration_note":"concept studies and three funded teams to selection of the TRW/Ball design","figure":null,"duration_days":null,"duration_span_note":null,"duration_covers_json":"[2]","ai_acts":0,"capabilities_json":"[]","capabilities":[],"duration_covers":[2],"capability_links":[]},{"path_id":"path-telescope-elapsed-time","seq":3,"link":"Phase B start and funding authorisation","blocker":"Appropriation, and conditions attached to it. A design can be complete, reviewed and approved and still wait on money. On JWST specifically this interval is the run-up from design selection to long-lead construction rather than a discrete appropriation event; the clean instances are ESO's and Rubin's.","ai_type":"Coordination and institutions","maturity":"Speculative","is_binding":0,"evidence":"ELT: ESO Council approved construction in June 2012 on the condition that contracts above 2 million euros could be awarded only once the total cost of 1,083 million euros (2012 prices) was 90% funded; the green light followed in December 2014. Roughly two and a half years of pure funding latency with the design frozen. Rubin: Final Design Review December 2013, National Science Board conditional approval May 2014, construction award August 2014.","rationale":"The ELT case is the cleanest instance available anywhere in this analysis: the delay is explicitly and only about money, with every technical question already settled. Nothing in the capability set acts on it. Gate C's correction, recorded rather than argued away: JWST's own 2002-2004 interval contains no identifiable appropriation event, only the run-up to a 3 March 2004 construction start, so the two years assigned here are carrying a blocker demonstrated on other programmes. The evidence field cites Rubin and the ELT for that reason, and it should be read that way.","duration_years":2,"duration_span":"2002 → 2004","duration_note":"design selection to the start of long-lead construction","figure":null,"duration_days":null,"duration_span_note":null,"duration_covers_json":"[3]","ai_acts":0,"capabilities_json":"[]","capabilities":[],"duration_covers":[3],"capability_links":[]},{"path_id":"path-telescope-elapsed-time","seq":4,"link":"Design maturation","blocker":"Technology readiness for long-lead items, and the design-review sequence.","ai_type":"Design search","maturity":"Working now","is_binding":0,"evidence":"JWST: TRW/Ball design selected 2002, construction of long-lead components begins 2004, instrument critical design reviews 2006, mission critical design review 2010.","rationale":"This is the link where design and optimization search is most obviously useful and most obviously mature, which is exactly why it does not set the pace. Two years from contractor selection to instrument critical design reviews. Gate C found the interval understates the activity: this link's own evidence runs to the mission critical design review, which STScI dates to 3 March 2010, with the preliminary design review in April 2008 and project confirmation in April 2009 — all inside the window the chain labels Fabrication. Design maturation on JWST ran to 2010. The interval is left at 2004-2006 because the phases genuinely overlap and no milestone pair separates them cleanly, but the overlap runs in one direction: it understates how much of the elapsed time design work occupies, and so understates the AI-acts column.","duration_years":2,"duration_span":"2004 → 2006","duration_note":"long-lead start to instrument critical design reviews and flight construction","figure":null,"duration_days":null,"duration_span_note":null,"duration_covers_json":"[4]","ai_acts":1,"capabilities_json":"[\"Leveraging Commercial Component Advances\"]","capabilities":["Leveraging Commercial Component Advances"],"duration_covers":[4],"capability_links":[{"name":"Leveraging Commercial Component Advances","url":"https://www.gap-map.org/capabilities/leveraging-commercial-component-advances/","initiatives":[{"title":"Argus Array","url":"https://evryscope.astro.unc.edu/"},{"title":"Morning Star Missions to Venus","url":"https://venuscloudlife.com/"},{"title":"The Dragonfly\nTelephoto Array","url":"https://www.dragonflytelescope.org/"}]}]},{"path_id":"path-telescope-elapsed-time","seq":5,"link":"Fabrication","blocker":"Long-lead optics, cryogenic qualification, and single-source suppliers.","ai_type":"Physical build","maturity":"2-5 years","is_binding":0,"evidence":"JWST: construction of long-lead components begins 2004; all 18 primary mirror segments complete and cryogenically tested in 2011. Seven years.","rationale":"All three of Convergent's capabilities attached to this gap act here or at the next link.","duration_years":5,"duration_span":"2006 → 2011","duration_note":"flight instrument construction to all 18 mirror segments complete and cryo-tested","figure":null,"duration_days":null,"duration_span_note":null,"duration_covers_json":"[5]","ai_acts":0,"capabilities_json":"[\"Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes\",\"Space Telescope Factory\",\"Leveraging Commercial Component Advances\"]","capabilities":["Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes","Space Telescope Factory","Leveraging Commercial Component Advances"],"duration_covers":[5],"capability_links":[{"name":"Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes","url":"https://www.gap-map.org/capabilities/leveraging-modular-assembly-reduced-launch-costs-for-space-telescopes/","initiatives":[]},{"name":"Space Telescope Factory","url":"https://www.gap-map.org/capabilities/space-telescope-factory/","initiatives":[]},{"name":"Leveraging Commercial Component Advances","url":"https://www.gap-map.org/capabilities/leveraging-commercial-component-advances/","initiatives":[{"title":"Argus Array","url":"https://evryscope.astro.unc.edu/"},{"title":"Morning Star Missions to Venus","url":"https://venuscloudlife.com/"},{"title":"The Dragonfly\nTelephoto Array","url":"https://www.dragonflytelescope.org/"}]}]},{"path_id":"path-telescope-elapsed-time","seq":6,"link":"Integration and test","blocker":"Serial single-string assembly on hardware that cannot be duplicated, so little can be parallelised and every anomaly stops the line — but not only that. Gate C established that a substantial share of this block is funding latency and defect rework rather than assembly: the 2011 replan, after the House Appropriations Committee recommended termination, rebaselined the project and moved launch by 52 months, and the 2018 Independent Review Board followed sunshield tears and loose fasteners attributed to workmanship error.","ai_type":"Physical build","maturity":"Speculative","is_binding":0,"evidence":"JWST: mirror segments mounted into the backplane 2015-2016, optical and spacecraft elements mated 2019, final environmental testing complete 2020. Nine years from mirror completion. GAO-13-4 (3 December 2012), p.4: 'the JWST project was reauthorized, but not before it was recommended for termination by the House Appropriations Committee... NASA announced that the project would be rebaselined at $8.835 billion — a 78 percent increase — and would launch in October 2018 — a delay of 52 months.' The interval endpoint is confirmed: STScI's 'final integrated testing completed, 1 December 2017' is the element-level OTIS cryo-vacuum test, while full-observatory environmental testing completed 6 October 2020.","rationale":"The single longest link in the JWST chain, and the one where the capability set is most clearly pointed at the right place — this is what a space telescope factory is aimed at. Gate C contested the blocker and the contest succeeds in part: roughly 4.3 years of this nine-year block are a near-termination and replan, and a further tranche is workmanship rework. On this chain's own taxonomy a replan is Coordination and institutions and defect detection is Measurement and sensing, neither of which is physical assembly. ai_acts is left at 0 because no current capability shortens a congressional replan either, but the stated reason for the zero was wrong and is corrected here. It also exposes a limit of the model: a strictly sequential eight-link chain cannot represent funding delay recurring mid-build, and this chain places funding once, early.","duration_years":9,"duration_span":"2011 → 2020","duration_note":"segments complete to final environmental testing complete","figure":null,"duration_days":null,"duration_span_note":null,"duration_covers_json":"[6]","ai_acts":0,"capabilities_json":"[\"Space Telescope Factory\",\"Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes\"]","capabilities":["Space Telescope Factory","Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes"],"duration_covers":[6],"capability_links":[{"name":"Space Telescope Factory","url":"https://www.gap-map.org/capabilities/space-telescope-factory/","initiatives":[]},{"name":"Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes","url":"https://www.gap-map.org/capabilities/leveraging-modular-assembly-reduced-launch-costs-for-space-telescopes/","initiatives":[]}]},{"path_id":"path-telescope-elapsed-time","seq":7,"link":"Launch","blocker":"Vehicle availability and launch window.","ai_type":"Physical build","maturity":"Speculative","is_binding":0,"evidence":"JWST: final environmental testing complete 2020, launch 25 December 2021. Roughly one year. Ariane 5 selected as launch vehicle in 2005.","rationale":"Not where the time is. Reduced launch cost, which two of the three capabilities attached to this gap depend on, acts on the cost axis and not on this one. No AI capability shortens a launch campaign.","duration_years":1,"duration_span":"2020 → 2021","duration_note":"testing complete to launch on 25 December 2021","figure":null,"duration_days":null,"duration_span_note":null,"duration_covers_json":"[7]","ai_acts":0,"capabilities_json":"[\"Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes\"]","capabilities":["Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes"],"duration_covers":[7],"capability_links":[{"name":"Leveraging Modular Assembly & Reduced Launch Costs for Space Telescopes","url":"https://www.gap-map.org/capabilities/leveraging-modular-assembly-reduced-launch-costs-for-space-telescopes/","initiatives":[]}]},{"path_id":"path-telescope-elapsed-time","seq":8,"link":"Commissioning","blocker":"On-orbit alignment and calibration of a segmented optic.","ai_type":"Measurement and sensing","maturity":"Working now","is_binding":0,"evidence":"JWST: launch 25 December 2021, first full-colour images and start of science operations 12 July 2022. Six and a half months.","rationale":"Not binding, and the shortest link in the chain by an order of magnitude. Wavefront sensing on eighteen segments is the most technically demanding cognitive task in the whole sequence and it took under seven months, which is the chain in miniature.","duration_years":0.5,"duration_span":"Dec 2021 → Jul 2022","duration_note":"launch to first full-colour images and start of science operations","figure":null,"duration_days":null,"duration_span_note":null,"duration_covers_json":"[8]","ai_acts":1,"capabilities_json":"[]","capabilities":[],"duration_covers":[8],"capability_links":[]}]}],"audits":[{"gap_id":"1c1cb37e-2a00-800c-aafb-ec36c16025a8","dimension":"measurability","original":"Verification contested","audit":"Directly measurable","agreed":0,"adjudicated":null,"auditor_note":"test: Does the field agree in advance on what observation would settle the question? Yes: zero resistance plus Meissner flux expulsion plus an independently reproduced isotope effect at a stated T and P is an accepted verdict criterion, and Tc and critical pressure are observables with an unambiguous direction of improvement. The disputes (retracted CSH work, background-subtraction fights) are about data integrity and diamond-anvil-cell sample quality, not about whether the criterion decides the case. | nearest alternative: Verification contested | Close call. Tier 3 is tempting because the literature is genuinely contested, but tier 3 requires that the field would still disagree after a clean measurement. Here a clean, reproduced Meissner measurement would end the argument, which is why fraud detection rather than verifier ambiguity is what Convergent flags. Rejected tier 3 on that ground.","gap_name":"Uncertainty and Noise in the Science of Room-Temperature Superconductivity"},{"gap_id":"1c1cb37e-2a00-800c-aafb-ec36c16025a8","dimension":"ai_type","original":"Coordination and institutions","audit":"Prediction and modeling","agreed":0,"adjudicated":null,"auditor_note":"test: Is the binding constraint generating candidate compositions, or evaluating them accurately enough to spend a scarce megabar experiment? Evaluating. Chemical intuition already supplies more superhydride candidates than can be tested; what decides which of them are worth a diamond-anvil run is predicted electron-phonon coupling and Tc, which is now done with ML interatomic potentials and DFT-based screening at throughput no ab initio pipeline could reach. The hydride predictions that were later confirmed experimentally came from this route. | nearest alternative: Coordination and institutions | maturity: Working now | Convergent's only linked capability, 'Conduct Rigorous, Non-Fraudulent Experiments', is an incentive and replication problem, which argues Coordination. Rejected because AI has no traction on fraud incentives, whereas prediction directly answers the gap's own stated question (whether metallic hydrogen superconducts at reachable pressures). Design search was the third candidate, rejected for the same reason as Coordination: proposal is not the bottleneck.","gap_name":"Uncertainty and Noise in the Science of Room-Temperature Superconductivity"},{"gap_id":"1c1cb37e-2a00-800c-aafb-ec36c16025a8","dimension":"outcome","original":"A superconductivity claim can be settled in weeks. The current alternative is years of argument and a retraction.","audit":"It becomes knowable which hydrogen-rich compositions, if any, superconduct near room temperature at pressures a device could actually sustain, turning a contested and fragmented literature into a screenable design space.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Stated as what becomes knowable (which compositions, at what pressures) rather than the activity (running better experiments).","gap_name":"Uncertainty and Noise in the Science of Room-Temperature Superconductivity"},{"gap_id":"1c1cb37e-2a00-804b-bc6b-de7f4bc3adce","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Could a measurement in principle be taken with the field still disagreeing about what it showed? Yes. Critical-slowing-down early-warning indicators (rising variance, rising lag-1 autocorrelation, spatial skewness) are actively computed on real ecological time series, and there is live literature arguing they produce false positives and miss real transitions in noisy, spatially extended systems. The gap statement itself concedes the second half: we do not know 'what metrics to use to determine that restoration is effective'. | nearest alternative: Proxy only | Proxy only would say species counts and biomass are measurable while the tipping point is not. Rejected because candidate discriminating observables are not merely absent, they exist and are being built and argued over, which is the tier 3 wording. Counterfactual required was considered on 'prevent' (a collapse that did not happen) and rejected: regime shifts are observed events (fishery collapses, lake eutrophication, reef bleaching), so an observation set exists.","gap_name":"Inability to Anticipate or Prevent Ecosystem Tipping Points"},{"gap_id":"1c1cb37e-2a00-804b-bc6b-de7f4bc3adce","dimension":"ai_type","original":"Prediction and modeling","audit":"Prediction and modeling","agreed":1,"adjudicated":null,"auditor_note":"test: Is the blocker seeing the system or modelling its dynamics? The gap names 'the underlying dynamics, feedback loops, and thresholds' as the object, and anticipation requires a model that maps current state to distance-from-threshold. Detection alone cannot do this: you can count every bee and still not know where the pollination network breaks. Maturity is 2-5 years because learned ecosystem dynamics models validated against observed regime shifts do not yet exist at the scale or with the skill the gap requires, unlike the detection layer. | nearest alternative: Measurement and sensing | maturity: 2-5 years | Two of Convergent's three capabilities are monitoring, so Sensing is the obvious pick. Rejected because bioacoustic, eDNA and remote-sensing classification are largely working now and feed the anticipation problem rather than constituting it. This is a defensible disagreement point.","gap_name":"Inability to Anticipate or Prevent Ecosystem Tipping Points"},{"gap_id":"1c1cb37e-2a00-804b-bc6b-de7f4bc3adce","dimension":"outcome","original":"An ecosystem under stress gives a usable warning before it collapses, and management gets to act first.","audit":"Which ecosystems are approaching an irreversible transition becomes knowable early enough to intervene, and whether a restoration actually worked becomes checkable against a threshold rather than asserted.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Covers both halves of the gap statement (anticipation and restoration verification) in one sentence.","gap_name":"Inability to Anticipate or Prevent Ecosystem Tipping Points"},{"gap_id":"1c1cb37e-2a00-8072-a36f-f7ea8adace81","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Could a measurement be taken with the field still disagreeing about what it showed? This is the paradigm case. Reasoning and planning benchmarks exist and are run constantly, and every strong result is immediately followed by an argument over construct validity: contamination, whether the benchmark measures generality or pattern coverage, whether held-out abstraction tests probe cognition. The dispute survives the measurement, which is the tier 3 test. | nearest alternative: Directly measurable | Directly measurable is superficially right: benchmark scores are numbers with a direction of improvement. Rejected because the contested thing is whether the score is about the construct at all, not whether the score is accurate. Distinguish this from gap 1c1cb37e-2a00-8012 (theorem proving), where a proof checker makes the verdict machine-decidable and the tier drops to 1.","gap_name":"AI is Still Narrow in its Reasoning and Planning"},{"gap_id":"1c1cb37e-2a00-8072-a36f-f7ea8adace81","dimension":"ai_type","original":"Reading and synthesis","audit":"Reading and synthesis","agreed":1,"adjudicated":null,"auditor_note":"test: The gap's object is AI reasoning and planning itself, so the type is fixed by subject matter; the real call is maturity. Would this be closed on a 2-5 year horizon? The gap defines closure as breadth comparable to human cognition, and every route Convergent lists (mimic human evolution in silico, developmentally realistic training, holistic brain-inspired architectures, Bayesian cognitive architectures) is an unvalidated research programme rather than a scaling curve. Incremental breadth is shipping now, but 'no longer narrow compared to human cognition' has no demonstrated path, so Speculative. | nearest alternative: 2-5 years (same type) | maturity: Speculative | Taxonomy fit issue: this gap is reflexive. Every other gap asks which AI would move a science problem; here the science problem is AI. The type table has no entry for 'AI capability research itself', and Reading and synthesis is the least-bad container rather than a clean match. Flagging rather than inventing a value.","gap_name":"AI is Still Narrow in its Reasoning and Planning"},{"gap_id":"1c1cb37e-2a00-8072-a36f-f7ea8adace81","dimension":"outcome","original":"An AI system handles a domain it was not trained on without someone rebuilding the scaffolding around it.","audit":"It becomes knowable whether general planning and reasoning emerge from scaling current training or require architecturally distinct, developmentally grounded systems, and if the latter, such systems become buildable.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Framed as a disjunction that gets resolved, because the gap's own uncertainty is about which route works, not only about the endpoint.","gap_name":"AI is Still Narrow in its Reasoning and Planning"},{"gap_id":"1c1cb37e-2a00-80a7-ad9a-c209405e462d","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Applying the taxonomy's explicit tiebreaker: could a measurement in principle be taken, with people still disagreeing about what it showed? Yes, and it is taken routinely. Scheming, sandbagging and deceptive-alignment evaluations, and interpretability probes for internal goal representations, are run today, and the field argues hard about whether a passing or failing result tells you anything about a deployed system's behaviour under distribution shift. The instrument exists; the verdict is what is contested. | nearest alternative: Counterfactual required | This is the hardest tier call in the batch and I expect disagreement here. The tier 4 reading is strong: catastrophic loss of control is an event that has not occurred, so safety-in-deployment has no observation set, exactly like the calibration example of research that never happened. I went tier 3 only because the taxonomy instructs that when torn, the existence of a takeable-in-principle measurement decides it, and eval suites plus mechanistic interpretability are precisely that. If the adjudicator prefers tier 4 here I would not argue hard, and would accept a downgrade to guess.","gap_name":"AI Could Go Rogue"},{"gap_id":"1c1cb37e-2a00-80a7-ad9a-c209405e462d","dimension":"ai_type","original":"Reading and synthesis","audit":"Reading and synthesis","agreed":1,"adjudicated":null,"auditor_note":"test: Which capability, if it arrived, would let you tell a safe system from an unsafe one? Six of Convergent's eight sub-capabilities (automated interpretability, guaranteed safe architectures, understanding AI psychology, data-poisoning mitigation, non-agentic AI scientists) are AI systems doing reasoning, proof and code at a scale humans cannot match over model internals and specifications. That is the Reading and synthesis row, applied reflexively. Maturity 2-5 years: automated interpretability agents already produce usable feature explanations, while formally guaranteed safe architectures do not exist at frontier scale. | nearest alternative: Coordination and institutions | maturity: 2-5 years | Coordination has a real case: hardware governance is on Convergent's list, and even perfect interpretability closes nothing if no institution requires its use. Rejected because the R&D gap as stated is the inability to see inside and to guarantee, which is a technique deficit; the governance layer is the second-order blocker. Reasonable people land the other way.","gap_name":"AI Could Go Rogue"},{"gap_id":"1c1cb37e-2a00-80a7-ad9a-c209405e462d","dimension":"outcome","original":"You can check a claim about an AI system against its internals. Today the evidence is behavior on tests the system may recognize.","audit":"Whether a given model harbours goals or dispositions its developers did not intend becomes checkable before deployment, rather than inferable only after a failure.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Deliberately an epistemic outcome (what becomes checkable) rather than 'AI becomes safe', which would restate the gap.","gap_name":"AI Could Go Rogue"},{"gap_id":"1b4cb37e-2a00-807c-88c6-d543174bdbd2","dimension":"measurability","original":"Proxy only","audit":"Verification contested","agreed":0,"adjudicated":null,"auditor_note":"test: Could a measurement be taken with the field still disagreeing? Yes. Candidate observables for 'is this model system representative' exist and are being built: predictivity of held-out human neural recordings, representational similarity to human cortex, behavioural match in embodied testbeds. The field does not agree that hitting them establishes validity, because a system can match recorded responses through a different mechanism, and the organoid-versus-human-cortex argument is exactly a dispute about what a match would prove. | nearest alternative: Proxy only | Proxy only would say we can measure adjacent things (expression similarity, firing statistics) but not representativeness. Rejected on the tier 3 wording 'a candidate observable exists or is being built': digital replicas are being benchmarked against human data right now, so the deficit is verdict agreement, not instrument availability.","gap_name":"Current “Model Systems” for Brain Function are Not Representative of the Real Human Brain"},{"gap_id":"1b4cb37e-2a00-807c-88c6-d543174bdbd2","dimension":"ai_type","original":"Prediction and modeling","audit":"Prediction and modeling","agreed":1,"adjudicated":null,"auditor_note":"test: Is the model system that closes this gap a learned emulator or a physical construct? Four of Convergent's seven capabilities (data-driven AI models of brain systems, digital replicas of a whole model-organism brain, embodied testbeds, behaviour-space datasets) are learned emulators standing in for an inaccessible human brain, which is the ML surrogates row verbatim. The wet capabilities (ex vivo human brains, in vitro cortex) are biology where AI is supporting rather than substantive. Maturity 2-5 years: whole-brain digital replicas of a model organism are in progress, not delivered. | nearest alternative: Running experiments | maturity: 2-5 years | Running experiments would fit if organoid and in vitro cortical model development were the primary route, since that is bench-scale closed-loop work. Rejected because the gap's framing (digital reconstructions, embodied simulations) puts the emulator half first, and because a self-driving lab does not solve 'this model is not human'.","gap_name":"Current “Model Systems” for Brain Function are Not Representative of the Real Human Brain"},{"gap_id":"1b4cb37e-2a00-807c-88c6-d543174bdbd2","dimension":"outcome","original":"Results from model systems survive the jump to the human brain.","audit":"Human-specific brain mechanisms that no rodent or dish currently reproduces become experimentally accessible, so results stop being stranded in the model organism and become testable claims about human brains.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | The outcome is transferability of findings, which is what a non-representative model system destroys.","gap_name":"Current “Model Systems” for Brain Function are Not Representative of the Real Human Brain"},{"gap_id":"1c1cb37e-2a00-801b-bbfd-c904de35f2b4","dimension":"measurability","original":"Proxy only","audit":"Proxy only","agreed":1,"adjudicated":null,"auditor_note":"test: Is there a specific measurement the field would accept as decisive, or only an accumulation of indicators? For quantum gravity, the tier 3 calibration case, the field can name the experiment it is arguing about. Here it cannot: there is no observation whose result would settle whether a polity's collective decision-making is better informed. What can be measured is adoption, turnout, participant counts, cryptographic audit properties and post-deliberation opinion shift, all of which are inputs or adjacent effects. That is the Proxy only definition. | nearest alternative: Verification contested | The tier 3 case is that deliberative-poll instruments exist and their construct validity is disputed. Rejected because the dispute is not organised around a candidate discriminator anyone expects to be decisive; it is the ordinary condition of a construct with no agreed operationalisation. I flag that this is the second-hardest tier call in the batch, and that a labeler using the literal 'could a measurement be taken' tiebreaker would land on tier 3.","gap_name":"Our Platforms for Civic Engagement and Democratic Decision-Making Don’t Take Advantage of 21st Century Scalable Technology"},{"gap_id":"1c1cb37e-2a00-801b-bbfd-c904de35f2b4","dimension":"ai_type","original":"Coordination and institutions","audit":"Reading and synthesis","agreed":0,"adjudicated":null,"auditor_note":"test: What does AI actually contribute here, as opposed to what the gap needs overall? The gap needs cryptography for voting (not AI at all) and language understanding for deliberation. The deliberation half is eliciting, clustering and fairly summarising the positions of a large population and finding statements that bridge factions, which is language synthesis. Working now: consensus-statement generation and opinion-clustering deployments already run at population scale, so the constraint is deployment rather than capability. | nearest alternative: Coordination and institutions | maturity: Working now | Coordination is arguable, since the actual blocker is that election authorities and legislatures do not adopt these tools. Rejected because Coordination in this taxonomy covers AI applied to allocation and review, and here AI is doing linguistic work while the institution is the passive adopter. Taxonomy fit note: 'Improved Voting and Auditing Protocols' is cryptography with no AI content, so one of the two capabilities under this gap is outside the taxonomy's frame entirely.","gap_name":"Our Platforms for Civic Engagement and Democratic Decision-Making Don’t Take Advantage of 21st Century Scalable Technology"},{"gap_id":"1c1cb37e-2a00-801b-bbfd-c904de35f2b4","dimension":"outcome","original":"Deliberation stops being limited to the number of people who fit in a room. Public preferences get elicited and aggregated at population scale, verifiably.","audit":"What a large public actually thinks after being informed and hearing the strongest objections becomes elicitable at population scale, rather than approximated by polling a sample on unconsidered opinions.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Names the thing that becomes knowable (considered population-scale preferences) rather than the activity (better civic platforms).","gap_name":"Our Platforms for Civic Engagement and Democratic Decision-Making Don’t Take Advantage of 21st Century Scalable Technology"},{"gap_id":"1c1cb37e-2a00-80b5-8594-d53ea33db290","dimension":"measurability","original":"Proxy only","audit":"Proxy only","agreed":1,"adjudicated":null,"auditor_note":"test: Is there an observable that moves when and only when the gap closes? No. Supplier concentration, stockpile days, single-points-of-failure counts and drill recovery times all move without robustness necessarily improving, because reducing one dependency commonly relocates the fragility rather than removing it. Those quantities are structural inputs to robustness; robustness itself shows up only under severe shocks, which are rare and not chosen by the analyst. | nearest alternative: Directly measurable | Directly measurable is defensible if you accept concentration indices as the gap's own quantity. Rejected on the input-versus-outcome test above. Counterfactual required was also weighed and rejected: real disruptions (pandemic, semiconductor shortage, grid failures) supply an observation set, so the gap is not about something that never happened. The exception is Convergent's 'Civilization Reboot Toolkit', which taken alone is genuinely tier 4 since no observation set for a civilizational restart exists; the gap as written is broader than that component.","gap_name":"Fragile Supply Chains and Lack of Backup for Critical Infrastructure"},{"gap_id":"1c1cb37e-2a00-80b5-8594-d53ea33db290","dimension":"ai_type","original":"Coordination and institutions","audit":"Coordination and institutions","agreed":1,"adjudicated":null,"auditor_note":"test: If a perfect robust-supply-chain optimizer existed tomorrow, would supply chains become robust? No, because each firm optimizes its own cost and no actor owns system-level resilience; the optimum the optimizer finds is not the one anyone is paid to implement. Run it the other way: with today's tools plus strategic reserves, mandated redundancy and shared visibility requirements, fragility falls immediately. The blocker is an organisation rather than a technique, which is the Coordination row verbatim. | nearest alternative: Design search | maturity: 2-5 years | Convergent's named capability is literally 'Optimize Circular and Robust Supply Chains', so Design search is the obvious pick and I expect the original labeler took it. I rejected it under the taxonomy's own instruction that the primary is the type that would move the gap most, not the one most obviously applicable. Maturity 2-5 years: n-tier supplier graph mapping from trade and text data is working now, but AI-assisted allocation and reserve policy is not.","gap_name":"Fragile Supply Chains and Lack of Backup for Critical Infrastructure"},{"gap_id":"1c1cb37e-2a00-80b5-8594-d53ea33db290","dimension":"outcome","original":"A shock to one node stops cascading into lost industrial capability, because the failure points are mapped and the fallbacks are already in place.","audit":"Which local failures propagate into loss of a critical societal function becomes knowable before the failure, and substitutable, restartable production paths for critical goods become buildable in advance.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Two-clause outcome covering both the fragility-mapping and the backup-mechanism halves of the gap.","gap_name":"Fragile Supply Chains and Lack of Backup for Critical Infrastructure"},{"gap_id":"1b3cb37e-2a00-80cd-9213-c8eb788cc657","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Is there an observable quantity with a direction of improvement that is the gap itself? Yes, several, and they are unambiguous: cubic millimetres reconstructed at synapse resolution, cost and imaging time per unit volume, segmentation and proofreading error rate, fraction of cells with molecular annotation. Nobody disputes what a completed wiring diagram is or how to tell a better one from a worse one; the gap is cost and throughput, which are numbers. | nearest alternative: Proxy only | Proxy only rejected because reconstructed volume and cost per volume are the gap's own quantities, not adjacent effects. Note the contrast with the model-systems gap in the same field: 'is this map complete and accurate' is decidable, 'is this model representative of a human brain' is not, and that is why the two neuroscience gaps land in different tiers.","gap_name":"Most Brain Circuitry is Still Invisible"},{"gap_id":"1b3cb37e-2a00-80cd-9213-c8eb788cc657","dimension":"ai_type","original":"Measurement and sensing","audit":"Measurement and sensing","agreed":1,"adjudicated":null,"auditor_note":"test: Where in the pipeline does AI change the achievable scale? Both places, and both are the same row. Turning terabyte-to-petabyte electron microscopy volumes into a segmented, synapse-annotated graph is detection and reconstruction from instrument data, and the route to faster microscopy is largely reconstruction from sparser, lower-dose acquisition rather than a faster stage. Working now: automated segmentation carried a whole-brain insect connectome to completion, with humans proofreading rather than tracing. | nearest alternative: Physical build | maturity: Working now | Physical build is arguable via 'Faster Electron Microscopy for Connectomics' as an instrument-engineering problem (automated sectioning, multibeam stages). Rejected because that row is about robotics for fabrication and field deployment, and because the scaling gains AI supplies here are computational reconstruction, not manipulation. Deliberately kept distinct from Running experiments, which does not apply: nothing here chooses its own experiments.","gap_name":"Most Brain Circuitry is Still Invisible"},{"gap_id":"1b3cb37e-2a00-80cd-9213-c8eb788cc657","dimension":"outcome","original":"Complete wiring diagrams with molecular annotation, affordable at brain scale. Circuit-level explanations of behavior and disorder become testable.","audit":"Synapse-resolution wiring diagrams with molecular cell identity become obtainable for mammalian brains at a price a normal lab can pay, making circuit-level explanations of behaviour and circuit-level signatures of brain disorders testable rather than hypothesised.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | The outcome is that circuit-level claims become testable, which is what invisibility currently prevents.","gap_name":"Most Brain Circuitry is Still Invisible"},{"gap_id":"1c1cb37e-2a00-8012-aeb4-c0c2e9c20f46","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Is the verdict machine-decidable? Yes, uniquely so in this batch. A proof assistant kernel accepts or rejects; there is no version of this field in which a verified proof is disputed, and progress is counted in benchmark pass rates, formalized theorem counts and open conjectures closed. This is the cleanest tier 1 in the batch and the sharpest contrast with the other Computation gaps, where the measurement exists but the verdict does not follow from it. | nearest alternative: Verification contested | Verification contested would be the reflex for anything about AI reasoning, and it is exactly wrong here: formal verification is what removes the contest. Interesting that the reasoning gap and the theorem-proving gap sit in the same field and different tiers purely because one has a checker.","gap_name":"Proving Math Theorems is Challenging for Both Humans and AI"},{"gap_id":"1c1cb37e-2a00-8012-aeb4-c0c2e9c20f46","dimension":"ai_type","original":"Reading and synthesis","audit":"Reading and synthesis","agreed":1,"adjudicated":null,"auditor_note":"test: What is the new ingredient, given that automated proof search has existed for decades? Tactic search and SMT solving are old and did not close this gap; what changed is language-model generation of candidate lemmas, informal-to-formal statement translation, and reinforcement learning against prover feedback. That is the reasoning-and-synthesis row, and the taxonomy names mathematical reasoning in it explicitly. Working now: systems already reach competition-olympiad level with machine-checked proofs. | nearest alternative: Design search | maturity: Working now | Design search fits the shape of proof search over a defined space, and Convergent's capability name ('RL from Interactive Theorem Prover Feedback') has an optimization flavour. Rejected because the search machinery was never the missing piece; the proposal distribution was.","gap_name":"Proving Math Theorems is Challenging for Both Humans and AI"},{"gap_id":"1c1cb37e-2a00-8012-aeb4-c0c2e9c20f46","dimension":"outcome","original":"Formal verification stops costing more than the proof it verifies, and machine-assisted proof turns routine at research level.","audit":"Open conjectures become attackable at scale with machine-checked certainty, and large results become verifiable without anyone trusting a referee's reading of a hundred-page argument.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Second clause is the underrated outcome: formalization changes who has to be trusted, not just how fast proofs arrive.","gap_name":"Proving Math Theorems is Challenging for Both Humans and AI"},{"gap_id":"1c1cb37e-2a00-803c-a293-e7aa6da74419","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Does an observable with a direction of improvement exist for the gap itself? Yes: kcat, kcat/KM, success rate of designs that fold and function, thermostability under the target extreme condition, and incorporation efficiency of non-canonical residues. A designed enzyme is assayed and either turns over substrate or does not; no one disputes what the assay showed. | nearest alternative: Proxy only | Proxy only rejected: in vitro activity is the quantity of interest, not an adjacent effect. The 'novelty' half of the gap (is this design truly non-biomimetic?) is fuzzier, but it is not what determines whether the gap has closed.","gap_name":"Protein Design Has Been Limited to Static, Bio-mimetic Structures"},{"gap_id":"1c1cb37e-2a00-803c-a293-e7aa6da74419","dimension":"ai_type","original":"Design search","audit":"Design search","agreed":1,"adjudicated":null,"auditor_note":"test: Is the blocker predicting properties of given candidates, or generating candidates outside the natural distribution? Generating. Structure prediction on natural-like sequences is solved well enough that it is no longer what limits novel function; the gap explicitly says designs are stuck being bio-mimetic, which is a statement about the proposal distribution, and inverse design methods (backbone generation conditioned on a catalytic geometry, sequence design onto that backbone) are what move it. Maturity 2-5 years, not Working now: de novo binders and static scaffolds work today, but long-range dynamic and allosteric mechanisms have not been recapitulated de novo, and that is precisely the 'static' limitation this gap names. | nearest alternative: Prediction and modeling | maturity: 2-5 years | This is the pressure point the protocol predicts (design search versus surrogates), and the two are genuinely coupled because the design loop scores candidates with predictors. Split on which side the gap statement locates the failure: it says designs are too conservative, not that scoring is inaccurate, so the generator is primary. Maturity checked against current literature rather than assumed.","gap_name":"Protein Design Has Been Limited to Static, Bio-mimetic Structures"},{"gap_id":"1c1cb37e-2a00-803c-a293-e7aa6da74419","dimension":"outcome","original":"Enzymes get designed for jobs evolution never had to solve. Today protein engineering starts from something nature already built.","audit":"Enzymes with catalytic mechanisms, conformational dynamics and chemistries that no natural protein exhibits become buildable, opening reaction space that evolution never explored.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | 'Reaction space evolution never explored' is the buildable outcome; 'better protein design tools' would have been the activity.","gap_name":"Protein Design Has Been Limited to Static, Bio-mimetic Structures"},{"gap_id":"1c1cb37e-2a00-8066-824e-d4ec0d434c3d","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Is there an agreed observable with a direction of improvement? Yes, and fusion has unusually well-policed ones: energy confinement time, the triple product, Q, disruption rate per shot, and time sustained before an instability terminates the discharge. The field argues about which machine concept will get there, not about what the numbers mean when a shot ends. | nearest alternative: Proxy only | Proxy only rejected because confinement time and Q are the gap's own quantities, not adjacent ones. Note the contrast with the superconductivity gap in the same field: both are tier 1, and both have contested claims about specific results, which is a reminder that contested results are not the same thing as contested verification.","gap_name":"Robust and Compact Plasma Confinement for Fusion is Still Not Solved"},{"gap_id":"1c1cb37e-2a00-8066-824e-d4ec0d434c3d","dimension":"ai_type","original":"Prediction and modeling","audit":"Design search","agreed":0,"adjudicated":null,"auditor_note":"test: The gap asks for confinement that is both robust and compact, and compactness is set at design time by the magnetic geometry, not at run time by the controller. Optimized-stellarator and coil-set design is inverse design over a defined configuration space, and it is the route by which AI enlarges the set of confinement geometries that are buildable at all; actuator-trajectory planning for a discharge falls into the same row as 'experiment planning over a defined space'. Working now: both optimized-geometry design pipelines and learned magnetic control on real tokamaks have been demonstrated on hardware. | nearest alternative: Prediction and modeling | maturity: Working now | Fast learned MHD and turbulent-transport emulators are indispensable and could be argued as primary. Rejected because they serve the search loop; the leverage is in which configurations get searched, not in emulating one configuration faster. Taxonomy fit issue worth recording: real-time closed-loop control of a physical system has no home in the seven types. It is not Running experiments (the controller does not choose experiments) and not Physical build. I placed it under Design search as least-bad; a control row would be a defensible taxonomy revision.","gap_name":"Robust and Compact Plasma Confinement for Fusion is Still Not Solved"},{"gap_id":"1c1cb37e-2a00-8066-824e-d4ec0d434c3d","dimension":"outcome","original":"Plasma stays confined for as long as the reactor needs, not as long as the instability allows. That is the last physics obstacle before sustained operation.","audit":"Confinement geometries that hold a burning plasma stably at a size and cost a utility could build become identifiable before a device is constructed, making the engineering case for fusion decidable on hardware rather than in argument.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Compactness is the load-bearing word: a stable plasma in an arbitrarily large machine is not what closes this gap.","gap_name":"Robust and Compact Plasma Confinement for Fusion is Still Not Solved"},{"gap_id":"1c1cb37e-2a00-80b4-a5fa-cd97b7b1bed6","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Does an observable exist for each of the three named sub-problems? Yes: separation factor and energy per kilogram for acquiring and concentrating, recovery yield and purity for waste streams, and maximum energy product plus coercivity at operating temperature for substituting neodymium. Every one has an undisputed direction of improvement, and a separation either achieves the factor or does not. | nearest alternative: Proxy only | Proxy only rejected: these are the gap's own quantities. The economic framing ('cost of materials is dominated by cost to obtain elements') might suggest cost is a proxy, but cost here is a direct readout of separation energy and yield.","gap_name":"We Have a Limited Ability to Acquire, Concentrate and Substitute Chemical Elements in Processes"},{"gap_id":"1c1cb37e-2a00-80b4-a5fa-cd97b7b1bed6","dimension":"ai_type","original":"Design search","audit":"Running experiments","agreed":0,"adjudicated":null,"auditor_note":"test: Is the shortage candidates or measured results? Candidates. Generative materials screening already produces far more predicted compositions and extractant ligands than anyone can test, and predicted selectivity for lanthanide separation is unreliable because solvation thermodynamics between adjacent f-block elements are nearly degenerate, so the loop only closes by measuring. The required work is bench-scale and in a well-defined parameter space (pH, ligand equivalents, solvent, temperature), which is exactly where closed-loop platforms operate today and where high-throughput rare-earth precipitation screens already separate Nd/Dy and La/Nd at lab scale. | nearest alternative: Design search | maturity: Working now | Explicitly checked against Physical build, the pair the protocol warns about, and it is not close: this is liquid handling and small-sample synthesis on a bench, not fabrication, assembly or field deployment. Design search is the real nearest alternative and was rejected on the candidates-versus-measurement test above. Maturity Working now is qualified: most deployed platforms are supervised rather than fully autonomous, so 'Working now' means the bench-scale narrow-domain sense the taxonomy itself specifies. Separately, industrial scale-up of a separation process is plant engineering that no row in the taxonomy covers.","gap_name":"We Have a Limited Ability to Acquire, Concentrate and Substitute Chemical Elements in Processes"},{"gap_id":"1c1cb37e-2a00-80b4-a5fa-cd97b7b1bed6","dimension":"outcome","original":"Separation gets cheap and constrained elements get designed around. What a material costs stops being set by what its elements cost.","audit":"Which functions genuinely require a scarce element and which can be delivered by an abundant substitute becomes knowable, turning element supply from a geological constraint into a design variable.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Captures the gap's own claim that the critical-minerals framing masks a scientific question about substitution.","gap_name":"We Have a Limited Ability to Acquire, Concentrate and Substitute Chemical Elements in Processes"},{"gap_id":"1c1cb37e-2a00-801a-bbfd-cd37c0419eca","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Would a well-controlled null result close the question for the people asking it? No — the gap statement itself says that despite decades of measurement and no good evidence, 'other mechanisms or parameter combinations' remain, which is exactly the structure of a question no measurement settles. Calorimetric excess-heat and He-4 observables exist and have been taken repeatedly; what is disputed is whether any of them (positive or null) is dispositive. | nearest alternative: Directly measurable | Genuinely hard call. Excess heat per unit input and nuclear-ash yield are observable quantities with a direction, which is the tier-1 test verbatim. Rejected because the field's disagreement is not about instrument precision but about whether a measurement in a given parameter regime licenses any general verdict. Not tier 4: an instrument can be pointed at a cell today.","gap_name":"Incomplete Resolution of the Possibility of Low-Energy Nuclear Reactions"},{"gap_id":"1c1cb37e-2a00-801a-bbfd-cd37c0419eca","dimension":"ai_type","original":"Running experiments","audit":"Running experiments","agreed":1,"adjudicated":null,"auditor_note":"test: Is the binding constraint proposing which conditions to try, or physically running enough well-controlled trials to cover the space? The gap says the space is underexplored, not that nobody knows what to try — so the constraint is closed-loop throughput of loading/electrochemistry/calorimetry runs at bench scale, which is the autonomous-experimentation category. | nearest alternative: Design search | maturity: 2-5 years | Design search would only help if candidate generation were the bottleneck; here candidate conditions are cheap to name and expensive to test. Not Physical build: this is bench-scale apparatus in a lab, not fabrication or field deployment. Maturity 2-5 years — self-driving labs work now in chemistry and formulation, but no closed-loop rig currently couples the calorimetry and materials handling this needs at the required control standard.","gap_name":"Incomplete Resolution of the Possibility of Low-Energy Nuclear Reactions"},{"gap_id":"1c1cb37e-2a00-801a-bbfd-cd37c0419eca","dimension":"outcome","original":"The question closes, either way, and stops absorbing attention.","audit":"Whether anomalous low-energy nuclear effects exist anywhere in the accessible materials-and-loading space becomes a bounded, closed question rather than one that survives every null result.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Outcome is exclusion, not discovery — the informative result here is most likely a defensible negative, which is still something that becomes knowable.","gap_name":"Incomplete Resolution of the Possibility of Low-Energy Nuclear Reactions"},{"gap_id":"1c1cb37e-2a00-804f-bf69-f232177933ce","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Does a candidate discriminator between living and nonliving already exist and get argued over? Yes — assembly-index measurements from mass spectrometry, entropy-production and dissipation measures, and information-theoretic criteria are all published as candidate criteria, and there is live literature denying that any of them separates life from abiotic chemistry. A measurement can be taken; the field will not agree what it showed. | nearest alternative: Proxy only | Proxy only would say dissipation rates are adjacent effects and 'aliveness' itself is unmeasurable. Rejected because the candidate quantities are advanced as direct criteria, not as proxies, and the dispute is over their verdict rather than their relevance.","gap_name":"Understanding Life as a Far-From-Equilibrium Physical Phenomenon"},{"gap_id":"1c1cb37e-2a00-804f-bf69-f232177933ce","dimension":"ai_type","original":"Prediction and modeling","audit":"Prediction and modeling","agreed":1,"adjudicated":null,"auditor_note":"test: What stands between the current state and a usable formal framework — generating candidate definitions, or being able to compute the behaviour of a nonequilibrium many-body system large enough for a candidate definition to be tested on it? The latter: candidate frameworks already outrun the systems we can simulate, and learned emulators of nonequilibrium dynamics are what change that. | nearest alternative: Reading and synthesis | maturity: 2-5 years | LLM reasoning is the tempting pick because the gap names 'formal frameworks', i.e. theory. Rejected because proposing formalisms is not the scarce step; adjudicating them against simulated dynamics at biological scale is. Maturity 2-5 years: neural potentials and learned coarse-grained dynamics are working now at small scale, whole-organism nonequilibrium scale is not.","gap_name":"Understanding Life as a Far-From-Equilibrium Physical Phenomenon"},{"gap_id":"1c1cb37e-2a00-804f-bf69-f232177933ce","dimension":"outcome","original":"There is an operational line between living and nonliving matter, so origin-of-life research and biosignature detection get a testable target.","audit":"A physical criterion that separates living from nonliving matter becomes computable on real systems, turning 'is this alive' from a definitional argument into a measurement — including for candidate biosignatures with no shared ancestry with us.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Stated as what becomes decidable, not as 'we would understand life better'.","gap_name":"Understanding Life as a Far-From-Equilibrium Physical Phenomenon"},{"gap_id":"1c1cb37e-2a00-808e-bcac-eb5bcdc8f23c","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Could a measurement in principle be taken with the field still disagreeing about what it showed? Yes, and it repeatedly has been: artificial-life systems have existed for decades and the open-endedness literature has never agreed on a criterion by which any of them counts as complex evolved computation. There is an object to point at; the verdict is what is contested. | nearest alternative: Counterfactual required | Hard call, and the one most likely to be a coin flip. The counterfactual reading — 'the forms of computation evolution could have produced but did not' — is real, and with n=1 you cannot separate necessary from contingent features. Rejected because the listed capability is to build the second instance, and once built it is observable; tier 4 requires there to be nothing to point an instrument at.","gap_name":"Biological Life is Our Only Working Example of Complex Evolved Computation"},{"gap_id":"1c1cb37e-2a00-808e-bcac-eb5bcdc8f23c","dimension":"ai_type","original":"Prediction and modeling","audit":"Design search","agreed":0,"adjudicated":null,"auditor_note":"test: Is the missing thing a cheaper way to compute a known dynamics, or a way to find substrates and dynamics that generate open-ended complexity at all? The latter — the search over candidate computational substrates is the whole problem, which is generative proposal over a space rather than emulation of an expensive simulator. | nearest alternative: Prediction and modeling | maturity: Speculative | The capability name ('Learning and GPUs') pulls toward surrogates/large-scale learned simulation; rejected because throughput is not the reported blocker, the absence of any substrate showing open-ended evolution is. Maturity Speculative: no artificial system has produced complexity growth resembling biology, and nothing in the current trajectory dates that within five years. Strains the taxonomy slightly — the category says 'over a defined space', and open-ended search is by construction not over a defined space.","gap_name":"Biological Life is Our Only Working Example of Complex Evolved Computation"},{"gap_id":"1c1cb37e-2a00-808e-bcac-eb5bcdc8f23c","dimension":"outcome","original":"We get a second example of complex evolved computation to generalize from. At present biology is the only one.","audit":"A second, non-biological instance of evolved complex computation exists, so which properties of biological computation are universal and which are contingencies of carbon chemistry becomes separable for the first time.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | The outcome is escaping n=1, not 'new computing paradigms', which is downstream and speculative.","gap_name":"Biological Life is Our Only Working Example of Complex Evolved Computation"},{"gap_id":"1c1cb37e-2a00-80ce-ba28-fe5226876f66","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Could a measurement be taken with the field still disagreeing about what it showed? Yes — sub-scale field trials (marine cloud brightening plumes, glacier-scale interventions, stratospheric perturbation experiments) are being built and instrumented right now, and the entire dispute is whether a detectable sub-scale response tells you anything about planetary-scale deployment against natural variability. | nearest alternative: Counterfactual required | The hardest tier call in this batch. The quantity people actually care about — what the climate would have done with and without a planetary intervention — is a counterfactual with no observation set, which is a serious tier-4 argument. Rejected on the taxonomy's own instruction: an instrument can in principle be pointed at a sub-scale trial, and the disagreement is about extrapolation from it, which is the tier-3 signature. Low confidence; I would not defend this over tier 4 strongly.","gap_name":"Intervening in Earth Systems at Scale is Largely Untested"},{"gap_id":"1c1cb37e-2a00-80ce-ba28-fe5226876f66","dimension":"ai_type","original":"Physical build","audit":"Prediction and modeling","agreed":0,"adjudicated":null,"auditor_note":"test: If robotic deployment hardware appeared tomorrow, would the gap close? No — nobody could say what the intervention would do, and no one would authorise it. The scarce object is a fast enough emulator of the earth system to run the large ensembles that detection, attribution and side-effect bounding require; that is learned emulation of expensive simulation. | nearest alternative: Physical build | maturity: 2-5 years | Physical build is genuinely applicable (pumps on glaciers, dispersal fleets, ocean deployment) and I considered it primary. Rejected because that engineering is not AI-limited, while the evaluation problem is squarely AI-addressable. Maturity 2-5 years: learned weather emulators are operational now, climate-scale and intervention-scenario emulators with credible regional skill are not.","gap_name":"Intervening in Earth Systems at Scale is Largely Untested"},{"gap_id":"1c1cb37e-2a00-80ce-ba28-fe5226876f66","dimension":"outcome","original":"Emergency climate intervention is either a real option or ruled out, decided at meaningful scale and while nobody is under duress.","audit":"Emergency climate interventions become a costed, side-effect-bounded option set that can be evaluated before a tipping point is crossed, instead of a class of actions whose consequences are first learned by performing them.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Deliberately not 'we could stop glacier melt' — what becomes knowable is whether and at what cost, which is the decision-relevant object.","gap_name":"Intervening in Earth Systems at Scale is Largely Untested"},{"gap_id":"1c1cb37e-2a00-800b-8598-ea123934c801","dimension":"measurability","original":"Proxy only","audit":"Verification contested","agreed":0,"adjudicated":null,"auditor_note":"test: Do candidate observables exist that people measure and then argue about? Yes, continuously: jailbreak attack-success rates, dangerous-capability evaluations and red-team pass rates are reported for every frontier release, and the standing dispute is whether clearing them predicts anything about misuse in deployment. Measurement happens; agreement on the verdict does not. | nearest alternative: Counterfactual required | Tier 4 is arguable — the quantity of ultimate interest is harm that did not occur, which has no observation set. Rejected because the gap as written is about whether safeguards hold, and safeguard robustness is measured today under active dispute. Tier 1 also considered and rejected: attack-success rate has a direction but no one treats it as settling the question.","gap_name":"AI Could Be Misused"},{"gap_id":"1c1cb37e-2a00-800b-8598-ea123934c801","dimension":"ai_type","original":"Coordination and institutions","audit":"Reading and synthesis","agreed":0,"adjudicated":null,"auditor_note":"test: Is the named blocker a technique or an organisation? The two capabilities attached are jailbreak resistance and decentralized training — both are engineering programs, and jailbreak resistance in particular advances through models attacking and critiquing models (automated red-teaming, adversarial training, model-based monitoring), which is the LLM reasoning category applied reflexively. | nearest alternative: Coordination and institutions | maturity: Working now | Poorest taxonomy fit in the batch: the taxonomy is built for AI-as-instrument-for-science, and here AI is the object of study rather than the tool. Coordination and institutions is a defensible primary — access policy, liability and compute governance plausibly bind harder than technique — and I would accept it on adjudication. Maturity Working now: automated red-teaming and adversarial hardening are in production today, whatever their sufficiency.","gap_name":"AI Could Be Misused"},{"gap_id":"1c1cb37e-2a00-800b-8598-ea123934c801","dimension":"outcome","original":"Releasing a capable model stops meaning releasing its worst uses along with it.","audit":"Whether a model's safeguards actually hold against a determined adversary becomes checkable before release rather than discovered after it, and capable systems become trainable without concentrating misuse-relevant control in a handful of actors.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Covers both attached capabilities; the pre-deployment/post-deployment inversion is the thing that becomes knowable.","gap_name":"AI Could Be Misused"},{"gap_id":"1c1cb37e-2a00-8072-976d-fb7598c7401e","dimension":"measurability","original":"Proxy only","audit":"Directly measurable","agreed":0,"adjudicated":null,"auditor_note":"test: Is there an observation set for the thing that did not happen? Unusually, yes: decadal surveys enumerate the prioritized planetary and astrobiology missions by name, so unflown-versus-flown, cost per mission and time from prioritization to launch are all countable, with an unambiguous direction of improvement. | nearest alternative: Counterfactual required | Hard call, and structurally close to the taxonomy's own tier-4 calibration example about organizational structures constraining what R&D gets done. The difference that decided it: for organizational forms there is no list of the research that would have happened, whereas here the missions are named in advance by an existing prioritization process. Only the discoveries those missions would have made are counterfactual, and those are downstream of the gap rather than the gap.","gap_name":"Major Planetary Science and Astrobiology Missions Are Not Realized by Existing Government Space Agencies"},{"gap_id":"1c1cb37e-2a00-8072-976d-fb7598c7401e","dimension":"ai_type","original":"Coordination and institutions","audit":"Coordination and institutions","agreed":1,"adjudicated":null,"auditor_note":"test: Is the blocker a missing technique or a missing organisation? The gap names the blocker outright — existing government agencies do not realize these missions and 'alternative, independent initiatives are needed'. Nothing about the physics or the hardware is stated as unsolved; who decides, funds and carries risk is. | nearest alternative: Physical build | maturity: 2-5 years | Physical build is the tempting technical read — autonomous spacecraft and cheap fabrication lower the cost of an independent mission. Rejected because cost is not the stated blocker and launch costs have already fallen by an order of magnitude without the missions appearing. Maturity 2-5 years: AI for portfolio allocation and review exists but has not been shown to change who funds and flies a planetary mission.","gap_name":"Major Planetary Science and Astrobiology Missions Are Not Realized by Existing Government Space Agencies"},{"gap_id":"1c1cb37e-2a00-8072-976d-fb7598c7401e","dimension":"outcome","original":"Europa and Enceladus plume sampling fly on a schedule set by mission design. Today it is set by position in an agency queue.","audit":"High-priority astrobiology targets — ocean-world plumes, Venus cloud chemistry — become reachable on philanthropic budgets and timescales, so whether there is life elsewhere in this solar system stops being gated on a decadal queue.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Names what becomes answerable rather than 'more missions get flown', which would restate the gap.","gap_name":"Major Planetary Science and Astrobiology Missions Are Not Realized by Existing Government Space Agencies"},{"gap_id":"1c1cb37e-2a00-80e2-a028-ebaa685df3da","dimension":"measurability","original":"Proxy only","audit":"Verification contested","agreed":0,"adjudicated":null,"auditor_note":"test: Have measurements actually been taken, with the field then disagreeing about what they showed? Yes, at the sharpest possible scale: platform-scale randomized experiments on feed ranking and their effects on beliefs and polarization have been run and published, and the interpretation fight over them is unresolved. Candidate observables exist; agreement that they settle 'epistemic health' does not. | nearest alternative: Proxy only | Proxy only is close — engagement, sharing rates and fact-check volumes are plainly adjacent to the quantity of interest. Rejected because belief accuracy and behavioural outcomes are measured directly in this literature; the failure is agreement about the verdict, not availability of the measurement. Tier 4 considered (what people would otherwise have believed) and rejected for the same reason.","gap_name":"Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation "},{"gap_id":"1c1cb37e-2a00-80e2-a028-ebaa685df3da","dimension":"ai_type","original":"Reading and synthesis","audit":"Coordination and institutions","agreed":0,"adjudicated":null,"auditor_note":"test: If the best available contextualization and detection models were handed to the ecosystem tomorrow, would the gap close? No — most of that capability already exists and is not deployed, because ranking is optimized for engagement. Conversely, a governance and incentive change with 2026-level technology already produces measurable effect, as crowd-sourced note systems on major platforms show. The blocker is the incentive structure, which is the coordination category by definition. | nearest alternative: Reading and synthesis | maturity: Working now | LLM reasoning covers most of the named capabilities individually (contextualization engines, System 2 recommenders, misinformation detection) and would be the pick if adoption were free. Maturity Working now: the mechanism design and the model quality required both exist today.","gap_name":"Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation "},{"gap_id":"1c1cb37e-2a00-80e2-a028-ebaa685df3da","dimension":"outcome","original":"Platform operators get an objective other than engagement that they can actually implement: information environments designed to reward being right.","audit":"Whether a claim reaching a reader is true, and whether its spread is organic or coordinated, becomes visible at the moment of exposure rather than in a fact-check published days later to a different audience.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Point-of-exposure visibility is the thing that becomes possible; 'better information environment' would restate the gap.","gap_name":"Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation "},{"gap_id":"1b4cb37e-2a00-80b8-a4f8-ee697b40331b","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Is there an observable with a direction of improvement? Yes and it is already the field's scoreboard: prediction error on held-out perturbations and unseen cell types, measured against matched experimental readouts. A model is better or worse by a number that everyone computes the same way. | nearest alternative: Verification contested | Considered seriously: there is a live argument that current perturbation-prediction benchmarks are uninformative because trivial baselines match learned models. Rejected because that is a dispute about benchmark construction, resolvable by better benchmarks, not a dispute about whether measurement can settle the question — nobody argues that a model correctly predicting unseen perturbations would fail to demonstrate the capability.","gap_name":"Cellular and Biomolecular States Are\n  Highly Multimodal and Complex"},{"gap_id":"1b4cb37e-2a00-80b8-a4f8-ee697b40331b","dimension":"ai_type","original":"Prediction and modeling","audit":"Prediction and modeling","agreed":1,"adjudicated":null,"auditor_note":"test: Is the missing thing an extraction of signal from raw instrument output, or a learned model that predicts state and behaviour? The gap names 'predictive modeling of cell behavior' and the attached capability is a universal latent-variable model of cellular state — an emulator of the cell standing in for the experiment, which is the surrogate category. | nearest alternative: Measurement and sensing | maturity: 2-5 years | Sensing is a real contender given the multimodal-measurement framing (segmentation, deconvolution, cross-modal registration of omics layers). Rejected because those are inputs to the model rather than the gap itself. Maturity 2-5 years: single-cell foundation models exist now, but generalization to unseen perturbations and cell types is not there.","gap_name":"Cellular and Biomolecular States Are\n  Highly Multimodal and Complex"},{"gap_id":"1b4cb37e-2a00-80b8-a4f8-ee697b40331b","dimension":"outcome","original":"One model of cell state that predicts the response to a perturbation it has never seen, replacing a stack of per-assay descriptions.","audit":"A cell's response to a perturbation it has never seen becomes predictable in silico, so wet-lab capacity shifts from screening candidates to confirming them.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Prediction of the unseen perturbation is the specific thing that becomes possible; 'better understanding of cells' would not be an outcome.","gap_name":"Cellular and Biomolecular States Are\n  Highly Multimodal and Complex"},{"gap_id":"1c1cb37e-2a00-8026-a0fa-ff9d396d425a","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Is there a ground truth against which a candidate answer can be scored? Yes — fraction of unknown spectra assigned to the correct structure, benchmarked against independently confirmed structures, with an unambiguous direction. The field agrees in advance what counts as a correct assignment. | nearest alternative: Proxy only | Proxy only would apply if only database coverage or spectral quality were measurable. Rejected because assignment accuracy on held-out unknowns measures the gap itself. No tier-3 tension at all: there is no dispute about what a correct structure determination is.","gap_name":"Limited ability to identify molecular structures through spectroscopy"},{"gap_id":"1c1cb37e-2a00-8026-a0fa-ff9d396d425a","dimension":"ai_type","original":"Measurement and sensing","audit":"Measurement and sensing","agreed":1,"adjudicated":null,"auditor_note":"test: Is the task recovering a hidden object from instrument output, or emulating an expensive calculation? The gap frames it explicitly as an inverse problem on spectral data with information loss — reconstruction of structure from instrument signal, which the taxonomy places squarely under sensing and signal processing. | nearest alternative: Prediction and modeling | maturity: 2-5 years | Surrogates is close because the standard attack is a fast forward model (structure to predicted spectrum) used inside a search. Rejected because the forward direction is comparatively solved and the binding step is the inversion, plus one attached capability (microwave spectroscopy for 1:1 mapping) is instrument-side, reinforcing the sensing read. Maturity 2-5 years: NMR and MS elucidation models are strong on drug-like space now, general de novo elucidation of true unknowns is not solved.","gap_name":"Limited ability to identify molecular structures through spectroscopy"},{"gap_id":"1c1cb37e-2a00-8026-a0fa-ff9d396d425a","dimension":"outcome","original":"Read the structure straight off the spectrum. Crystallization and purification drop out of the critical path for structure determination.","audit":"An unknown compound in a complex mixture — a metabolite, an environmental contaminant, a natural product — becomes identifiable from its spectrum alone, with no authentic standard required.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | The 'no authentic standard' clause is what makes it an outcome rather than a speed-up: it removes a hard precondition, not just effort.","gap_name":"Limited ability to identify molecular structures through spectroscopy"},{"gap_id":"1c1cb37e-2a00-8044-b547-ddba3202ee63","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Can the unknown fraction itself be quantified rather than only gestured at? Yes — coverage and rarefaction estimators give a defensible figure for the share of diversity still unsampled, alongside straightforward counts of taxa, genomes and volume of ocean and atmosphere sampled, all with a clear direction. | nearest alternative: Proxy only | Proxy only is arguable on 'unknown unknowns' grounds — you cannot count what you have not seen. Rejected because estimating unseen diversity from sampling curves is standard practice and directly targets the gap, not an adjacent effect.","gap_name":"Much of the Biosphere Remains Uncharted and Vulnerable to Information Loss"},{"gap_id":"1c1cb37e-2a00-8044-b547-ddba3202ee63","dimension":"ai_type","original":"Physical build","audit":"Physical build","agreed":1,"adjudicated":null,"auditor_note":"test: Running experiments versus physical build, decided explicitly: does the system choose and run experiments in a closed loop at bench scale, or does it deploy hardware into the field? The attached capabilities are ocean exploration systems, dense ocean DNA sampling and atmospheric extremophile mapping — platforms that must be built and deployed into the deep ocean and stratosphere, which the taxonomy assigns to physical build and field deployment. | nearest alternative: Measurement and sensing | maturity: 2-5 years | Sensing (metagenomic assembly, annotating sequence dark matter) is real but is not the throughput limit; getting samplers to depth and altitude at density is. Maturity 2-5 years: autonomous underwater vehicles and robotic samplers work now, sustained global-density autonomous sampling does not.","gap_name":"Much of the Biosphere Remains Uncharted and Vulnerable to Information Loss"},{"gap_id":"1c1cb37e-2a00-8044-b547-ddba3202ee63","dimension":"outcome","original":"Most of Earth's biology is microbial and uncatalogued. It gets recorded before it is lost, and turns into something searchable.","audit":"The unsampled majority of Earth's genetic diversity enters the record before habitat change erases it, making lineages and biochemistries that currently have no representative in any database available as objects of study.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | The irreversibility framing (information loss) is part of the gap, so the outcome names what stops being lost, not merely what is catalogued.","gap_name":"Much of the Biosphere Remains Uncharted and Vulnerable to Information Loss"},{"gap_id":"1c1cb37e-2a00-8089-a828-f143cf6f5c86","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Is there an observable with an agreed direction? Yes, several and all conventional: existence and efficacy of a licensed vaccine for each named disease, time and cost from candidate to licensure, and burden averted. Nobody disputes what a successful trial endpoint means for these pathogens. | nearest alternative: Proxy only | Considered Counterfactual required on the 'under-provisioning' framing — the shortfall is against products that should exist but do not. Rejected because the named diseases give an explicit target list, so the gap is a census, not an unobservable. Proxy only rejected because trial efficacy measures the gap directly.","gap_name":"Under-Provisioning of Antibiotics, Vaccines and Other Interventions for Major Global Health Challenges"},{"gap_id":"1c1cb37e-2a00-8089-a828-f143cf6f5c86","dimension":"ai_type","original":"Design search","audit":"Design search","agreed":1,"adjudicated":null,"auditor_note":"test: If funding were unlimited tomorrow, would these products exist? No — TB, Group A Strep and hepatitis C vaccines have absorbed sustained non-market funding and still failed on immunological grounds, so the binding constraint is proposing immunogens and regimens that elicit protection, which is inverse design over a candidate space rather than allocation. | nearest alternative: Coordination and institutions | maturity: 2-5 years | Coordination is the obvious read of the word 'under-provisioning' and is correct for some of the portfolio (low-cost therapeutics for low-resource settings, malnutrition interventions). Rejected as primary because the flagship named diseases are scientific failures, not funding failures. Maturity 2-5 years: de novo binder and immunogen design work now in narrow cases, designing protection against an antigenically variable chronic pathogen does not.","gap_name":"Under-Provisioning of Antibiotics, Vaccines and Other Interventions for Major Global Health Challenges"},{"gap_id":"1c1cb37e-2a00-8089-a828-f143cf6f5c86","dimension":"outcome","original":"Tuberculosis and Group A Streptococcus have resisted vaccines for decades. They get effective countermeasures, and malnutrition's developmental mechanism turns into a design target.","audit":"Pathogens that have defeated vaccinology for decades — TB, Group A Strep, hepatitis C — move from 'no candidate mechanism of protection is known' to having a development timeline.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Names the state change (unknown to scheduled) rather than 'more vaccines exist', which restates the gap.","gap_name":"Under-Provisioning of Antibiotics, Vaccines and Other Interventions for Major Global Health Challenges"},{"gap_id":"1c1cb37e-2a00-80c4-b958-d7161b660b3e","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Is there an observable with a direction? Yes and it is routinely tracked in the industry: design and documentation hours per project, code-compliance error and rework rates, and cost and schedule variance from tender to completion. There is no dispute about what improvement looks like. | nearest alternative: Proxy only | Proxy only rejected because labour hours and rework are the gap itself ('complex and labor-intensive'), not an adjacent effect. No tier-3 tension: nobody argues about whether a design satisfies a building code.","gap_name":"Designing Buildings is Hard"},{"gap_id":"1c1cb37e-2a00-80c4-b958-d7161b660b3e","dimension":"ai_type","original":"Design search","audit":"Design search","agreed":1,"adjudicated":null,"auditor_note":"test: Is the output a document or a design satisfying hard constraints? A building must simultaneously satisfy structural, code, cost, energy and site constraints, so the task is constrained generative search over a defined space, not synthesis of text about buildings. | nearest alternative: Reading and synthesis | maturity: 2-5 years | LLM reasoning genuinely covers a large share of the labour (specifications, code interpretation, drawing annotation, contract documents) and would be the pick if paperwork were the whole gap. Rejected because the attached capability is design of buildings and construction plans, i.e. the artifact itself. Physical build rejected outright: the gap is design and planning, not on-site construction robotics. Maturity 2-5 years — generative layout and structural optimization tools exist, end-to-end constructible code-compliant generation does not.","gap_name":"Designing Buildings is Hard"},{"gap_id":"1c1cb37e-2a00-80c4-b958-d7161b660b3e","dimension":"outcome","original":"Architects can explore many building designs cheaply and pick on cost, energy and buildability. Today the design is settled early and everyone lives with it.","audit":"Hundreds of fully costed, code-compliant building designs become generatable and comparable per project, so how much of the design space gets explored stops being set by available drafting labour.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Weakest outcome in the batch — this gap is closer to an efficiency gain than to something becoming knowable, so the outcome is phrased as the constraint that is removed.","gap_name":"Designing Buildings is Hard"},{"gap_id":"1c1cb37e-2a00-8027-8eda-e6f4a472c3ad","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Could a measurement in principle be taken, with the field still disagreeing about what it showed? Yes. Gravitationally-induced entanglement between mesoscopic masses, tabletop torsion/interferometric tests, and Lorentz-violation searches in multi-messenger timing are all concrete, proposed, partly-built observables — and there is live literature arguing a classical gravitational field reproduces the entanglement signature, so the verdict would be disputed even after a successful run. | nearest alternative: Counterfactual required | Rejected tier 4 because there is no shortage of things to point an instrument at; the shortage is agreement on what a positive result would license. This gap is the taxonomy's own calibration anchor for tier 3, and I reached the same value from the discriminating test rather than from the anchor.","gap_name":"Quantum Gravity is Experimentally Hard to Constrain\n"},{"gap_id":"1c1cb37e-2a00-8027-8eda-e6f4a472c3ad","dimension":"ai_type","original":"Physical build","audit":"Design search","agreed":0,"adjudicated":null,"auditor_note":"test: If AI made building giant detectors cheap but the field still had no configuration whose outcome it agreed would discriminate quantum from classical gravity, would the gap close? No. If AI proposed a tabletop configuration with an agreed discriminating signature, would it close? Largely yes. So the binding constraint is experiment proposal over a space of configurations, not fabrication and not data extraction. | nearest alternative: Physical build | maturity: Speculative | Physical build is the taxonomy's explicit exemplar for 'would build a detector, a beamline, an observatory' and the gap text names cost and scale of construction, so it is genuinely close. I rejected it on the test above. Measurement and sensing is a real secondary (squeezing bounds out of existing LIGO and gamma-ray-burst timing data, Working now) but moves the gap only marginally. Maturity Speculative: no AI system today proposes physically novel experimental schemes in fundamental physics.","gap_name":"Quantum Gravity is Experimentally Hard to Constrain\n"},{"gap_id":"1c1cb37e-2a00-8027-8eda-e6f4a472c3ad","dimension":"outcome","original":"Quantum gravity gets an experimental constraint that the field agrees is one, and becomes a partly empirical subject.","audit":"Whether gravity is quantised becomes an empirically decidable question rather than a theoretical preference, because a laboratory-scale observable with an agreed interpretation would exist.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Deliberately names the epistemic state that changes, not the activity (running more experiments).","gap_name":"Quantum Gravity is Experimentally Hard to Constrain\n"},{"gap_id":"1c1cb37e-2a00-8055-b34e-fe7956440682","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Is there a candidate direct observable being built, whose interpretation the field would dispute? Yes — Convergent's own linked capabilities include 'Metrics to Track Human Influence' and 'Early Warning Signs for Human Disempowerment', i.e. the observable is under construction. Nobody would accept any single index (labour share, share of consequential decisions taken without human review, measured institutional responsiveness) as settling whether disempowerment is occurring. | nearest alternative: Counterfactual required | Hardest tier call in this batch. The tier-4 reading is strong: the quantity of interest is arguably the world in which humans retained influence, which never happens. I rejected it because the claimed process is asserted to be underway now and produces present-tense observables; the disagreement is over interpretation, not over the absence of anything to measure. Proxy only was the third candidate and fails because the contest is precisely about whether the candidate direct measures are direct. Confidence here is closer to guess than confident.","gap_name":"Labor-Replacing AI Could Lead to Human Disempowerment"},{"gap_id":"1c1cb37e-2a00-8055-b34e-fe7956440682","dimension":"ai_type","original":"Coordination and institutions","audit":"Coordination and institutions","agreed":1,"adjudicated":null,"auditor_note":"test: Is the blocker a missing technique or a missing organisational arrangement? Five of the seven linked capabilities (voting and auditing protocols, deliberative democracy tools, oversight interventions, influence metrics, early-warning indicators) are institutional artefacts, not model capabilities. A better model does not restore human influence if the competitive incentive to remove humans is unchanged. | nearest alternative: Reading and synthesis | maturity: Speculative | LLM reasoning is the obvious reading of 'AI Agents to Advocate for Human Interests' and 'Technology to Augment Humans', and it is Working now — which is exactly why it is the wrong primary: those tools already exist in usable form and the gap is unmoved. Maturity Speculative rather than 2-5 years because AI-mediated deliberation and auditing exist at pilot scale (Polis/vTaiwan, collective-constitution exercises, the Habermas-machine line of work) but no demonstrated mechanism maintains human influence against the competitive pressure the gap describes.","gap_name":"Labor-Replacing AI Could Lead to Human Disempowerment"},{"gap_id":"1c1cb37e-2a00-8055-b34e-fe7956440682","dimension":"outcome","original":"Erosion of human influence over consequential decisions becomes visible while it is still reversible.","audit":"Erosion of human influence over economic and political decisions becomes detectable while it is still reversible, instead of being recognisable only in retrospect.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Outcome is detectability-in-time, which is what the linked early-warning and metrics capabilities actually buy; 'better governance' would have restated the gap.","gap_name":"Labor-Replacing AI Could Lead to Human Disempowerment"},{"gap_id":"1c1cb37e-2a00-8099-b97a-ce85f02ab5a5","dimension":"measurability","original":"Verification contested","audit":"Verification contested","agreed":1,"adjudicated":null,"auditor_note":"test: Has a measurement already been taken that the field declines to treat as settling the question? Yes — the 2007-2010 two-dimensional electronic spectroscopy results on the FMO complex were measured, then largely reattributed to vibronic coherence, and the field does not agree that any observed coherence lifetime demonstrates functionally relevant quantum effects. So the instrument exists and the verdict is contested. | nearest alternative: Directly measurable | Coherence lifetime is an observable with a direction of improvement, which makes tier 1 tempting. Rejected because the gap as written is about quantum effects influencing biomolecular interactions, and the mapping from any measured coherence to functional influence is exactly what is disputed. Radical-pair magnetoreception is the same shape: measurable magnetic-field effects, contested inference to a quantum mechanism.","gap_name":"Lack of Direct Measurement of Quantum Effects in Biological Systems"},{"gap_id":"1c1cb37e-2a00-8099-b97a-ce85f02ab5a5","dimension":"ai_type","original":"Measurement and sensing","audit":"Measurement and sensing","agreed":1,"adjudicated":null,"auditor_note":"test: Is the missing thing a prediction of what the effect should look like, or the ability to pull a weak, fast, warm-and-noisy coherence signature out of instrument data? The description says direct experimental evidence is needed, and the single linked capability is 'Quantum Biology Measurements' — measurement, not theory. The failure mode in this field is that the signature and the decoherent background are not separable by existing analysis. | nearest alternative: Prediction and modeling | maturity: 2-5 years | Surrogates for open-quantum-system dynamics in a protein environment are a real secondary and would sharpen what to look for, but they cannot supply the direct evidence the gap demands. Maturity 2-5 years: ML denoising and reconstruction on ultrafast spectroscopy is adjacent-domain working, but ML that discriminates vibronic from electronic coherence to the field's satisfaction is not demonstrated.","gap_name":"Lack of Direct Measurement of Quantum Effects in Biological Systems"},{"gap_id":"1c1cb37e-2a00-8099-b97a-ce85f02ab5a5","dimension":"outcome","original":"Whether quantum coherence does anything functional in biology gets settled in a lab.","audit":"Whether evolution exploits quantum coherence becomes answerable, and if it does, quantum design rules become available for engineering enzymes, light-harvesting and biological sensors.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Names the design rules that become usable, not the experiments that become possible.","gap_name":"Lack of Direct Measurement of Quantum Effects in Biological Systems"},{"gap_id":"1c1cb37e-2a00-80b5-8cf4-c8195273fa45","dimension":"measurability","original":"Counterfactual required","audit":"Counterfactual required","agreed":1,"adjudicated":null,"auditor_note":"test: Is there an observation set for the quantity of interest? No. The quantity is the R&D that rigid structures prevented — projects never proposed because no institution could host them. You can count what was funded; you cannot enumerate the suppressed alternatives, and no instrument could be pointed at them. Contrast the tier-3 test: there is no candidate measurement whose verdict people would argue about, because there is no candidate measurement. | nearest alternative: Proxy only | Proxy only is defensible — scientist time-on-admin, grant success rates, and organisational-form diversity are all measurable adjacent effects. Rejected because those measure the input side while the gap's claim is about suppressed output, which has no observation set. This gap is the taxonomy's tier-4 calibration anchor and the test reproduces it.","gap_name":"A Limited Set of Rigid Organizational Structures for Organizing and Funding Research Constrains the Forms of R&D That Get Done"},{"gap_id":"1c1cb37e-2a00-80b5-8cf4-c8195273fa45","dimension":"ai_type","original":"Coordination and institutions","audit":"Coordination and institutions","agreed":1,"adjudicated":null,"auditor_note":"test: Is the blocker a technique or an organisation? All three linked capabilities (decentralised funding, reduced fundraising burden, novel research-organisation structures) are allocation and incentive designs. The taxonomy assigns 'anything where the blocker is an organisation rather than a technique' to this type by definition, and this gap is self-described as a meta-bottleneck about incentive structures. | nearest alternative: Reading and synthesis | maturity: 2-5 years | LLM reasoning is the strongest rival and is Working now: agents that absorb grant-writing and administrative load directly attack 'scientists spending a lot of time not doing science'. Rejected as primary because it relieves the symptom while leaving the incentive structure that produces the rigidity intact. Maturity 2-5 years: AI-assisted review, matching and allocation are running as pilots, not as institutions.","gap_name":"A Limited Set of Rigid Organizational Structures for Organizing and Funding Research Constrains the Forms of R&D That Get Done"},{"gap_id":"1c1cb37e-2a00-80b5-8cf4-c8195273fa45","dimension":"outcome","original":"Long-horizon, heavily coordinated research that no current institutional form can house becomes fundable and organisable.","audit":"Research programmes that no current institution can host — long-horizon, tightly coordinated, or cross-disciplinary at scale — become fundable and therefore attemptable.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | The outcome is a class of work becoming possible, not scientists having more time.","gap_name":"A Limited Set of Rigid Organizational Structures for Organizing and Funding Research Constrains the Forms of R&D That Get Done"},{"gap_id":"1c1cb37e-2a00-800f-9e2a-cf61f6212c7e","dimension":"measurability","original":"Proxy only","audit":"Directly measurable","agreed":0,"adjudicated":null,"auditor_note":"test: Does an observable quantity exist with a direction of improvement, for the gap itself rather than an adjacent effect? Yes — the gap is a tooling deficit, and tooling capability is directly observable: cost and throughput per qualitative interview, coverage and response rate of survey instruments, and archaeological sites detected per unit area surveyed (a quantity satellite-ML archaeology already reports and improves). | nearest alternative: Proxy only | Genuinely mixed gap. The 'identify and prioritise important questions that will have an impact' clause is not directly measurable — research impact is proxy-measured and arguably counterfactual — but two of the three linked capabilities and most of the description text are instrument deficits with direct metrics. I labelled on the dominant content and flag the heterogeneity; a labeller weighting the prioritisation clause would land on Proxy only and would not be unreasonable.","gap_name":"Underdevelopment of Modern Tools in the Social Sciences"},{"gap_id":"1c1cb37e-2a00-800f-9e2a-cf61f6212c7e","dimension":"ai_type","original":"Reading and synthesis","audit":"Reading and synthesis","agreed":1,"adjudicated":null,"auditor_note":"test: Which capability class do the linked capabilities fall into, and which one does the description's own emphasis name? Two of three ('Infrastructure for Problem-Driven Research', 'Modernize Survey Infrastructure') are synthesis-and-language problems — coding open-ended responses, conducting and analysing qualitative instruments, mapping a literature to identify tractable questions — and the description explicitly says AI-enabled qualitative methods could super-charge the field. | nearest alternative: Measurement and sensing | maturity: Working now | Satellite-ML archaeology is unambiguously sensing and is the most mature single item here, but it covers one of three capabilities and is a secondary in a gap whose stated core is qualitative data and question selection. Maturity Working now: LLM coding of qualitative data and LLM-administered interviewing are in production use today; the gap is adoption and infrastructure, not capability.","gap_name":"Underdevelopment of Modern Tools in the Social Sciences"},{"gap_id":"1c1cb37e-2a00-800f-9e2a-cf61f6212c7e","dimension":"outcome","original":"Qualitative and survey evidence scales the way quantitative data already does, so the how and why of social outcomes stops being the underpowered half of the field.","audit":"Causal accounts of social outcomes that depend on unrecorded qualitative context become testable at population scale, and archaeological landscapes never walked by a surveyor become findable.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Two clauses because the gap is genuinely two gaps; each names a thing that becomes knowable rather than easier.","gap_name":"Underdevelopment of Modern Tools in the Social Sciences"},{"gap_id":"1c1cb37e-2a00-8085-bd56-cdf566ac7e42","dimension":"measurability","original":"Proxy only","audit":"Counterfactual required","agreed":0,"adjudicated":null,"auditor_note":"test: Is there an observation set for the quantity of interest? No. The gap is the absence of a field, and its content is the transport, energy and civil-engineering work that is not being done. That is the same shape as the taxonomy's tier-4 anchor (research that never happened), and unlike tier 3 there is no candidate measurement whose verdict anyone is arguing about — nobody is running the measurement at all. | nearest alternative: Proxy only | Hard call. Proxy only is arguable (publications, funding and person-years in terraforming-adjacent work are all countable), and so, on a literal reading, is Directly measurable (a field either exists or it does not). I resolved by consistency with the taxonomy's own institutional-absence anchor, which takes the same countable-inputs objection and still lands on tier 4.","gap_name":"Lack of a Dedicated Field for Planetary Terraforming"},{"gap_id":"1c1cb37e-2a00-8085-bd56-cdf566ac7e42","dimension":"ai_type","original":"Coordination and institutions","audit":"Coordination and institutions","agreed":1,"adjudicated":null,"auditor_note":"test: Which blocker does the gap name? It names the absence of an established field, and both linked capabilities are 'Directed Work on ...' — that is agenda-setting, funding and coordination, which the taxonomy assigns to this type by definition ('the blocker is an organisation rather than a technique'). | nearest alternative: Physical build | maturity: 2-5 years | The closest call on ai_type in this batch, and I want the disagreement recorded rather than smoothed. The counter-test is strong: if a fully funded terraforming institute existed tomorrow, the transport, energy and planetary civil-engineering challenges would remain physically unattainable, which says the real binding constraint is autonomous large-scale off-world construction (Physical build, Speculative). I labelled the blocker the gap description actually names. A labeller applying the binding-constraint test instead would say Physical build / Speculative, and I would not call that wrong. Maturity 2-5 years for the coordination reading: AI-assisted roadmapping and research-agenda synthesis are near-term, not present, capabilities.","gap_name":"Lack of a Dedicated Field for Planetary Terraforming"},{"gap_id":"1c1cb37e-2a00-8085-bd56-cdf566ac7e42","dimension":"outcome","original":"Terraforming gets a research community and a shared problem list, so its transport, energy and civil engineering questions get worked on deliberately.","audit":"Whether a planet can be deliberately made habitable becomes a question with an engineering answer — a costed, physically constrained one — instead of a speculative one.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | The outcome is a feasibility verdict becoming available, not a field being founded (which would restate the gap).","gap_name":"Lack of a Dedicated Field for Planetary Terraforming"},{"gap_id":"1c1cb37e-2a00-80ea-93f0-c8ed8e7f2155","dimension":"measurability","original":"Proxy only","audit":"Verification contested","agreed":0,"adjudicated":null,"auditor_note":"test: For the components that dominate the gap, could a measurement be taken with the field still disputing the verdict? Yes, twice over. Decoded animal communication has no agreed criterion for what would count as having decoded it — there is active disagreement over whether statistical structure in sperm whale codas or elephant calls demonstrates semantic content or only correlational regularity. And for Hadean conditions, a recovered zircon or lunar-regolith fragment can be measured, but its inference to primordial planetary conditions is contested at every step. | nearest alternative: Directly measurable | The nanostructure-imaging component alone would be tier 1 (resolution, with a direction of improvement), which is the reason tier 1 is the nearest alternative. Rejected because the two components that carry the gap's ambition — animal communication and life's genesis — fail the agreed-verdict test. Composite gap; the tier is the tier of its hardest load-bearing part.","gap_name":"We Can Learn More from Nature’s Biological Designs"},{"gap_id":"1c1cb37e-2a00-80ea-93f0-c8ed8e7f2155","dimension":"ai_type","original":"Measurement and sensing","audit":"Measurement and sensing","agreed":1,"adjudicated":null,"auditor_note":"test: Counting linked capabilities by class: advanced nanostructure imaging, ML animal-communication decoding, massive-scale zircon screening, and ancient-protein reconstruction are all extraction of structure from instrument data — four of six. The remaining two (Europa/Enceladus missions, raking lunar regolith) are physical deployment. The gap's recurring phrase is that things 'remain hidden due to current imaging limitations' — a detection-threshold problem. | nearest alternative: Reading and synthesis | maturity: Working now | Treating animal communication as a sequence-modelling or translation problem is the obvious LLM reading and I rejected it because that work bottlenecks upstream, on segmentation, source separation and detection in field recordings, before any language model has units to model. Physical build is a genuine secondary here (Europa lander, lunar regolith raking) at Speculative maturity and I record it rather than fold it into sensing. Maturity Working now for the primary: ML detection and denoising in bioacoustics, microscopy and mineral screening are deployed today.","gap_name":"We Can Learn More from Nature’s Biological Designs"},{"gap_id":"1c1cb37e-2a00-80ea-93f0-c8ed8e7f2155","dimension":"outcome","original":"Biological nanostructures, animal communication and the chemistry of early Earth become readable as engineering precedent.","audit":"Biological structure and signal currently below the detection threshold — nanoscale architectures, the information content of animal calls, and the chemistry of the pre-3.5 Ga Earth — enter the observational record and become available as design precedent.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Composite gap; the outcome names what enters the record rather than what researchers do.","gap_name":"We Can Learn More from Nature’s Biological Designs"},{"gap_id":"1b4cb37e-2a00-80e6-8b25-fd9e770fbb3f","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Does the gap name its own observable? Yes — resolution, imaging depth in intact 3D tissue, number of simultaneously resolved molecular species, and duration of live acquisition before phototoxicity, each with an unambiguous direction of improvement. No interpretive dispute exists about whether a higher-resolution in-situ molecular map is better. | nearest alternative: Proxy only | Proxy only would apply if the gap were 'we do not understand molecular organisation in tissue' and imaging metrics were the adjacent stand-in. It is not: the gap is stated as an imaging capability deficit, so the imaging metrics are the quantity itself.","gap_name":"Limited Ability to Image Molecules in Their Native Contexts"},{"gap_id":"1b4cb37e-2a00-80e6-8b25-fd9e770fbb3f","dimension":"ai_type","original":"Measurement and sensing","audit":"Measurement and sensing","agreed":1,"adjudicated":null,"auditor_note":"test: Classifying the five linked capabilities: molecular 3D scanning, live-cell subcellular imaging and foundation models, and multiplexed live-cell dynamics recordings are all reconstruction from instrument data — three of five. Live 3D imaging is photon-budget-limited, and learned denoising, deconvolution and super-resolution are what convert an unusably low-dose acquisition into a usable one. | nearest alternative: Design search | maturity: Working now | 'Binders for Every Epitope' is a serious rival primary: de novo binder design is arguably the strongest current AI-for-science capability, and comprehensive molecular mapping is coverage-limited as much as optics-limited. I rejected it because binders solve specificity and would still leave the depth and live-3D constraints the gap explicitly names. This is a close call and I would accept Design search / Working now on adjudication. Maturity Working now: content-aware restoration and learned reconstruction are standard practice in live-cell microscopy.","gap_name":"Limited Ability to Image Molecules in Their Native Contexts"},{"gap_id":"1b4cb37e-2a00-80e6-8b25-fd9e770fbb3f","dimension":"outcome","original":"Molecular identity and position get read inside intact three-dimensional tissue, so spatial maps describe real organs.","audit":"Molecular composition becomes mappable inside intact living tissue, so molecular state can be tied to cellular behaviour without destroying the context that produced it.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | The load-bearing clause is 'without destroying the context' — that is what becomes knowable that was not before.","gap_name":"Limited Ability to Image Molecules in Their Native Contexts"},{"gap_id":"1c1cb37e-2a00-8031-b79b-c13cca53b9c9","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Are the two quantities the gap names directly assayable with agreed directions? Yes — delivery efficiency into live cells, and the magnitude of the perturbation the probe causes (viability, transcriptional response, altered kinetics). Both are routinely measured and neither is interpretively contested. | nearest alternative: Proxy only | Rejected Proxy only because perturbation is not an adjacent effect standing in for something unmeasurable — it is the harm the gap is defined by, and it is measured directly.","gap_name":"Difficulty Delivering Physical Probes for Imaging into Living Cells"},{"gap_id":"1c1cb37e-2a00-8031-b79b-c13cca53b9c9","dimension":"ai_type","original":"Design search","audit":"Measurement and sensing","agreed":0,"adjudicated":null,"auditor_note":"test: Does the gap ask for a better way to get a probe in, or for a way to obtain the same information without one? The description's closing clause is explicit: enable high-dimensional biosensing without invasive probes. That is inference of molecular content from intrinsic contrast — phase, Raman, autofluorescence lifetime — which is signal extraction, not delivery. | nearest alternative: Design search | maturity: 2-5 years | Design search is the natural reading of 'Metabolically Incorporated Labels' and of probe chemistry generally, and it would be the right primary if the gap were framed as needing better probes. It is framed as needing to stop using them. Physical build was considered and dismissed early: this is subcellular delivery chemistry, not fabrication robotics. Maturity 2-5 years: label-free ML inference of molecular identity is demonstrated at low dimensionality and does not yet reach the high-dimensional readout the gap requires.","gap_name":"Difficulty Delivering Physical Probes for Imaging into Living Cells"},{"gap_id":"1c1cb37e-2a00-8031-b79b-c13cca53b9c9","dimension":"outcome","original":"Watch biochemistry happen inside a living cell without disturbing the process you are watching.","audit":"Molecular dynamics inside living cells become observable without altering the cell being observed, removing the perturbation confound from live-cell measurement.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Names the confound that disappears, which is the epistemic gain; 'better imaging' would have restated the gap.","gap_name":"Difficulty Delivering Physical Probes for Imaging into Living Cells"},{"gap_id":"1c1cb37e-2a00-8057-ae20-c47fd74df08d","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Is there an agreed pass/fail observable? Yes — autonomous self-replication of a system assembled from defined, purified components, with intermediate metrics along the way (number of reconstituted functions, generations of division sustained, fraction of behaviour predicted by the design model). The field would not dispute a vesicle that divided for many generations from a specified parts list. | nearest alternative: Verification contested | Tier 3 has some purchase — there is real argument about what counts as a minimal or truly bottom-up cell, and about whether JCVI-syn3.0-style work qualifies. Rejected because that is a definitional boundary dispute, not a dispute about whether an experiment would settle anything; the functional criteria are agreed even where the label is not.","gap_name":"Synthetic Biology Platforms Are Over-Reliant on Evolved Cells That We Don’t Fully Understand or Control"},{"gap_id":"1c1cb37e-2a00-8057-ae20-c47fd74df08d","dimension":"ai_type","original":"Design search","audit":"Running experiments","agreed":0,"adjudicated":null,"auditor_note":"test: Is the missing thing a design we cannot compute, or an empirical search over a combinatorial space no forward model can predict? Bottom-up cells fail on the emergent behaviour of reconstituted mixtures — component stoichiometry, crowding, membrane composition, energy regeneration — which no model predicts, so the loop has to close on the bench. That is closed-loop formulation optimisation, the taxonomy's stated home ground for this type. | nearest alternative: Design search | maturity: 2-5 years | This is one of the two distinctions the protocol flags. Design search (in-silico minimal-genome and protein-component design) is the rival and I rejected it because there is no forward model good enough to make pure in-silico design meaningful here. I also explicitly checked Physical build and rejected it: this is bench-scale liquid handling and reconstitution, not fabrication or assembly robotics, and collapsing the two would destroy the maturity distinction. Maturity 2-5 years: active-learning optimisation of cell-free expression systems is working now in narrow form, but closed-loop search over whole synthetic-cell assembly is not.","gap_name":"Synthetic Biology Platforms Are Over-Reliant on Evolved Cells That We Don’t Fully Understand or Control"},{"gap_id":"1c1cb37e-2a00-8057-ae20-c47fd74df08d","dimension":"outcome","original":"Cells get built from specified parts, so synthetic biology stops inheriting mechanisms nobody has characterized.","audit":"Living systems become buildable from specified parts, making cellular behaviour predictable from design rather than inferred after the fact from evolved organisms.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | The outcome is predictability-from-design, which is what distinguishes a bottom-up cell from a smaller evolved one.","gap_name":"Synthetic Biology Platforms Are Over-Reliant on Evolved Cells That We Don’t Fully Understand or Control"},{"gap_id":"1c1cb37e-2a00-809e-8e26-e1c9f14f5143","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Does a binary observable with a direction exist? Yes — for a given compound, did a diffraction-quality crystal form and was a structure solved. Aggregated over a compound library this gives an uncontested success rate, and no one disputes what a solved structure shows. | nearest alternative: Proxy only | Rejected Proxy only because crystallisation outcome is the quantity itself, not an input standing in for something hidden. The clearest tier-1 case in the batch.","gap_name":"Many Molecules Can’t Easily Be Crystallized"},{"gap_id":"1c1cb37e-2a00-809e-8e26-e1c9f14f5143","dimension":"ai_type","original":"Prediction and modeling","audit":"Prediction and modeling","agreed":1,"adjudicated":null,"auditor_note":"test: Is the blocker throughput of trying conditions, or the absence of a forward model of nucleation and growth? High-throughput robotic crystallisation screening already exists at scale and has not closed the gap, which rules out throughput. The description names the remedy itself: improved computational models of crystal growth — a learned emulator of a process too expensive to simulate from first principles. | nearest alternative: Design search | maturity: 2-5 years | This is the protocol's other flagged pressure point (surrogates vs design search). Design search would be right if the ask were for a proposer of co-formers, additives or solvent systems; the ask is for a forward model that says what will happen, and the proposer is only useful once that model exists. Running experiments was the third candidate and is rejected by the existing-screening-robots argument above. Maturity 2-5 years: crystal structure prediction is working now for small rigid organics, but predicting crystallisability and growth kinetics for recalcitrant molecules is not.","gap_name":"Many Molecules Can’t Easily Be Crystallized"},{"gap_id":"1c1cb37e-2a00-809e-8e26-e1c9f14f5143","dimension":"outcome","original":"Crystallization becomes something you plan. The trial-and-error step that currently gates structure determination for hard molecules goes away.","audit":"Atomic structures become determinable for the large class of molecules that currently resist crystallisation, bringing compound families invisible to structure-based work into it.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Names the compound classes that become tractable rather than the crystallisation attempts that get cheaper.","gap_name":"Many Molecules Can’t Easily Be Crystallized"},{"gap_id":"1c1cb37e-2a00-80cd-9ae4-ee6fa18c02ed","dimension":"measurability","original":"Directly measurable","audit":"Directly measurable","agreed":1,"adjudicated":null,"auditor_note":"test: Does an uncontested observable with a direction exist for the gap as stated? Yes — titre, rate and yield at production scale, cost per kilogram, and the ratio of achieved to scale-down-predicted productivity. Scale-up failure is quantified as a matter of routine industrial practice. | nearest alternative: Proxy only | Rejected Proxy only because industrial-scale productivity is the quantity of interest, not a stand-in for it.","gap_name":"Poor Scalability of Bioreactors Limits Biomanufacturing"},{"gap_id":"1c1cb37e-2a00-80cd-9ae4-ee6fa18c02ed","dimension":"ai_type","original":"Physical build","audit":"Prediction and modeling","agreed":0,"adjudicated":null,"auditor_note":"test: Is a design space defined and awaiting a proposer, or do we have candidate designs we cannot evaluate without prohibitively expensive pilot builds? Bioreactor scale-up fails at evaluation: mixing time, oxygen transfer and shear distribution at 100,000 L cannot be known without building the vessel, and coupled CFD-plus-kinetics simulation is too costly to sit inside a design loop. A learned emulator is what puts evaluation inside the loop. | nearest alternative: Design search | maturity: 2-5 years | Design search is close and is the natural reading of 'Modularization of Bioreactors' — inverse design of geometry and module topology. Rejected on the evaluation-vs-proposal test: the proposer is cheap and already exists in engineering practice; the evaluator is what is missing. Physical build was considered because the artefact is hardware, and rejected because the blocker is engineering knowledge, not the ability to fabricate vessels. Maturity 2-5 years: CFD surrogates are working in adjacent process industries but not validated across the scale range that matters here.","gap_name":"Poor Scalability of Bioreactors Limits Biomanufacturing"},{"gap_id":"1c1cb37e-2a00-80cd-9ae4-ee6fa18c02ed","dimension":"outcome","original":"Bioprocess economics currently fall apart at exactly the volume where they start to matter commercially. Yields hold from bench to tank.","audit":"Bioproducts that are economical at litre scale become economical at industrial scale, making biomanufacturing a live route to commodity chemicals and materials rather than to high-value niches only.","agreed":0,"adjudicated":null,"auditor_note":"test: n/a | nearest alternative: n/a | Names the product class that becomes buildable, not the reactors that get better.","gap_name":"Poor Scalability of Bioreactors Limits Biomanufacturing"}],"audit_summary":{"n_sampled":36,"n_population":103,"dimensions":{"measurability":{"n":36,"agreed":28,"raw_disagreement":0.2222222222222222,"strata":[{"stratum":"Directly measurable","population":72,"sampled":15,"agreed":15,"disagreement":0},{"stratum":"Proxy only","population":19,"sampled":9,"agreed":2,"disagreement":0.7777777777777778},{"stratum":"Verification contested","population":11,"sampled":11,"agreed":10,"disagreement":0.09090909090909091},{"stratum":"Counterfactual required","population":1,"sampled":1,"agreed":1,"disagreement":0}],"weighted_disagreement":0.15318230852211437,"disagreements":[{"gap":"Uncertainty and Noise in the Science of Room-Temperature Superconductivity","original":"Verification contested","audit":"Directly measurable"},{"gap":"Current “Model Systems” for Brain Function are Not Representative of the Real Human Brain","original":"Proxy only","audit":"Verification contested"},{"gap":"AI Could Be Misused","original":"Proxy only","audit":"Verification contested"},{"gap":"Major Planetary Science and Astrobiology Missions Are Not Realized by Existing Government Space Agencies","original":"Proxy only","audit":"Directly measurable"},{"gap":"Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation ","original":"Proxy only","audit":"Verification contested"},{"gap":"Underdevelopment of Modern Tools in the Social Sciences","original":"Proxy only","audit":"Directly measurable"},{"gap":"Lack of a Dedicated Field for Planetary Terraforming","original":"Proxy only","audit":"Counterfactual required"},{"gap":"We Can Learn More from Nature’s Biological Designs","original":"Proxy only","audit":"Verification contested"}]},"ai_type":{"n":36,"agreed":24,"raw_disagreement":0.3333333333333333,"strata":[],"weighted_disagreement":null,"disagreements":[{"gap":"Uncertainty and Noise in the Science of Room-Temperature Superconductivity","original":"Coordination and institutions","audit":"Prediction and modeling"},{"gap":"Our Platforms for Civic Engagement and Democratic Decision-Making Don’t Take Advantage of 21st Century Scalable Technology","original":"Coordination and institutions","audit":"Reading and synthesis"},{"gap":"Robust and Compact Plasma Confinement for Fusion is Still Not Solved","original":"Prediction and modeling","audit":"Design search"},{"gap":"We Have a Limited Ability to Acquire, Concentrate and Substitute Chemical Elements in Processes","original":"Design search","audit":"Running experiments"},{"gap":"Biological Life is Our Only Working Example of Complex Evolved Computation","original":"Prediction and modeling","audit":"Design search"},{"gap":"Intervening in Earth Systems at Scale is Largely Untested","original":"Physical build","audit":"Prediction and modeling"},{"gap":"AI Could Be Misused","original":"Coordination and institutions","audit":"Reading and synthesis"},{"gap":"Limited Tools for Improving Individual, Social and Societal Epistemics in the Face of Misinformation ","original":"Reading and synthesis","audit":"Coordination and institutions"},{"gap":"Quantum Gravity is Experimentally Hard to Constrain ","original":"Physical build","audit":"Design search"},{"gap":"Difficulty Delivering Physical Probes for Imaging into Living Cells","original":"Design search","audit":"Measurement and sensing"},{"gap":"Synthetic Biology Platforms Are Over-Reliant on Evolved Cells That We Don’t Fully Understand or Control","original":"Design search","audit":"Running experiments"},{"gap":"Poor Scalability of Bioreactors Limits Biomanufacturing","original":"Physical build","audit":"Prediction and modeling"}]}}},"relabel":{"n":103,"type_agreed":77,"maturity_agreed":63,"v1":{"Physical build":{"n":17,"now":0},"Measurement and sensing":{"n":15,"now":3},"Prediction and modeling":{"n":20,"now":6},"Coordination and institutions":{"n":15,"now":1},"Reading and synthesis":{"n":10,"now":6},"Design search":{"n":21,"now":4},"Running experiments":{"n":5,"now":2}},"v2":{"Physical build":{"n":8,"now":1},"Measurement and sensing":{"n":19,"now":9},"Prediction and modeling":{"n":22,"now":3},"Coordination and institutions":{"n":15,"now":0},"Running experiments":{"n":9,"now":5},"Reading and synthesis":{"n":10,"now":5},"Design search":{"n":16,"now":2},"Real-time control":{"n":4,"now":1}}},"decisions":[{"phase":"phase-2","decision":"Maturity is scored relative to the specific gap, not to the capability class in general.","rationale":"Two blind auditors independently flagged the taxonomy as ambiguous between the two readings, which give different answers. Both auditors and the original labeler had already converged on the for-this-gap reading, so fixing the wording changes no existing label but makes the maturity disagreement rate interpretable.","runner_up":"Score maturity per capability class globally, which is simpler to apply but collapses the distinction between autonomous experimentation in chemistry and in orbital assembly.","confidence":"high","reversal_condition":"If a future batch shows labelers cannot apply the for-this-gap reading consistently, switch to global class maturity and record the loss."},{"phase":"phase-2","decision":"Keep a single primary AI type and a single tier per gap, and name composite gaps as a stated limitation rather than splitting them.","rationale":"Two gaps bundle several unrelated research programmes whose sub-components would take different types and tiers. Splitting them would mean creating gap records Convergent did not write, which violates the additive-only constraint and is exactly the schema redesign the brief forbids.","runner_up":"Allow multiple tiers per gap, or decompose composite gaps into sub-gaps.","confidence":"medium","reversal_condition":"If Convergent respond that they would welcome decomposition, split them in a later version."},{"phase":"phase-2","decision":"Publish the audit disagreement rate as measured, including the population-weighted figure, without re-adjudicating labels to improve it.","rationale":"The brief requires the artifact to be auditable rather than confident. A tuned agreement rate would defeat the purpose of running the audit, and a near-zero rate would be evidence the audit was not independent rather than evidence the labels are good.","runner_up":"Adjudicate every disagreement to a single answer and report only the post-adjudication labels.","confidence":"high","reversal_condition":"Never — if this is reversed the audit stops measuring anything."},{"phase":"phase-2","decision":"Downgrade every 'Proxy only' tier to confidence 'guess', including gaps the auditor never sampled.","rationale":"The stratum ran 78 percent disagreement against 0 percent for Directly measurable and 9 percent for Verification contested, and all three auditors independently reported the category was repeatedly the nearest alternative and almost never won. A category that two independent labelers applying the same written definition cannot agree on has not earned confident anywhere, and downgrading only the sampled failures would let an unreliable category keep its confidence wherever the auditor happened not to look.","runner_up":"Downgrade only the eight sampled disagreements, which understates the problem and implies the unsampled Proxy only labels are sound.","confidence":"high","reversal_condition":"If Proxy only is redefined sharply enough that a re-run audit brings its disagreement rate near the other strata, restore confidence on relabeled rows."},{"phase":"phase-2","decision":"Do not add a 'Control of physical systems' category to the AI type taxonomy in this version; record it as a finding instead.","rationale":"A blind auditor found that closed-loop control of a physical system has no home: fusion plasma control chooses no experiments so it is not autonomous experimentation, and builds nothing so it is not physical build. The case is real, but adding an eighth category mid-run would invalidate the 103 labels already applied and the audit measured against them. Naming it is the contribution; acting on it belongs in a later version.","runner_up":"Add the eighth category now and relabel all 103 gaps against it.","confidence":"medium","reversal_condition":"If a re-run is scheduled anyway, add the category first and relabel from scratch."},{"phase":"phase-2b","decision":"Withdraw the working-now maturity gradient as a finding. Report maturity per gap with its disagreement rate, and make no aggregate claim resting on it.","rationale":"An independent relabel of all 103 gaps did not reproduce it. Coordination and institutions inverted from 7 percent working-now to 67 percent, physical build stopped being zero, and maturity agreement between the two passes was 63 of 103 with 23 of 40 disagreements moving the same direction. The cause is an unresolved ambiguity between availability and efficacy readings of Working now, which diverge completely for institutional capability. The gradient was the most striking thing this analysis produced, which is exactly why publishing it after it failed replication would be indefensible.","runner_up":"Keep the gradient with a caveat, on the grounds that v1 applied the efficacy reading consistently. Rejected: consistency within one labeler is not reliability, and a reader cannot tell which reading produced a published number.","confidence":"high","reversal_condition":"If a forced-choice maturity rubric is written and a third independent pass reproduces an ordering, the gradient can be reinstated on that measurement."},{"phase":"phase-2b","decision":"Adjudicate relabel disagreements mechanically: agreement keeps the label as confident, disagreement takes v2 and is flagged guess. No case-by-case adjudication.","rationale":"I authored v1, so I am not a neutral adjudicator and picking winners case by case would reintroduce exactly the bias the blind pass was built to remove. v2 wins ties because it alone had the complete eight-category taxonomy. A guess flag is the honest record of two independent labelers failing to converge on the same text.","runner_up":"Adjudicate each of the 57 disagreements on the merits, which yields more confident labels but launders my prior judgement through a process that looks independent and is not.","confidence":"high","reversal_condition":"If a third independent labeler is run, majority vote across three passes would be a legitimate adjudication rule and would restore confidence to the cases where two of three agree."},{"phase":"phase-2b","decision":"Define the ai-as-object frame by a criterion (the gap would still exist if AI did not) rather than by enumeration, and add Labor-Replacing AI Could Lead to Human Disempowerment as a fourth member.","rationale":"The first revision listed three gaps instead of stating a test. A blind relabeler immediately found a fourth meeting the same description and correctly noted that by the letter of the rule it stayed ai-as-instrument. An enumeration cannot be applied to a gap nobody thought of, which is the thing a labeler has to do.","runner_up":"Keep the enumeration and extend it as cases arise, which fails the same way again on the next unanticipated gap.","confidence":"high","reversal_condition":"If the criterion admits gaps that are clearly not about AI, tighten it rather than reverting to a list."},{"phase":"phase-3","decision":"Recorded the quantum gravity gap as an honest null even though a real, published, improving quantity exists (the lower bound on the QG energy scale from photon dispersion: Fermi-LAT E_QG,1 > 7.6 E_Planck, GRB 090510, 2013; extended by LHAASO on GRB 221009A, 2024).","rationale":"The bound constrains linear-in-energy Lorentz violation, which most leading quantum-gravity programmes do not predict, and the community has no target and no shared claim about what any particular value of it would establish. A number that improves without agreement on what its improvement means is not a progress indicator for the gap. This is precisely what the Verification contested tier asserts, so recording it as an indicator would have contradicted our own tier label.","runner_up":"Record the Fermi/LHAASO bound as the indicator with confidence 'guess'. This would have been the more flattering choice — a tier-3 gap with a real number is a more interesting headline than a null — and it is why the decision is logged rather than assumed.","confidence":"medium","reversal_condition":"A published community statement, roadmap or review that names a target value for E_QG or an equivalent parameter and says what reaching it would settle. Reverse immediately if one exists."},{"phase":"phase-3","decision":"Used a 2025 arXiv preprint as the source for the retraction rate rather than a peer-reviewed paper or the Nature news analysis.","rationale":"Nature's site returns an authentication redirect and could not be fetched, and the plan forbids sourcing a number from a search snippet. The preprint gives an explicit denominator (Web of Science, 2000-2024) which the widely quoted figures do not, and arXiv is checkable by engine/validate-indicators.mjs, which verified the title. The row is marked 'guess' and its rationale states that published estimates disagree by roughly a factor of three.","runner_up":"Leave the fraud gap without an indicator. Rejected because a contested number with its dispute stated is more useful to a reader than a blank, and because the dispute is the substance of the Proxy only tier.","confidence":"medium","reversal_condition":"A peer-reviewed rate with a stated denominator becomes fetchable; replace the source and re-run the validator."},{"phase":"phase-3","decision":"Left target_value NULL on five of the six non-null indicators, populating it only where a programme or roadmap states one (brain volume, silicon scaling).","rationale":"The plan requires a target_basis for every target and rules out invented round numbers. For elapsed telescope time, materials-per-day, trial cost and retraction rate, no funder, roadmap or community body has published a figure. In two of those cases a plausible-looking number was available and rejected on inspection: the 2.5 million USD median for US-funded phase 3 trials is a different population from industry pivotal trials, and a lower retraction rate is not unambiguously better because the quantity measures detection.","runner_up":"Derive targets from the best observed case in each series. Rejected: it manufactures an improvement factor out of a sampling difference and would read as analysis rather than as the arithmetic it is.","confidence":"high","reversal_condition":"A funder or roadmap publishes a target for any of these quantities."},{"phase":"phase-3","decision":"Persisted the Phase 3 search transcript to research-log/searches/phase-3.json, and added ingest scripts so indicators, new gaps, critical paths, decisions and runs all rebuild from files.","rationale":"db/gapmap.sqlite and research-cache/ are both gitignored, by deliberate earlier design. Without this, every Phase 3-6 row and every search behind the two nulls would exist only in an untracked binary and an untracked cache, and the claim that the nulls are provable rather than asserted would be false in the repository a stranger actually receives.","runner_up":"Un-ignore research-cache/. Rejected: it commits several megabytes of third-party search results to hold a few hundred kilobytes of evidence, and the titles and URLs are the part that matters.","confidence":"high","reversal_condition":"None expected. If cache contents themselves become disputed, commit the cache."},{"phase":"phase-4","decision":"Dropped the fifth candidate gap — queue and turnaround time at shared nanofabrication user facilities — rather than shipping four with one weak.","rationale":"Five logged searches turned up no published wait-time or turnaround data for NNCI or comparable facilities, so the gap could not be grounded in an observed rate limit. It also had the weakest funding check of the five: NSF's National Nanotechnology Coordinated Infrastructure exists precisely to provide open access across 16 sites, so proposing access latency as an unowned gap would have required arguing against a programme built for it, on no evidence.","runner_up":"Include it with confidence 'guess' and an honest funding check. Rejected: the plan asks for 3-5 gaps, four are well grounded, and an ungrounded fifth costs more credibility than the count gains.","confidence":"high","reversal_condition":"Facility-level turnaround statistics become available, from NNCI reporting or a user survey."},{"phase":"phase-4","decision":"Proposed the coating thermal noise gap even though coatings research is funded, and said so in the funding check rather than around it.","rationale":"The NSF LSC Center for Coatings Research and Italy's ETIC project both fund this work as a detector subsystem inside one instrument programme. The gap proposed is the cross-instrument materials framing — the same loss angle bounds optical clocks and cavity-stabilised lasers, and no programme owns it at that level. Writing 'not clear of funding' in the funding_check and stating exactly what is funded is more useful than a claim of novelty a reader could puncture in one search.","runner_up":"Drop it as already funded. Rejected: it would discard the strongest tier-1 candidate over a framing question, and the framing is the contribution.","confidence":"medium","reversal_condition":"A programme is found that funds low-mechanical-loss optical coatings as a materials target across instrument classes."},{"phase":"phase-4","decision":"Added ai_type, maturity and tier columns to new_gaps rather than writing proposed gaps into gap_ai_types and gap_measurability.","rationale":"Those two tables hold foreign keys into gm_gaps. A proposed gap is deliberately not one of theirs, and giving it a row in the same tables would make the two indistinguishable in every downstream query and export. The plan requires new gaps to carry the same three labels; this carries them without blurring the boundary the whole project rests on.","runner_up":"Relax the foreign keys so both kinds of gap share the label tables. Rejected: it trades the clearest structural guarantee in the schema for a small convenience in the export.","confidence":"high","reversal_condition":"None expected."},{"phase":"phase-5","decision":"The Notion connector, unavailable at the start of this run, was authorised mid-run, so chain 2 is built from its named source of record after all.","rationale":"Recorded because a decision to substitute public sources was logged earlier in this same run and then reversed; the ledger should show the reversal rather than quietly drop it. Public sources gathered before the connector returned are kept and cited alongside the Notion page, since they are citable in the artifact and the page is not.","runner_up":"Proceed on public sources only. No longer necessary.","confidence":"high","reversal_condition":"None."},{"phase":"phase-5","decision":"Gave the chain 1 headline two accountings rather than one, and reported the weaker of the two as the honest figure.","rationale":"The expectation recorded in advance was that closing every cognitive link would change the total duration 'very little'. Counting JWST's seven-year science-case period as compressible cognitive work makes the saving nine and a half of thirty-two years, which is under a third but not 'very little'. Counting it as community consensus formation — which is what the milestone record shows it was — makes the saving two and a half years. Both are stated. Tuning the classification to protect the prediction would have been the easy move and would have destroyed the value of recording the prediction in advance.","runner_up":"Report only the realistic accounting. Rejected: the generous one is the number a skeptical reader would compute, and pre-empting it is worth more than winning on it.","confidence":"high","reversal_condition":"Evidence that the 1989-1996 period was analysis-limited rather than consensus-limited."},{"phase":"phase-5","decision":"Split chain 2's reviewer recruitment link into matching and willingness inside a single link rather than making them two links.","rationale":"They are sequentially inseparable — an editor cannot recruit before identifying — so two links would imply an ordering that does not exist. But the whole finding of the chain lives in the split: matching is a prediction problem AI is applicable to and, on the published evidence, not even best at; willingness is labour supply that no matching system touches. The split is carried in the blocker and rationale fields of one link.","runner_up":"Two links, 'matching' then 'recruitment'. Rejected as a false serialisation.","confidence":"medium","reversal_condition":"A venue is found where identification and invitation are genuinely separate stages with separate durations."},{"phase":"phase-5","decision":"Verified Aaron Tohuvavohu against both the export and the live site before relying on the association.","rationale":"The plan flagged it as needing checking. Confirmed twice: resource 1c3cb37e-2a00-80a1-8ddf-fb19d0b8b0ee, type Individual, is cited by the capability 'Space Telescope Factory', which is attached to the telescope gap; and the name appears in the acknowledgments list on gap-map.org/about.","runner_up":"Rely on the plan's statement. Rejected — the plan itself asked for the check.","confidence":"high","reversal_condition":"The site's acknowledgments change."},{"phase":"phase-6","decision":"Cut all three 3ie figures from the cover note rather than softening them: the count of evidence gap maps, the Development Evidence Portal totals, and the absolute-gap versus synthesis-gap terminology.","rationale":"The plan flagged them as needing verification and they did not verify. 3ie's own gap maps page states no total; their own blog posts give portal figures that disagree by roughly a factor of three (3,745 impact evaluations in one, 'more than 11,000' in another) and nothing found states 21,800 or 1,700; and neither the gap maps page nor the working paper page uses the terms 'absolute gap' or 'synthesis gap'. The substance of the distinction is on their page verbatim and is kept as a quote in the notes, so a later comparison has a defensible form.","runner_up":"Keep the figures with an 'approximately'. Rejected outright: the cover note's entire argument is that this work checks things, and an unverified number in it would be the one thing a reader could puncture.","confidence":"high","reversal_condition":"3ie publishes a stated total, or the Development Evidence Portal exposes counts to a non-JavaScript client."},{"phase":"phase-6","decision":"Rendered the two chains as HTML/SVG in the artifact and shipped the Mermaid source alongside, rather than running a Mermaid renderer in the page.","rationale":"Each diagram is eight boxes and an arrow. A client-side Mermaid runtime would have been the single largest dependency in the build, for output that is less accessible than the HTML version — which also carries the binding flag as a word in the link table, so the distinction never rests on colour. The Mermaid source is in docs/critical-paths.md and in a copy block under each chain, so nothing is lost for anyone who wants to paste it elsewhere.","runner_up":"Add mermaid as a dependency and render client-side. Rejected on build weight and accessibility, not on effort.","confidence":"high","reversal_condition":"A reviewer wants the diagrams to match Convergent's own tooling exactly."},{"phase":"phase-6","decision":"Wired engine/adjudicate.mjs into engine/rebuild.mjs after finding that a clean rebuild silently dropped every Phase 2 confidence downgrade.","rationale":"adjudicate.mjs writes confidence flags directly to SQLite and had only ever been run by hand. Because db/gapmap.sqlite is gitignored, a rebuild from files produced 5 tier guesses and 10 AI-type guesses where the published findings say 24 and 20 — including the blanket downgrade of every 'Proxy only' assignment, which is one of the project's headline results. The artifact was reading those wrong numbers before this was caught. The script is idempotent, so running it inside rebuild is safe.","runner_up":"Bake the downgrades into the label files. Rejected: the labels are what the labeler wrote, and overwriting them would erase the fact that adjudication changed something.","confidence":"high","reversal_condition":"None. This was a defect."},{"phase":"phase-6","decision":"Kept the JAMA article page as the source URL for the clinical trial indicator even though it returns HTTP 403 to programmatic clients, and recorded that in the row's rationale.","rationale":"It is the page the number was actually read off, the DOI verifies against Crossref, and a human clicking the link sees the paper. Swapping to a URL that passes an automated check but is not where the number was read would make the link check pass and the provenance worse.","runner_up":"Point at the PubMed record, which returns 203. Rejected: the PubMed abstract does not carry the IQR or the design breakdown the rationale relies on.","confidence":"medium","reversal_condition":"JAMA stops blocking, or an open-access copy with the same figures appears."},{"phase":"maturity-repair","decision":"Replaced mechanical adjudication with adjudication on the merits for the maturity field only, re-deciding 17 of 54 reviewed gaps in research-log/relabel-adjudication.json. Type is still adjudicated mechanically by taking v2.","rationale":"The rule \"on disagreement take v2\" was chosen so the author of v1 could not launder his own judgement, and it worked for that. But a rule that cannot be wrong removes the bias and every check on validity together, and v2 read \"Working now\" as availability rather than efficacy, so the rule adopted the wrong reading 25 times in one direction. It produced \"AI Could Be Misused\" as Coordination and institutions / Working now, which no reader would accept. Efficacy is now pinned in methodology/taxonomy.md with the DNA-synthesis-screening example, and each re-adjudicated entry carries its own per-gap reasoning in its note, so the judgement is auditable row by row rather than resting on a rule.","runner_up":"Keep the mechanical rule and disclaim maturity in the artifact. Rejected: the labels are wrong, not merely uncertain, and a caveat a reader cannot act on is noise. The second runner-up was to re-run a third blind pass under the pinned definition, which is the more rigorous fix and was rejected on cost — it would relabel 103 gaps to correct a defect that provably reaches only 54, and every change here landed on a gap where the two existing passes already disagreed.","confidence":"high","reversal_condition":"A third independent pass under the pinned efficacy definition disagrees with this adjudication on more than about a fifth of the 54 reviewed gaps, or David rules the other way on the escalations in research-log/maturity-repair/escalations.md, in which case the affected entries revert individually rather than the method being abandoned."},{"phase":"maturity-repair","decision":"Made agreement between the two labelling passes an overridable default in engine/apply-relabel.mjs, rather than a branch that always takes v1. An explicit entry in relabel-adjudication.json now wins even where both passes agreed, and the rebuild logs the override count separately.","rationale":"David ruled on four escalations and two of them — Ephemeral Societal Data and Inadequate Emergency Climate Interventions — were gaps both passes had agreed on, which the code had no way to express. Unoverridable agreement is the same defect as the unoverridable 'on disagreement take v2' rule, one layer down: two independent passes sharing a misreading is exactly how 'AI Could Be Misused' came out as Working now. Agreement is evidence, not proof. Logging the override count on its own line keeps it from happening quietly.","runner_up":"Edit the v1 maturity in research-log/labels/*.json so those gaps fall into the disagreement branch. Rejected: it would falsify the record of what the first pass actually said, and move the v1-v2 agreement figure from 63/103 to 61/103, a number both docs/relabel-report.md and methodology/taxonomy.md cite.","confidence":"high","reversal_condition":"The override count printed by rebuild stops being small. If explicit overrides on agreed gaps become routine rather than exceptional, the problem is in the labelling passes and not the adjudication layer, and the fix belongs upstream."},{"phase":"maturity-repair","decision":"Did not add a '5-10 years' maturity value, despite 63 of 103 gaps now sitting in '2-5 years'. Recorded as a proposal for a later pass, together with a proposal to separate the time question from the has-a-path question that 'Speculative' actually asks.","rationale":"David raised it against 'AI is Still Narrow', and the diagnostic supports him — three fifths of the map in one bucket is barely a label, and 'Speculative' conflates 'slow' with 'no known route'. But adding a fourth value means re-reviewing all 63 gaps currently at '2-5 years', because a value nobody has applied to the whole set is worse than three honest ones: those 63 would silently mean '2-5 or 5-10, unexamined'. That is a full relabel, and this branch is a repair.","runner_up":"Add the value and apply it only to the gaps this pass touched. Rejected for exactly that reason — a partially applied enum value makes the distribution unreadable, which is a worse artifact than the pileup it fixes.","confidence":"medium","reversal_condition":"A pass is commissioned with a blind second reader to re-review all 63 gaps at '2-5 years', at which point the split should be done on both axes rather than by adding one bucket to one of them."},{"phase":"maturity-repair","decision":"Recorded, and did not build, a blocked-on-adoption versus blocked-on-capability dimension. Eleven of this pass's downgrades turn on adoption rather than on whether the technique exists.","rationale":"David's note on the clinical-trials ruling. DNA synthesis screening exists and is not adopted; archiving works and permission is withheld; adaptive trials run and the field does not take them up. Maturity absorbs all of this and reports 'not ready', which is the wrong diagnosis for a funder, because the intervention for an unadopted capability is not more research. This is a genuine addition to Convergent's map rather than a correction to our own labels.","runner_up":"Encode it now as a fourth dimension on the augmentation tables. Rejected: it would need its own definition, its own labelling pass over 103 gaps and its own audit, and it would arrive in the same artifact as a repair — which is how a contribution turns into a rewrite of someone else's map.","confidence":"medium","reversal_condition":"Convergent asks what would be worth adding next, in which case this is the strongest candidate on the evidence this pass produced."},{"phase":"maturity-repair","decision":"Added docs/future-work.md and an app page at /missing/ (\"What's missing\") listing the nine pieces of open work, and did NOT regenerate app/public/data.json to go with it. The page reads every number it prints out of data.json instead of hardcoding them.","rationale":"The artifact regeneration is its own track and is meant to run once after both the maturity repair and the review gates land; regenerating a large derived JSON now would conflict with the review-gate branch, which changes indicators, new-gaps and critical-paths. Reading from data.json means the page cannot contradict the rest of the site today, and it picks up the repaired maturity distribution automatically when the artifact track runs — the middle-bucket share it cites goes from 50% to 61% with no edit to the page.","runner_up":"Hardcode the post-repair numbers in the page copy. Rejected: it would make the site contradict itself until the artifact regenerates, which is worse than being briefly out of date, and it would need a second edit later that nobody would remember to make.","confidence":"high","reversal_condition":"The artifact track runs and the page still reads wrong, which would mean a number it needs is not in the data.json summary and should be added to engine/export-artifact.mjs rather than typed into the page."},{"phase":"revision-7","decision":"Rename all eight AI-capability categories for display so every one names a kind of work, not a kind of model. Reading and synthesis, Prediction and modeling, Design search, Measurement and sensing, Running experiments, Real-time control, Physical build, Coordination and institutions.","rationale":"Six of the eight were named after the AI that would do the work and two after the work itself, so a column that means 'what stands in the way' read as 'which AI would do it'. That made Coordination and institutions look like a category error when it was one of the two naming the thing consistently. David raised it; the diagnosis is his. Applied at the export boundary in engine/export-artifact.mjs, which walks the emitted object rewriting values, object keys and category references inside rationale strings, so summary buckets, cross-tabs, chain links, per-gap labels and the CSV all move together and no render site can be missed by hand.","runner_up":"Add a second attribute naming which AI capability could accelerate each gap, and show blocker and accelerator as two sections. Rejected for now because that attribute has never been labelled: it would need a full pass over all 103 gaps plus an audit, and for coordination gaps the honest answer is often 'none'. Recorded in docs/todo.md as an open question rather than closed.","confidence":"high","reversal_condition":"One constant in engine/export-artifact.mjs. The stored enum, the CHECK constraints and every research-log judgment are unchanged, so reverting is deleting the map."},{"phase":"revision-7","decision":"AI Could Be Misused is Speculative, not 2-5 years.","rationale":"David ruled on 2026-08-25. The gap has many partial mitigations -- evaluation regimes, deployment policy, convening -- and the maturity repair had moved it to 2-5 years on the strength of those existing as mechanisms. But the capability that would actually close the gap is alignment, and a substantial part of the field holds it may not be solvable at all. A capability whose feasibility is itself contested is what Speculative is for. This is the same distinction the maturity repair was built on: availability is not efficacy.","runner_up":"2-5 years, on the grounds that partial mitigation mechanisms exist and are being built. Rejected because partial mitigations do not close this gap, and the taxonomy asks what would move the gap rather than what is being attempted.","confidence":"high","reversal_condition":"A change in the state of alignment research such that the field stops treating the core problem as possibly unsolvable. Recorded in research-log/relabel-adjudication.json for this gap."},{"phase":"revision-6","decision":"Dropped is_binding from the telescope chain entirely, and renamed what survives on the publishing chain to \"carries the cost\".","rationale":"A blind cold reviewer pointed out that the telescope chain is strictly sequential, so removing any step shortens the total and every step is on the critical path. Marking four of eight as binding claimed a distinction the structure does not contain, and the marked set turned out to be exactly the steps no AI capability acts on, which made the headline finding restate its own labelling. The textbook sense of binding needs parallel paths and neither chain has them. What survives is the cost sense on chain 2, where three of seven steps genuinely account for a disproportionate share of reviewer and editor labor. That is a claim about distribution rather than about slack, so it gets its own words.","runner_up":"Keep the vocabulary on both chains and explain the difference in a footnote. Rejected because the circularity was real rather than presentational: no wording rescues a flag that has no discriminating power on the chain it is applied to. The arithmetic that replaced it -- AI acts on 9.5 of the 32.5 years -- is stronger and needs no vocabulary at all.","confidence":"high","reversal_condition":"A chain is built with genuine parallel structure, where some steps run alongside others. is_binding then recovers its classic meaning on a time axis and methodology/critical-path.md needs a third case."}],"runs":[{"phase":"phase-0","kind":"agent","started_at":"2026-08-23T01:44:00Z","ended_at":"2026-08-23T02:30:00Z","n_units":null,"note":"Repo scaffold, schema, importer, additive guardrail, integrity report, search client, baseline snapshot."},{"phase":"phase-1","kind":"agent","started_at":"2026-08-23T02:30:00Z","ended_at":"2026-08-23T02:41:00Z","n_units":103,"note":"Full-coverage labeling of all 103 gaps across 20 field batches: outcome, AI type and maturity, measurability tier."},{"phase":"phase-2","kind":"agent","started_at":"2026-08-23T02:41:00Z","ended_at":"2026-08-23T03:00:00Z","n_units":36,"note":"Stratified blind audit across 3 independent auditors, adjudication, findings report."},{"phase":"phase-2b","kind":"agent","started_at":"2026-08-23T03:00:00Z","ended_at":"2026-08-23T03:55:00Z","n_units":103,"note":"Taxonomy revision (8th type, frame dimension), full independent blind relabel of all 103 gaps by 5 labelers, mechanical adjudication, withdrawal of the maturity gradient."},{"phase":"phase-3","kind":"agent","started_at":"2026-08-23T17:46:30Z","ended_at":"2026-08-23T17:55:49Z","n_units":8,"note":"progress indicators: 8 rows across 4 tiers, 2 honest nulls, 29 logged searches. Local-session setup before this point (rebuild, network verification, ingest plumbing) is not counted."},{"phase":"phase-4","kind":"agent","started_at":"2026-08-23T17:55:49Z","ended_at":"2026-08-23T18:02:03Z","n_units":4,"note":"new gaps: 4 proposed, 1 candidate dropped; dedup over the full export plus funding checks."},{"phase":"phase-5","kind":"agent","started_at":"2026-08-23T18:02:03Z","ended_at":"2026-08-23T18:09:09Z","n_units":2,"note":"two critical paths, 15 steps, plus the cross-field intersection. Expectations committed before the analysis in a separate commit."},{"phase":"phase-6","kind":"agent","started_at":"2026-08-23T18:09:09Z","ended_at":"2026-08-23T18:36:38Z","n_units":1,"note":"artifact (Next.js static export), CSV keyed on their id and slug, findings summary, cover note draft; found and fixed the adjudication rebuild defect."},{"phase":"revision-1","kind":"agent","started_at":"2026-08-23T18:39:22Z","ended_at":"2026-08-23T19:52:15Z","n_units":null,"note":"rebuilt around the argument in Convergent's voice, after a blind cold review"},{"phase":"revision-2","kind":"agent","started_at":"2026-08-23T19:52:15Z","ended_at":"2026-08-23T20:42:25Z","n_units":null,"note":"second cold review: replotted the chart, fixed three dead links, stated three holes"},{"phase":"revision-3","kind":"agent","started_at":"2026-08-23T20:42:25Z","ended_at":"2026-08-23T21:23:15Z","n_units":null,"note":"cut page one to a fifth, split into five pages, rebuilt the map on their components"},{"phase":"revision-4","kind":"agent","started_at":"2026-08-23T21:23:15Z","ended_at":"2026-08-23T21:45:30Z","n_units":null,"note":"new opening, all 103 outcome sentences rewritten, indicators page"},{"phase":"revision-5","kind":"agent","started_at":"2026-08-23T21:45:30Z","ended_at":"2026-08-23T22:18:22Z","n_units":null,"note":"third cold review: two reasoning errors corrected, outcomes deticked, attributes page"},{"phase":"revision-6","kind":"agent","started_at":"2026-08-23T22:18:22Z","ended_at":null,"n_units":null,"note":"robotics as the AI analogue for physical build, time ledger, drafting cost corrected"},{"phase":"all","kind":"human-review","started_at":"2026-08-23T18:39:22Z","ended_at":"2026-08-23T22:43:00Z","n_units":null,"note":"David's own time reviewing, directing and rewriting, over the same window as the revision cycle. His estimate rather than an instrumented figure, and roughly equal to the agent's. The build phases above had no human review at all."}]}