# P(DOOM) INDEX — METHODOLOGY v0.1 Governing methodology for pdoomcount.org. Dated 2026-09-16. Clinical register: this document is a protocol, not an argument for a conclusion. v0.1 integrates `docs/MISSION.md`, the locked forks in `docs/METHODOLOGY-DRAFT-v0.md`, the verified seed ledger in `docs/RESEARCH-PASS-v1.md`, and the adopted findings in `docs/CRITIQUE-R1.md` / `docs/CRITIQUE-R1-VERDICTS.md`. The v0 draft remains in the repo as history; do not build from it. Companion extract: `docs/SOURCE-LEDGER-TIER-A.md` (verified values plus v0.1 scoring columns). **This document does not publish a live index number.** It states how a number is computed. Static snapshot v1.0 (as-of 2026-09-16) lives in `docs/SNAPSHOT-v1.md` and `snapshots/v1.json`; the frontpage mockup renders that snapshot. The design-preview figures that used to sit on the mockup (headline 4.7%, G2 1.2%, G3 3.9%, conditions ×1.08) were **ILLUSTRATIVE** and are not v0.1 outputs. --- ## 0. Document control | Field | Value | |---|---| | Version | Methodology v0.1 | | Rubric version | R0.1 (rule IDs in §12) | | Status | Governing for any static v1 snapshot compute | | Forks | Locked — see §1. Do not reopen in this version. | | Snapshot model | Static v1. Automation is v2. | | Hash stamp | R0.1 freeze: commit `051565a`, blob `0fae6bbcc552eec8ce70b36b0ecbab2c5a1c85dd` (§12.8). Rule-ID rows cite **R0.1**. A later edit to §11–§13 requires **R0.2** and a visible discontinuity. | Site judgments in this text are labeled **[SJ-…]**. Verified source facts are labeled **[VF]** and point at the research pass. Mixing the two unlabeled is a protocol error. --- ## 1. Locked forks (do not reopen) Ruled 2026-09-15 by the operator. The critique attacked M1 as instructed; amendments fold **around** these forks. 1. **Headline gauge = G1 catastrophe** [LOCK: G1]. G2 extinction is a displayed sub-gauge. G3 is displayed and is not folded silently into G1. 2. **News mechanism = M1 bounded conditioning multiplier** [LOCK: M1], with a public event ledger, per-category bounds, and movement caps. v0.1’s work is to make that multiplier attributable, bounded against forecaster disagreement, quantized, and symmetrically falsifiable. 3. **Source universe includes high outliers** [LOCK: include]. They are shown individually. UNVERIFIED or paraphrased outliers stay in the ledger and **do not score** (§5). Inclusion in the universe is not inclusion in the statistic. 4. **v1 ships as a static, fully-documented snapshot** [LOCK: static]. The live pipeline is a v2 work item. Front-page language says snapshot, as-of date, next refresh date. Also ratified via MISSION and unchanged: public site, no ads, no monetized fear, clinical language. --- ## 2. What the site publishes An **index of published judgments, conditioned by a documented news rubric**. Not a measurement. The front page says so. Published surfaces, each with uncertainty: - **G1 headline** — point, source band, conditioning envelope, unconditioned aggregate, conditions comparison against forecaster disagreement, down-move eligibility. - **G2** — same machinery, own mapping, own minimum-n outcome. - **G3** — labeled constructed estimate (§4.3). Never presented as a conditioned aggregate of published forecasts. - **Source ledger** — every candidate, scoring or not. - **Event ledger** — every multiplier movement, including zeros, clips, decays, resets, and down-triggers. - **Corrections log**. - **Method page** — this protocol, leave-one-out table, weighting-scheme range, mapping table, genesis row. Why this shape: the chance that AI causes a terminal catastrophe cannot be observed. A single author’s number is one credence. What can be built honestly is a reproducible aggregation of named estimates, with every site judgment written down, and with current conditions applied only inside a bound that is smaller than the disagreement already on the record. --- ## 3. Honesty rules, operationalized From MISSION, made checkable: 1. **Index, not measurement** — front page sentence, permanent. 2. **Attributable** — every scored value cites a ledger ID; every movement cites an event ID and a **rule ID** from R0.1; every site judgment has an `[SJ-…]` tag. 3. **Uncertainty always displayed** — a point with no band is not publishable. The band’s components are labeled. What the band excludes is stated (§9.3). 4. **Symmetric falsifiability** — “what would move this down” is not a caption. Numeric down-triggers in §13 have the same force as up-steps. A **down-move eligibility** line sits beside the number. 5. **No bombast** — clinical language; absolute percentage-point deltas; resolution floor (§14.3). 6. **Mundane-deployment path is first-class** — G3 is always on the page; categories C2 and C7 exist; they are capped so they cannot dominate the multiplier (§4.4, §11.4). --- ## 4. The gauges All gauges are AI-caused. Horizon for G1 and G2 mixtures: **2100**, with as-stated values shown beside any standardized use (§6.2). ### 4.1 G1 — Catastrophe (headline) [LOCK: G1] An AI-caused event that kills **≥10% of humanity** or **irreversibly collapses industrial civilization**, by 2100. This is broader than literal extinction. It is meant to capture both an alignment-lab failure and a mundane deployment cascade that ends in collapse. Those two routes are not extra gauges inside G1; they are covered by the definition. G3 exists so the second route is not only a clause in G1. ### 4.2 G2 — Extinction Literal human extinction from AI, or the FRI operationalization where it is the source’s own definition: **global population < 5,000**. Horizon for the mixture: 2100. Near-term extinction statements are a separate panel (§6.2). ### 4.3 G3 — Deployment cascade (constructed estimate) **[SJ-G3-CONSTRUCT]** G3 is the founding thesis as its own gauge: premature deployment of insufficiently-verified AI into critical systems (grid, military command, health, finance, logistics), producing compounding failures that neither market nor regulator catches in time. Collapse-shaped, not extinction-shaped by necessity. Always displayed. Never folded silently into G1. The verified Tier A table contains **no published forecast of this event** **[VF: RESEARCH-PASS-v1]**. Therefore G3 has **n = 0** scoring sources. **v0.1 rendering rule:** G3 is a **labeled constructed estimate**, not a conditioned aggregate. It must not look like G1/G2. Allowed on the G3 surface: - a short scenario narrative; - a Carlsmith-style factor list (deployment into named system classes; verification lag; liability lag; compounding; terminality), each qualitative or sourced as *indicator*, not as P(G3); - pointers to event-ledger rows in C2 and C7, which instrument the same thesis as *conditions*, not as a G3 probability. Not allowed on the G3 surface: - a percentage presented as an aggregate of published forecasts; - `headline_G3 = aggregate(TierA, TierB) × conditioning(TierC)`; - silent reuse of the G1 number. If a later version shows a G3 percentage, the label **SITE-CONSTRUCTED, NOT AN AGGREGATE** is mandatory, the construction recipe is a versioned appendix, and n=0 remains on the page until a primary numeric forecast of this definition exists. The mockup’s **3.9%** is illustrative only and is not that recipe. ### 4.4 Overlap (G1 collapse limb, G3, C2, C7) **[SJ-OVERLAP]** The mundane-deployment thesis appears in three places by design: G1’s collapse limb, G3’s whole definition, and scoring categories C2 and C7. That is overweighting if left unnamed. Controls: - G3 is not a second G1 number (section 4.3). - C2+C7 together may use at most **2/7 of the overall conditioning log-span** (§11.4), equal to two categories’ share of seven, not more. - G1’s mixture is published forecasts of catastrophe/existential catastrophe, not a constructed G3. The thesis is first-class. It is not allowed to count three times as if it were three independent quantitative inputs. --- ## 5. Scoring eligibility — primary numeric utterances **[SJ-UTTER]** Only a **primary numeric utterance** scores. A primary numeric utterance is a percentage, odds, or clearly numeric interval spoken or written by the named person or instrument, in a citable primary (book, paper, survey table, recorded interview, dated post). Excluded from every statistic, retained in the ledger with flags: - third-party conversions (LeCun’s asteroid comparison converted to <0.01% **[VF]**); - UNVERIFIED high percentages without a pinned primary quote (Yudkowsky **[VF]**); - qualitative institutional language with no number (International AI Safety Report **[VF]**); - a different question treated as if it were this one (Hinton “50%” good-vs-bad **[VF]**). LOCK include-outliers means those rows are **shown**, not deleted. It does not mean a journalist’s paraphrase becomes a data point. Interval utterances (Hinton “10% to 20%”) score as intervals, not as a silently chosen midpoint. Bound utterances (Hubinger “>10%”) score as bounds, not as 10%. Carlsmith’s later “>10%” is a bound companion to the main-text product ~5%; the mixture point for that row is the main-text **5%** **[SJ-CARL-POINT]**, because it is the only exact product; the bound is displayed beside it. --- ## 6. Normalization Every source probability is mapped to a gauge **before** it enters a mixture. Where the definition does not map cleanly, the row stays in the ledger with a reason and is excluded from that mixture. ### 6.1 Gauge mapping **[SJ-MAP]** Mapping is a site judgment. It can move the headline more than a typical source. It is therefore a published table with decision-maker (this document), an ambiguity flag, and a sensitivity line — not an implicit cleanup. | Rule | Native G1 (catastrophe / existential catastrophe / collapse-or-worse) | Native G2 (literal extinction / FRI <5,000) | Ambiguous | Does not score | |---|---|---|---|---| | Examples | XPT catastrophe 2.13% / 12%; Carlsmith existential catastrophe; Ord existential risk from unaligned AI; AI Impacts extinction-*or*-severe-disempowerment (flagged) | XPT extinction 0.38% / 3%; Hinton human extinction; Hubinger “kill all humans” | Bengio “catastrophic” / p(doom) with no operationalization | LeCun conversion; Yudkowsky UNVERIFIED | | Mixture | May enter **G1-2100** if horizon class allows | May enter **G2-2100** if horizon class allows | Ledger; not in 2100 mixtures until a later mapping version | Ledger only | **No face-value dump of G2 numbers into G1.** Extinction ⊆ catastrophe, so P(G1) ≥ P(G2) for nested events in one mind, but mixing Hinton’s 10–20% *extinction* with XPT’s 12% *catastrophe* as if they named the same event is the inflation the critique named. **Published G2→G1 factor (sensitivity only):** XPT is the only verified instrument that reports both definitions on one panel **[VF]**. - Superforecasters: 2.13 / 0.38 ≈ **5.61** - Domain experts: 12 / 3 = **4.00** **[SJ-M-FACTOR]** v0.1 records **m = 4**, the domain-expert ratio, as the sensitivity factor. It is **not** applied to the headline. The required sensitivity line is: “G1 if native-G2 rows were up-mapped by m=4” versus “G1 if native-G2 rows were included at face value” versus **the actual headline (native-G1 rows only)**. Per-source mapping lives in `docs/SOURCE-LEDGER-TIER-A.md`. Changing a map is a methodology version bump. ### 6.2 Horizon standardization Verified horizons **[VF]**: Hubinger 10y, Hinton ~30y, Carlsmith ~2070, XPT 2100, Ord ~100y, AI Impacts open / 100y variant, Bengio unspecified. They are not all “2100.” **[SJ-HORIZON]** v0.1 does **not** invent a hazard rate. Converting 10-year probabilities into 2100 probabilities would be a larger judgment than most sources. Two columns are always shown: **as-stated** and **mixture-eligible**. Horizon classes: | Class | Rule | Mixture | |---|---|---| | **H-century** | Source horizon ∈ [2070, 2100], or “this century” / 100 years, or a 100-year variant of the same instrument | Eligible for the matching 2100 gauge | | **H-near** | Source horizon < 50 years | **Excluded** from 2100 mixtures. Shown on a near-term panel as-stated. Inequality note: for an absorbing catastrophe, P(by t_near) ≤ P(by 2100); v0.1 does not use that inequality to impute a 2100 point. | | **H-unspecified** | No calendar horizon | Excluded from 2100 mixtures. Ledger only. | As-stated and mixture-eligible values sit side by side on the method page. This standardization rule is **[SJ-HORIZON]**, not a source claim. ### 6.3 Minimum-n **[SJ-MIN-N]** | n scoring rows mapped to the gauge and horizon class | Display | |---|---| | n ≥ 3 | Full aggregate (headline statistic + band) | | n = 2 | **Thin panel**: both numbers and an unweighted midpoint, labeled “thin panel, not a full aggregate.” No IQR presented as if it were a sample spread. | | n < 2 | **Insufficient sources.** No point estimate. | | n = 0 published forecasts | G3 path: constructed estimate (§4.3). | v0.1 expected membership (not a computed headline): G1-2100 has five scoring rows (A-XPT-SF catastrophe, A-XPT-DE catastrophe, A-CARL-2022, A-ORD-2020, A-AII-2024) → full aggregate eligible. G2-2100 has two scoring rows (the two XPT extinction medians) → thin panel. G3 has zero → constructed. --- ## 7. Tiers ### 7.1 Tier A — published forecasts (backbone) Named, dated, cited. Seed: `docs/SOURCE-LEDGER-TIER-A.md`. Each scoring row carries: value, gauge map, horizon class, family, weight, COI note, utterance type. Weights are published here and **fixed between quarterly reviews**. They are not tuned to hit a preferred headline. ### 7.2 Tier B — live aggregates (slot; not in the v1 snapshot) Metaculus community questions in the AI catastrophe/extinction family, and prediction markets where active. The research pass left question IDs and values **UNVERIFIED**. **[SJ-TIERB-V1]** Static v1 **excludes** Tier B from the mixture. The slot remains in the protocol. When a later version verifies IDs, each series needs: question ID, operational definition, horizon, liquidity/volume, and a **two-tier flag** if it overlaps Tier A people or questions. Markets (play-money or thin books) are **out of v1**. A future recency function, if any, must be numeric and versioned; v0.1 defines none, so none is applied. ### 7.3 Tier C — news-conditioning rubric (M1) A fixed set of indicator categories. Events move G1/G2 **only** through pre-declared rules, quantized steps, and bounds. G3 does not receive this product as if it were an aggregate. Categories (R0.1 — seven, same list as the draft, now with declared sign and precedence): | ID | Category | Sign (declared) | Role | |---|---|---|---| | C1 | Frontier capability leaps | Up when an *external* instrument clears a leap threshold. Vendor-only claims do not score. | Capability | | C2 | Deployment into sensitive systems | Up for premature/sensitive adoption; down for documented rollback | Deployment-cascade engine (capped with C7) | | C3 | Governance / regime change | Down for binding constraints; up for weakening or collapse of constraints. Lab press releases are not “governance.” | Governance | | C4 | Safety / interpretability research | Down for independent replicated results; up for independent stall. Lab safety *spend* alone does not score (ambiguous). | Safety | | C5 | Racing dynamics | Up for verified race intensification; down for verified coordination that reduces the race | Race | | C6 | Incident record | Up for a qualifying incident; down for a declared clear window | Incidents | | C7 | Autonomy dependency | Up for measured handoff of critical operations without human-in-loop verification; down for restored verification | Deployment-cascade engine (capped with C2) | C2 and C7 are the founding thesis’s live instrumentation. They are not a third G3 probability. --- ## 8. Weights ### 8.1 Class weights and the survey–superforecaster tie Draft v0 loaded “surveys and superforecasters” as the heaviest classes. Those classes disagree by a large factor (XPT superforecasters 0.38% extinction vs domain experts 3% vs surveys 5–10% **[VF]**). Whichever intra-pair tie-break wins dominates the headline. **[SJ-TIE]** v0.1 sets **equal class weight** for the current AI Impacts wave and the XPT family as a whole (both **3.0**). That is the tie-break: equality. Residual disagreement is shown as a **weighting-scheme range**, not absorbed into a silent winner. Calibration is **not** used as a cross-task license for the superforecaster panel to dominate; both XPT panels answered the same AI questions in the same tournament, and that is the comparability note. | Class | Class weight | Task-comparability note | |---|---|---| | Broad expert survey (current wave) | 3.0 | Many respondents; question is extinction-or-disempowerment, not FRI catastrophe. | | XPT family (both panels combined) | 3.0 | Operational FRI definitions; two panels, one tournament. | | Structured analysis | 2.0 | Long-form credence; not a population median. | | Named individual statement | 1.0 | One mind; modest weight; still visible because outliers are included. | ### 8.2 Instrument-family caps **[SJ-FAM]** Correlated rows are not independent. - Combined weight inside a family **≤** the class weight of that family. - FAM-AII: 2024 in the mixture at 3.0; 2023 weight 0.0 in the mixture; 2023 remains for substitute-wave sensitivity. - FAM-XPT: superforecasters 1.5 + domain experts 1.5 = 3.0. - A row that also appears in Tier B (later) is flagged two-tier and does not collect both weights. Per-source mixture weights: see the ledger. Sum of G1-2100 mixture weights under v0.1: 1.5+1.5+2.0+2.0+3.0 = **10.0**. ### 8.3 Named-statement line (required) **[SJ-NAMED-LINE]** Named statements can sit 2–5× above the survey class and pull a median even at small weight. The method page always shows: - headline **with** the named-statement class (for 2100: that class is empty under §6.2, so the line reads “no named-statement rows in this mixture”); - headline **without** that class; - the **difference in percentage points**. Near-term named statements (Hinton, Hubinger) appear on the near-term panel, not as a quiet tug on G1-2100. --- ## 9. Aggregation Per gauge that meets minimum-n: ``` unconditioned = W( {scoring rows mapped to this gauge and horizon class} ) headline = unconditioned × conditioning conditioning = clip( exp( Σ_c log m_c ) ) # §11 ``` `W` is defined next. The unconditioned aggregate is a **permanent second displayed line**. Conditioning is its own line. A reader can always separate the expert baseline from the news adjustment. ### 9.1 Headline statistic **[SJ-STAT]** The headline is the **untrimmed weighted median** of mixture-eligible values, weights as in §8. - **Untrimmed** because LOCK include-outliers forbids quietly trimming the same outliers the site promised to keep. There is no trim rule. - The **weighted mean** is shown beside the median. It is not the headline. - “Which is the number?” — the untrimmed weighted median. The mockup’s “trimmed weighted median” caption is illustrative and predates this rule; it must be corrected on the live page at snapshot build, not by relabeling the mockup in this docs pass. Weighted median: order the scoring values; the headline is the smallest value for which cumulative weight is at least half the total weight (standard weighted-median convention). Ties and even splits: the midpoint of the two central weighted observations. State the convention on the compute printout. ### 9.2 Leave-one-out and other required sensitivities On the method page, every published snapshot includes: 1. **Leave-one-out:** headline with each scoring row removed, one at a time. 2. **Substitute-wave:** G1 with A-AII-2023 (5%) in place of A-AII-2024 (10.0%). 3. **Weighting-scheme range:** v0.1 weights; equal weights; survey-only; XPT-only; structured-only. 4. **Mapping lines:** native-G1 only (headline); native-G2 at face value added; native-G2 up-mapped by m=4 added. 5. **Named-statement on/off** (§8.3). 6. **Unconditioned vs conditioned.** Small n makes the weighted median jumpy. Leave-one-out is how that is shown instead of pretended away. ### 9.3 The band **[SJ-BAND]** The source band is computed **only over same-gauge-mapped, mixture-eligible rows**. It is not an IQR dumped across extinction, disempowerment, and catastrophe as if they were one variable. | n in the mixture | Source-spread display | |---|---| | n ≥ 4 | Interquartile range of the mixture values (unweighted across rows for the spread display, so a heavy row cannot hide the disagreement). Point remains the weighted median. | | n = 3 | Full range (min–max), labeled “range, n=3,” not IQR. | | n = 2 | Thin panel: both points. No band costume. | | n < 2 | No number. | **Band composition (labeled):** - **Included in the source band:** disagreement among mixture-eligible rows for that gauge. - **Not included in the source band:** mapping uncertainty, weighting-scheme uncertainty, horizon-standardization uncertainty, conditioning. - **Conditioning envelope (separate, required):** unconditioned point × overall bound [0.65, 1.54], drawn and labeled so the largest single judgment in the system is inside the picture (critique A14). - **Mapping-induced shift (separate line):** headline versus the face-value and m=4 counterfactuals. After conditioning, the source band is multiplied by the **current** combined multiplier so the displayed IQR/range moves with conditions. The envelope uses the **bounds**, not the current multiplier. Both are on the page. --- ## 10. Conditioning bound, derived from forecaster disagreement M1 stays. The overall bound is the **primary** judgment. Per-category caps are derived so their worst-case product equals that overall bound. Decorative per-category caps that compose to ×4.77 while an overall ×1.6 quietly does all the work are not R0.1. ### 10.1 Reference disagreement **[VF]** Same instrument, same PDF, same horizon, two definitions: | | Superforecasters | Domain experts | Ratio D | |---|---|---|---| | AI catastrophe (G1-matched) | 2.13% | 12% | 12 / 2.13 ≈ **5.63** | | AI extinction (G2-matched) | 0.38% | 3% | 3 / 0.38 ≈ **7.89** | v0.1 uses the **G1-matched** pair as the reference for the G1 multiplier, because G1 is the headline. D_ref = 5.63. A ~4% baseline moved by the draft’s 0.7–1.6 span becomes 2.8–6.4% — a swing on the same order, in percentage points, as the 0.38–3 extinction dispute. The site’s job is to aggregate other people’s judgments. News may update for recency; it may not replace that dispute. ### 10.2 Derivation **[SJ-BOUND]** Principle: the **log-span** of the overall conditioning bound is at most **half** the log-span of D_ref. ``` D_ref = 12 / 2.13 ≈ 5.6338 S_max = √D_ref ≈ 2.3736 ``` Log-symmetric about 1 (equal up and down in ratio; critique A3): ``` H = √S_max ≈ 1.5406 → rounded 1.54 L = 1/H ≈ 0.6491 → rounded 0.65 overall multiplier ∈ [0.65, 1.54] span H/L ≈ 2.37 ``` Rounding to two decimals is **[SJ-BOUND-ROUND]**. Asymmetry 25% up / 15% down is withdrawn. If a later version wants alarm-side asymmetry, it needs a named justification and a rubric bump; v0.1 has none. ### 10.3 Permanent comparison line Beside the number, always: > Conditions currently multiply the aggregate by **×M** (span allowed: > ×0.65–×1.54, ratio 2.37). Reference disagreement (XPT catastrophe, > superforecasters vs domain experts) is **5.63×**. Conditions move the > number **less / comparably / more** than that disagreement: > compare |M − 1| against (√D_ref − 1), and compare M to D_ref. Worked wording at M = 1.00: “conditions currently move this **less** than the forecaster disagreement (neutral multiplier; disagreement remains 5.63×).” This line is not decorative. If a snapshot cannot print it, the snapshot is not publishable. ### 10.4 Per-category bounds, derived from the overall bound Seven categories. Worst-case product equals the overall bound: ``` 1.54^(1/7) ≈ 1.0636 0.65^(1/7) ≈ 0.9403 ``` Per-category multiplier ∈ **[0.94, 1.06]** before quantization. The overall clip remains primary if arithmetic ever exceeds it. **Clip order** when the overall bound binds: categories are listed by descending |log m_c|; the suppressed log-amount is written on the event row. The draft’s example steps 1.10, 1.15, 1.20, 1.25 belonged to a decorative 1.25 per-category cap. They are **not** R0.1 steps. Using them would make the overall bound the only real constraint again. --- ## 11. Multiplier mechanics ### 11.1 Quantized steps **[SJ-QUANT]** Default is **no movement** (1.00). Allowed per-category values: ``` {0.94, 0.97, 1.00, 1.03, 1.06} ``` There is no 1.17. Intermediate wishes round **toward 1.00** (the default). Evidence needed for each step is in §12.5. Worst-case 1.06⁷ ≈ 1.50, inside the 1.54 overall cap; 0.94⁷ ≈ 0.65. ### 11.2 Combination: log-sum, then one cap (order-independent) **[SJ-LOGSUM]** Do not multiply in an event-order sequence. ``` M_raw = exp( Σ_{c=1..7} log m_c ) M = clip( M_raw, 0.65, 1.54 ) ``` If a clip hits, ledger the suppressed amount `log(M_raw) − log(M)` and the clip-order list. Application order of *events* therefore cannot change M except via timestamps that the decay function already uses. ### 11.3 Cap precedence Numeric, in this order (later binds harder): 1. Quantized step on the event (§11.1). 2. C2+C7 share cap (§11.4). 3. Per-quarter cap: combined M may not move by more than a **log-symmetric factor 1.15** within a calendar quarter (up to ×1.15 or down to ×1/1.15 from the quarter-open M). **[SJ-QCAP]** 4. Overall bound [0.65, 1.54]. Whichever cap truncates, the row logs the **suppressed amount**. Mid-quarter reviews use the event’s **effective date** (publication date of the primary source), not the date a human typed the row, so logging delay is not an arbitrage. ### 11.4 C2+C7 share cap **[SJ-C27]** Let L_all = log(1.54/0.65). Then ``` |log m_C2 + log m_C7| ≤ (2/7) × L_all ≈ 0.246 ``` so m_C2 × m_C7 stays inside about **[0.78, 1.28]**. Excess is clipped toward 1.00 on C2 and C7 equally in log space, then logged. This is the numeric control on thesis overweight inside M1. ### 11.5 Reference period and the meaning of 1.00 **[SJ-REF]** Neutral (all m_c = 1.00, M = 1.00) means: *no news-conditioning relative to the information environment of the scored Tier A set as a whole.* Reference window: **2020-01-01 through 2026-09-15** (Ord through the research pass). Events inside that window do not score unless they are clearly *after* the relevant source’s own information set *and* still news relative to the set median date; v1 snapshot default is **M = 1.00 at genesis** with an empty qualifying event list, unless the genesis pass explicitly scores post-window or post-median events. The unconditioned aggregate is always shown, so “neutral” cannot hide the baseline. ### 11.6 Decay and reversion **[SJ-DECAY]** Multipliers ratchet without a decay rule. Each category reverts toward 1.00 in log space with **half-life 180 days**, floor/ceiling 1.00 (decay never crosses neutral). Effective daily: ``` log m(t) = log m(t0) × (1/2)^{Δt / 180d} ``` then re-quantized toward 1.00 onto the allowed set. Decay emits a ledger row on each snapshot date (static v1: one decay-to-date row at compute, not a theater of daily ticks). ### 11.7 Reset on Tier A refresh A Tier A refresh **does not** silently reset M to 1.00. If a reset is intended (new reference window), it is an event row **RESET-*** citing **R-RESET**, with before/after M and a new [SJ-REF] window. No reset, no row, no change. ### 11.8 External instruments for C1; vendor-only bar **[SJ-C1-EXT]** C1 scores only if an **external** instrument moves: - a named independent task-horizon series (e.g. METR-style time-horizon) **doubles** relative to the last scored C1 baseline, or - a named independent benchmark family shows a comparable leap, replicated or reported outside the vendor’s own blog. Vendor self-reports, demo videos, and un-audited evals **do not score**. They may be non-scoring annotations. --- ## 12. Rubric R0.1 — frozen rulebook This section is the versioned rulebook the draft lacked. Ledger rows cite a **rule ID**. Changing eligibility, precedence, signs, steps, or down-triggers requires **R0.2** and a visible history discontinuity on the method page. ### 12.1 Eligibility checklist (every event) An event scores only if **all** of the following hold. Failed checks produce a **non-scoring** ledger row (still public) with the failed item named. | Check ID | Requirement | |---|---| | E1 | **Primary source**: first-party document, official dataset, statute, or named-instrument series. Not a roundup of another outlet’s roundup. | | E2 | **Dated**: publication or occurrence timestamp. | | E3 | **One fact, one row**: five outlets reporting one fact are citations on one row, not five deltas. Split/merge: if later reporting shows two facts, split with a correction; if two rows were one fact, merge and retract the extra delta. | | E4 | **Quantified or thresholded**: meets a named threshold in §12.5. Vibes do not score. | | E5 | **Single category**: exactly one scoring category via §12.2. Other categories may be annotations (non-scoring). | | E6 | **Rule ID**: the row names the R0.1 rule that authorizes the step. | | E7 | **Not vendor-only** if the category is C1. | | E8 | **Rubric version** R0.1 (or a later frozen ID). | ### 12.2 Precedence (one event, one category) First matching wins. Lower items become annotations. 1. **C6** if the event is an incident (harm, near-miss, or documented failure in deployment or evaluation at the severity threshold). 2. **C2** if it is a deployment into a named sensitive system and not already scored as C6. 3. **C7** if it is a measured autonomy-without-verification share/handoff and not already C2/C6. 4. **C1** if it is a capability leap on an external instrument. 5. **C5** if it is racing (compute buildout, alliance fracture, Manhattan-style program) and not already C1. 6. **C3** if it is governance/regime change. 7. **C4** if it is safety/interpretability research results. ### 12.3 Direction (sign) Declared below. A sign flip is a rubric version bump, not a clever row. | Cat | Up | Down | |---|---|---| | C1 | External leap threshold met | Not from a failed demo. Down via decay or D-triggers elsewhere. | | C2 | Premature/sensitive deployment documented | Documented rollback / pause covering the same class | | C3 | Constraints weaken or die | Binding constraints in force (see D-GOV-TREATY) | | C4 | Independent stall or replicated backslide | Independent replicated safety/interpretability result | | C5 | Verified race intensification | Verified coordination that reduces race | | C6 | Qualifying incident | Clear window (D-INC-CLEAR) | | C7 | Measured increase in unverified autonomy | Restored human-in-loop requirement | ### 12.4 Rule IDs (movement) | Rule ID | Effect | |---|---| | R-NULL | Eligibility failed or default; m_c stays 1.00 relative to last state (no step). | | R-UP-1 | One up-step: 1.00→1.03, 1.03→1.06, already 1.06→clip, log suppressed. | | R-UP-2 | Two up-steps from 1.00 to 1.06 when §12.5 “strong” clause is met. | | R-DN-1 | One down-step, symmetric magnitudes. | | R-DN-2 | Two down-steps from 1.00 to 0.94 when a strong down-clause or a §13 trigger’s strong form is met. | | R-DECAY | Half-life reversion toward 1.00. | | R-CLIP | A cap in §11.3–11.4 truncated the step. | | R-RESET | Reference-window reset (§11.7). | | R-D-* | Mechanical down-trigger fired (§13). | Every scoring row: `event_id, source, date, category, rule_id, delta, m_before, m_after, M_before, M_after, suppressed, rationale`. ### 12.5 Evidence steps (what 1.03 vs 1.06 means) **[SJ-EVID]** Default R-NULL. The burden is on movement. | Step | Evidence | |---|---| | ±0.03 (R-UP-1 / R-DN-1) | One eligible event; independently confirmable; meets the category’s ordinary threshold (below). | | ±0.06 (R-UP-2 / R-DN-2) | Ordinary threshold **and** independent confirmation from a second primary, **or** a §13 named trigger. | Ordinary thresholds (R0.1): - **C1:** external task-horizon doubling, or equivalent named-benchmark leap (§11.8). - **C2:** documented adoption of frontier-class or safety-critical AI into grid, military command, clinical decision, core finance plumbing, or logistics control, with verification/liability gap named in the primary. - **C3:** statute, treaty, binding agency rule, or their repeal/expiry; not a blog post. - **C4:** paper with independent replication, or a pre-registered independent replication failure of a previously scored up-result (down would be elsewhere if we had scored that result; v1 may have none). - **C5:** public, quantified compute/alliance move of frontier scale; “we are racing” rhetoric does not score. - **C6:** incident in AIID / OECD / equivalent with severity ≥ **S2** (named harm or safety-critical near-miss; S1 chatter is non-scoring). Exact S2 codebook is a v1 compute annex; until that annex exists, C6 does not score upward (down-triggers may still run on the empty clock). - **C7:** measured share, mandate, or contract language handing a named critical operation to AI without human-in-loop verification. ### 12.6 Negative-evidence windows (mechanism, not prose) Loud events are discrete. Quiet good news is slow. Without clocks, M1 only moves up. | Clock | Rule | |---|---| | INC-CLOCK | Days since last C6 S2+ incident. At **365 consecutive days**, fire **D-INC-CLEAR** (R-DN-1 on C6). Retriggers every additional 365 clear days, subject to caps. | | GOV-CLOCK | If no qualifying C3 event (either sign) for 365 days, **no automatic move** (inaction is not a treaty). Governance down-moves require D-GOV-TREATY or an eligible C3 down-event. | | C1-CLOCK | If no C1 leap for 365 days, R-DECAY continues; no extra down-step. Capability not leaping is not negative evidence of safety. | The incident-free clock is the one the critique demanded as a mechanical peer to incident up-moves. ### 12.7 Hash stamp and discontinuity R0.1 is this §12 plus §11 and §13. After the commit that adds this file: - record `git hash-object docs/METHODOLOGY-v0.1.md` as **BLOB-R0.1**; - record the commit SHA as **COMMIT-R0.1**; - print both on the method page of any snapshot. Any edit to §11–§13 after that stamp is **R0.2** even if someone forgets to rename it. Snapshots must not mix event rows from two rubric IDs without a discontinuity marker on the sparkline. ### 12.8 Stamp R0.1 rule text is the §11–§13 bytes in commit `051565a` (blob below). This stamp table is metadata pointing at that freeze; editing only this subsection does not bump the rubric. Launch freeze on `main` still requires the operator to accept the stamp. | | | |---|---| | BLOB-R0.1 | `0fae6bbcc552eec8ce70b36b0ecbab2c5a1c85dd` (`git hash-object` of `docs/METHODOLOGY-v0.1.md` at COMMIT-R0.1) | | COMMIT-R0.1 | `051565a5751dfe3e87cb60298da91c6d3d26b9ea` | --- ## 13. Mechanical down-triggers Prose “what would move this down” is not a commitment. These triggers fire **unconditionally** when their predicates are met, with the same quantized force as up-steps. They do not wait for a human to feel that good news counts. Beside the number, a **down-move eligibility** line lists each trigger as MET / UNMET and the INC-CLOCK value. | Trigger ID | Predicate | Effect | |---|---|---| | **D-GOV-TREATY** | A legally binding international instrument with quantified compute or training-run caps is in force for at least two states that, together, hold an estimated majority of publicly reported frontier training compute. | R-DN-2 on **C3** (0.94 if at 1.00). | | **D-INC-CLEAR** | INC-CLOCK ≥ 365 days (section 12.6). | R-DN-1 on **C6**. | | **D-INT-REPL** | A safety or interpretability result that meets C4’s independent-replication clause. | R-DN-1 on **C4**. | | **D-DEP-PAUSE** | Documented pause, rollback, or restored human-in-loop covering a named critical-infrastructure class across a primary-source set of operators (regulator order or equivalent). | R-DN-1 on **C2** (and C7 annotation). | Unmet triggers are listed, not implied. At genesis, unless a predicate is already true on the as-of date, all four read UNMET and INC-CLOCK starts at the as-of date (or at the last S2+ incident if one is scored that day). --- ## 14. Display requirements A snapshot that omits these is not a v0.1 snapshot. ### 14.1 Beside G1 1. “Index of judgments, not a measurement.” 2. Point (untrimmed weighted median, conditioned). 3. Source band, labeled, same-gauge only. 4. Conditioning envelope [unconditioned × 0.65, unconditioned × 1.54]. 5. Unconditioned aggregate. 6. Current M and the **forecaster-disagreement comparison** (§10.3). 7. Down-move eligibility (§13). 8. Absolute percentage-point delta since last snapshot, not a relative “+12%.” 9. Static snapshot as-of date and **next refresh date**. Automation = v2. ### 14.2 G2 and G3 - G2: thin-panel or aggregate per minimum-n; near-term extinction panel as-stated; never a silent G1 copy. - G3: constructed-estimate treatment (§4.3). If the mockup-style percentage is still on a design preview, it stays labeled ILLUSTRATIVE. ### 14.3 Resolution floor **[SJ-RES]** Material change: **≥ 0.5 percentage points** in the headline. Smaller absolute moves are recorded in the ledger and labeled **no material change** on the sparkline. A 5.0% → 5.6% move is **+0.6 pt**, not “up 12%.” ### 14.4 Genesis v1’s founding number has no predecessor event. **[SJ-GENESIS]** The first snapshot emits row **GENESIS-001**: unconditioned aggregate, M (default 1.00), rationale (“founding static snapshot under methodology v0.1”), and day-one down-move eligibility. Subsequent movements cite events. This document is not GENESIS-001; compute is a later work item. --- ## 15. Cadence and v1 shape | Tier | v1 (locked static) | Later | |---|---|---| | A | Frozen as of the snapshot’s source cutoff | Quarterly refresh + RESET rule if the window moves | | B | Out of mixture (unverified) | After ID verification; still two-tier flagged | | C | Event ledger may be empty at genesis (M=1.00) | Event-driven inside this rubric | | Site | Static files: headline, band, method, both ledgers, JSON of the compute | Scheduled job is v2 | Tech shape is unchanged in spirit: a versioned compute script writes JSON; pages render it. This PR does not build that. --- ## 16. Operator prior and site incentive COI labeling that covers only sources is incomplete. | Row | Content | Control | |---|---|---| | **SITE-PRIOR-001** | Operator thesis, 2026-09-15: the most plausible route is mundane deployment of not-ready AI into sensitive systems faster than verification, liability, or regulation. This thesis motivated G3 and the existence of C2 and C7. | It does not set the G1 headline. C2+C7 log-share cap. G3 is not an aggregate. | | **SITE-INCENTIVE-001** | The product is more noticeable if the number moves. | Static v1; default R-NULL; quantized steps; resolution floor; public ledger of non-moves. | Both rows print on the method page. --- ## 17. Corrections and retractions | Policy | Rule | |---|---| | Correction | Timestamped corrections log: what was wrong, what it is now, effect on M and on the headline in percentage points. | | Retraction | A retracted event reverses its delta (subject to caps), citing the original event ID. The original row stays, marked RETRACTED. | | Source revision | If a source revises a scored number, the ledger records the movement; the mixture uses the latest primary numeric utterance. Hinton ~10% → 10–20% is the model. | --- ## 18. Non-comparability **[SJ-NOCOMP]** This index is not Metaculus question N, not a market price, not a personal p(doom) poll, and not Ord’s or Carlsmith’s number. Definitions, horizons, weights, and M1 conditioning differ. A comparability note lives on the method page. Matching a market is not a success criterion. --- ## Appendix A — Critique disposition (R1 → v0.1) All 41 findings adopted in principle (verdicts). Mapping to this text: | ID | Sev | v0.1 location | |---|---|---| | A1 | HIGH | §10.4, §11.2–11.3 overall primary; per-category derived | | A2 | HIGH | §10 entire; comparison line §10.3, §14.1 | | A3 | HIGH | Log-symmetric [0.65, 1.54] §10.2 | | A4 | HIGH | §11.1 quantized set | | A5 | HIGH | §11.5 reference window; unconditioned line §9 | | A6 | HIGH | §12.1 checklist; one-fact-one-row E3 | | A7 | HIGH | §12.2 one event, one category | | A8 | HIGH | §12.3 declared signs | | A9 | HIGH | §12.6 clocks; §13 D-INC-CLEAR | | A10 | MED | §11.3 numeric caps, effective dates, log space | | A11 | MED | §11.2 log-sum then single cap | | A12 | MED | §11.6 half-life 180d | | A13 | MED | §11.7 RESET row | | A14 | MED | §9.3 conditioning envelope | | A15 | MED | §11.8 vendor bar | | A16 | HIGH | §12 hash-stamped rubric, rule IDs | | A17 | LOW | §11.3 precedence; suppressed amount | | B18 | HIGH | §9.1 untrimmed weighted median declared | | B19 | HIGH | Untrimmed, because outliers are locked in | | B20 | MED | §9.2 leave-one-out | | B21 | HIGH | §9.3 same-gauge band | | B22 | MED | §9.3 composition labels | | B23 | MED | §8.2 family caps; two-tier flag | | C24 | HIGH | §4.4, §11.4 | | C25 | HIGH | §4.3 G3 constructed | | C26 | HIGH | §6.1 m=4 sensitivity, no naive up-map | | C27 | HIGH | §6.2 as-stated vs mixture-eligible | | C28 | MED | §6.3 minimum-n | | D29 | HIGH | Ledger mapping table; §6.1 | | D30 | HIGH | §5 primary numeric only | | E31 | HIGH | §8.1 equal 3.0 tie; weighting-scheme range | | E32 | HIGH | Task-comparability notes §8.1 and ledger | | E33 | MED | §8.3 named-statement line | | E34 | MED | §7.2 Tier B out of v1; no undeclared recency | | F35 | HIGH | §13 mechanical down-triggers | | F36 | HIGH | §14.4 GENESIS-001 | | F37 | MED | §16 site prior and incentive | | F38 | MED | §14.3 0.5 pt floor; absolute deltas | | F39 | MED | §18 | | F40 | LOW | §17 | | F41 | LOW | §15, §14.1 static snapshot language | **Three insist-amendments:** (1) §10, (2) §12, (3) §13. --- ## Appendix B — v0.1 mixture membership (no headline computed) G1-2100 scoring rows (n=5, full aggregate eligible): | ID | Value used | Weight | |---|---|---| | A-XPT-SF | 2.13% | 1.5 | | A-XPT-DE | 12% | 1.5 | | A-CARL-2022 | 5% (companion >10% not the point) | 2.0 | | A-ORD-2020 | 10% | 2.0 | | A-AII-2024 | 10.0% | 3.0 | G2-2100 scoring rows (n=2, thin panel): A-XPT-SF 0.38%, A-XPT-DE 3%. Near-term panel (not mixed): A-HINT-2024 10–20% / ~30y; A-HUB-2026 >10% / 10y. Ledger only: A-BENG-2023; A-AII-2023 (sensitivity substitute); A-LECUN-2023; A-YUD. This appendix is a membership list. Running W(·) is the static v1 compute in `compute/snapshot.py`; the printout is `docs/SNAPSHOT-v1.md`. --- ## Appendix C — Site-judgment index | Tag | One-line | |---|---| | SJ-G3-CONSTRUCT | G3 is constructed, not an aggregate | | SJ-OVERLAP | Thesis overweight named and capped | | SJ-UTTER | Primary numeric utterances only | | SJ-CARL-POINT | Carlsmith mixture point is 5%, >10% companion | | SJ-MAP | Mapping table; no G2 face-value dump into G1 | | SJ-M-FACTOR | m=4 from XPT experts, sensitivity only | | SJ-HORIZON | No hazard-rate conversion; H-century / H-near / H-unspecified | | SJ-MIN-N | 3 / 2 / <2 / 0 display rule | | SJ-TIERB-V1 | Tier B out of static v1 | | SJ-TIE | Survey class and XPT family both 3.0 | | SJ-FAM | Family combined-weight caps | | SJ-NAMED-LINE | With/without named statements | | SJ-STAT | Untrimmed weighted median is the number | | SJ-BAND | Same-gauge band; envelope separate | | SJ-BOUND | Overall [0.65, 1.54] from √D_ref | | SJ-BOUND-ROUND | Two-decimal rounding | | SJ-QUANT | {0.94, 0.97, 1.00, 1.03, 1.06} | | SJ-LOGSUM | Sum of logs, then clip | | SJ-QCAP | Quarterly factor 1.15 log-symmetric | | SJ-C27 | C2+C7 ≤ 2/7 of overall log-span | | SJ-REF | Neutral window 2020-01-01–2026-09-15 | | SJ-DECAY | 180-day half-life to 1.00 | | SJ-C1-EXT | External instruments only for C1 | | SJ-EVID | Default no movement; two step sizes | | SJ-RES | 0.5 percentage-point material floor | | SJ-GENESIS | GENESIS-001 required at first snapshot | | SJ-NOCOMP | Not comparable to markets / Metaculus / personal p(doom) |