HomeForecasts

Registered forecasts about the X algorithm

Everyone in this niche makes claims; nobody scores themselves. This page is our registry: falsifiable predictions about the algorithm, the repo, and this market — each with a stated resolution source a stranger could judge, an outside-view anchor, an auditable evidence ledger, and a deadline. When each resolves, the outcome and Brier score get published here, wins and losses alike.

PredictionOursShadowDeadlineStatus
A new commit is pushed to github.com/xai-org/x-algorithm (main branch HEAD chang…48%60%2026-09-15 23:59open
Numeric production ranking weights (the values the withheld params module suppli…14%15%2026-12-31 23:59open
X publishes its full codebase publicly WITH a named third party attesting produc…7%10%2026-12-31 23:59open
At least one of TweetHunter, Typefully, SuperX, Teract, Pounce, OpenTweet public…45%30%2026-11-01 23:59open
X announces a further increase to any per-item API price (per post read/write/lo…79%75%2026-12-31 23:59open
Rerunning our published scripts (phx_fetch.py + phx_analyze.py, frozen at build …78%80%2026-08-31open
CONDITIONAL: IF xai-org publishes a new runnable checkpoint by 2026-12-31 (artif…80%75%resolves when/ifopen

A new commit is pushed to github.com/xai-org/x-algorithm (main branch HEAD changes from 0bfc279) between 2026-08-01 and 2026-09-15

Framework: 48% · shadow baseline (gut, recorded first): 60% · CHANGE · driver: xai-release-cadence · window: 2026-08-01 → 2026-09-15 23:59 UTC

Anchor (hazard, p₀=0.38): k=2 public drops (Jan 20, May 15) over L=192 days -> lambda=0.0104/day; w=46 days; P=1-e^(-0.479)=0.38

Evidence ledger (each item bounded ±1.0 log-odds):

Resolves at: github.com/xai-org/x-algorithm commit history (main branch) — HEAD SHA on main differs from 0bfc279 on deadline day. Stranger-testable in one click.

Falsified by: HEAD still 0bfc279 on 2026-09-16

Numeric production ranking weights (the values the withheld params module supplies to ranking_scorer.rs, or their successor) appear in public code from X/xAI by 2026-12-31

Framework: 14% · shadow baseline (gut, recorded first): 15% · CHANGE · driver: weights-transparency · window: 2026-07-31 → 2026-12-31 23:59 UTC

Anchor (reference-class, p₀=0.12): (a) Musk public-deadline promises delivered on time and in full: low base rate (~15%; 2022 'open source the algorithm' promise delivered partial, ~11 months late); (b) revealed preference: two 2026 drops each deliberately withheld params

Evidence ledger (each item bounded ±1.0 log-odds):

Resolves at: github.com/xai-org (any repo) / official X engineering channels — A public file from X/xAI contains numeric weight values consumed by the production ranking formula (named constants for the engagement heads). Partial (demo/sample constants labeled illustrative) does NOT count.

Falsified by: No such file public by deadline

X publishes its full codebase publicly WITH a named third party attesting production-matches-GitHub, by 2026-12-31 (the July 15 pledge, delivered as stated)

Framework: 7% · shadow baseline (gut, recorded first): 10% · CHANGE · driver: weights-transparency · window: 2026-07-31 → 2026-12-31 23:59 UTC

Anchor (reference-class, p₀=0.08): Musk-deadline reference class as p2, further discounted for the two-part criterion (full publication AND third-party verification)

Evidence ledger (each item bounded ±1.0 log-odds):

Resolves at: official X/xAI announcement + the published repo(s) — Both parts required: (1) repo(s) presented as the full X codebase are public; (2) an announcement names a specific third-party verifier of production parity. Stranger reads the announcement and repo landing page only.

Falsified by: Either part missing at deadline

At least one of TweetHunter, Typefully, SuperX, Teract, Pounce, OpenTweet publicly markets a feature as scoring with or running the open-source Phoenix model (naming Phoenix or linking xai-org/x-algorithm in a product/feature description) by 2026-11-01

Framework: 45% · shadow baseline (gut, recorded first): 30% · CHANGE · driver: competitor-adoption · window: 2026-07-31 → 2026-11-01 23:59 UTC

Anchor (hazard-laplace, p₀=0.69): k=0 adoption events in 77 days since the runnable release; Laplace (k+1)/(L+2)=0.0127/day; w=92d -> 0.69. Flagged: k=0 makes this anchor weak by construction

Evidence ledger (each item bounded ±1.0 log-odds):

Resolves at: the six vendors' public marketing pages / product changelogs / official X accounts (archived on resolution day) — An archived public page or post from one of the six that describes a product feature and either uses the name 'Phoenix' for X's model or links the xai-org/x-algorithm repo. Generic 'AI-powered' or 'algorithm-optimized' claims do NOT count.

Falsified by: No qualifying page by deadline · Largest shadow-vs-framework divergence in the session (30% vs 45%) — the Laplace anchor on k=0 drags upward. This is the A/B doing its job; resolution will say which arm was right.

X announces a further increase to any per-item API price (per post read/write/lookup or per-content surcharge), effective in 2026, posted on devcommunity.x.com between 2026-08-01 and 2026-12-31

Framework: 79% · shadow baseline (gut, recorded first): 75% · SQ · driver: x-api-economics · window: 2026-08-01 → 2026-12-31 23:59 UTC

Anchor (hazard, p₀=0.77): k=2 pricing increases in 2026 (Feb pay-per-use default, Apr 20 link-post surcharge) over L=211 days -> lambda=0.0095/day; w=153d; P=1-e^(-1.45)=0.77

Evidence ledger (each item bounded ±1.0 log-odds):

Resolves at: devcommunity.x.com official announcements category — An official announcement post in the window states an increased per-item price (any item) taking effect in 2026. Price restructures that only decrease or hold prices do not count.

Falsified by: No qualifying announcement by deadline

Rerunning our published scripts (phx_fetch.py + phx_analyze.py, frozen at build commit) on >=300 previously-unfetched sports-corpus posts on or after 2026-08-14 yields median replies for question-mark posts >= 2x the median for non-question posts

Framework: 78% · shadow baseline (gut, recorded first): 80% · SQ · driver: engagement-content-correlates · window: run window 2026-08-14 → 2026-08-31

Anchor (reference-class, p₀=0.75): one prior measurement (4.0 vs 1.0 median on 1,311 posts) with small-integer medians; regression toward the mean expected; 2x threshold leaves headroom

Evidence ledger (each item bounded ±1.0 log-odds):

Resolves at: the rerun's lab.json output, published at /api/ on this site (scripts are public; anyone can reproduce) — feature_effects.question_mark: median_replies_with >= 2 x median_replies_without, on n>=300 posts none of which appear in the Jul 31 fetch files

Falsified by: ratio < 2x on the qualifying rerun

CONDITIONAL: IF xai-org publishes a new runnable checkpoint by 2026-12-31 (artifacts our frozen scripts run against with at most path fixes), THEN the within-author backtest on it yields median Spearman <= 0.10

Framework: 80% · shadow baseline (gut, recorded first): 75% · SQ · driver: mini-model-behavior-stability · window: resolves when/if the antecedent occurs, else VOID at 2026-12-31

Anchor (reference-class, p₀=0.78): the null replicated across two viewer contexts (warm rho=-0.003, cold rho=+0.008); ID-hash architecture makes post-quality signal structurally hard for a mini distillation

Evidence ledger (each item bounded ±1.0 log-odds):

Resolves at: our frozen phx_score.py + phx_analyze.py output on the new artifacts, published at /api/ — within_author_backtest.fav_spearman_median <= 0.10 with >=50 qualifying authors. VOID (not scored) if no qualifying checkpoint ships.

Falsified by: median rho > 0.10 on a qualifying run

Method

The discipline is borrowed from calibration research (hazard-rate anchors for recurring events, log-odds evidence ledgers, positive-event framing, resolution criteria designed before the forecast, shadow-baseline controls). Machine-readable registry: predictions.json. Think a number is wrong? The ledger shows exactly which evidence item to attack — tell us publicly.