The Threshold Report
← Front page

The Threshold Report - methodology - September 26, 2026

This page explains how The Threshold Report scores its lane records, the Security Dial, and The Threshold Board.

The whole system in one paragraph

The Threshold Report tracks real progress in AI capability and what that progress changes. Each lane starts at 100 and moves only when something becomes real: usable, deployed, approved, or released. Announcements alone score zero. If something matters but is not real yet, it stays in Upcoming, with the missing condition stated, until that condition is met. Every point comes from the fixed rules below.

The rules

What counts

An item scores when it becomes real at the standard required for its lane, not when it is merely demoed, promised, or announced.

A signed deal, lease, or contract for something that is not yet operating stays in Upcoming until it is operating. A regulatory approval or clearance counts as real.

If something has been announced or reported but is not yet real or available, it stays in Upcoming with a clear note on what still needs to happen.

The lanes

We track eight running totals:

Each lane starts at 100 on its start date and builds from developments reported in the Real Progress section.

The totals are not meant to be compared with one another. A higher number simply means more logged movement since that lane started, not that one field is more capable than another.

How much one event moves a lane:

Event Lane effect
Milestone - a capability crosses a line for the first time +10
Notable - a clear improvement to an existing capability +5
Minor - access expands, prices fall meaningfully, or a smaller step becomes real +2
Regression - something real is withdrawn or degraded Takes back the points it earned, or -3 if it predates the record
Restoration - a logged regression is reversed Returns the points that were removed
Not scored - real news with no capability change, such as incidents, certifications, partnerships, or policy context 0, with the reason stated

A single release can move more than one lane. If it meaningfully improves Text, Audio, and Video, for example, each lane is scored separately.

There is currently no daily or per-issue cap on how much a lane can move. A short-lived cap was used from August 10 to August 18, 2026; that history is noted under Edge cases below.

The Threshold Board

The Threshold Board tracks plain-language questions about what AI can actually do. Each question has one of three statuses:

A status changes only when the evidence is sustained and the capability is genuinely available.

When a new question is added, it gets its first status on the next daily update. That first status is the starting point, not a flip.

A standing status carries no date. When a question flips, the flip is shown with the date it happened.

The Security Dial

The Security Dial is our running score of AI risk and readiness. It runs from Under Control to Chaos and moves only on developments reported in the Security section.

Every item is marked Positive or Negative, with the reason stated. Exploits, misuse, and weaker oversight push the dial toward Chaos. Real safety gains, such as deployed defenses, foiled attacks, rules that have actually taken effect, or meaningful alignment and interpretability progress, pull it toward Under Control.

Words alone do not move the dial. An open letter, pledge, or call for a pause can be worth reporting, but it scores nothing until someone acts on it. A rule has to take effect, a defense has to be deployed, or a lab has to verifiably change what it does.

How far one item moves the dial:

Impact What it looks like Needle movement
Minor Contained or single-organization impact, such as one exploited flaw or one deployed defense 4 points
Notable Sector-level impact, such as misuse at scale, a rule in force, or meaningful alignment progress 8 points
Major Systemic impact, such as widespread exploitation in the wild, a landmark rule, or a fundamental interpretability result 15 points

The tally below the dial shows the points on each side and how many items put them there, cumulatively, since the record began.

How the needle settles

The needle follows the running difference between the two sides: Chaos points minus Under Control points. Nothing decays, and nothing is hidden.

Near the middle, every four points of difference move the needle by one degree, so even a quieter week of Security news can show up. The farther the needle moves from the center, the less each additional point shifts it.

A difference of 400 points reads at about 76. A difference of 800 reads at about 96. The needle can move very close to either end without ever being pinned there.

That soft ceiling keeps one unusually loud period from locking the dial at an extreme while still allowing the long-term record to lean strongly in either direction.

Independence

No scored item is ever paid for.

Any paid listing on The Threshold Board will be clearly marked and disclosed. The publication does not accept payment to defend a position or influence a score.

Edge cases

If something is announced first and becomes real later, it enters the record on the day it became real, not the day it was announced.

If a capability later regresses, it can lose points under the same standard that originally earned them.

There is no current cap on how much a lane can move in a day or issue. From August 10 to August 18, 2026, a temporary per-lane cap applied per issue and overflow was logged at zero. Those entries remain as originally recorded. Since August 19, every qualifying item has scored at its full tier.

When a story genuinely falls between two scoring tiers, it gets the lower score. If something could reasonably be a Milestone (+10) or a Notable (+5), it gets +5.

A frontier lab releasing a meaningfully better model is a good example. Is it a genuinely new capability, or a clear improvement to an existing one? If the answer is debatable, the lower score wins.

That restraint is deliberate. The index is designed with an anti-hype bias: when the evidence does not clearly justify the higher score, we under-score. The goal is for each lane's progress slope to remain trustworthy.

Thanks for reading. If you find the tracker useful, visit regularly, subscribe, and follow The Threshold Report on social media. And if you think we missed something or classified it poorly, let us know! We welcome constructive comments from the community.