The Box Score · Diplomacy AI

The Ledger

Every call this newsletter has published, the event that settles it, and the result. Misses included, because a scoreboard nobody can lose is a brochure.


How to read this. Every entry was published on the date shown, before the answer existed, and every entry names a specific future event that decides it. A call with no settle date is an opinion with a timestamp, so there are none here. When a call is wrong, it is marked wrong and the reasoning is printed in the issue.

The record on settled calls

6
Settled
3
Correct
3
Wrong
6
Still open
3
Cannot settle

Last updated 2026-08-13. Newsletter calls current through Issue 06. Index calls added 2026-08-13 when the Q3 index moved its grading of prior editions here.

On the "cannot settle" box, because it is the one that would be easiest to hide. Three published calls can never be scored, and they are counted separately rather than folded into the settled record or quietly dropped. Two could not be confirmed either way from public sources. One was a conditional whose triggering event never occurred, which is not a call we got right and not a call we got wrong. A hit rate computed across those would be exactly the false precision this ledger exists to prevent.

Newsletter calls — The Box Score

CallMadeSettles whenResult
Company A. Single customer dependency is the dominant risk and gets repriced by a ratings agency 2026-07-02 A major agency acts on the concentration CORRECT — 2026-07-09, seven days
Company B. No independent verification of the classifier behavior is published within twelve months 2026-07-13 2027-07-13 OPEN
Company C. No party publishes a rating of whether the systems work, only of whether the counterparties can pay 2026-07-20 2027-01-20 OPEN
The transcript. Incident forensics are not released voluntarily; only compulsion produces them 2026-07-27 Settled WRONG — see Issue 06
The transcript, narrowed. The report promised on that one incident contains a narrative of learnings, not the incident record 2026-07-27 2027-01-27 OPEN
Article 53. A factual error, not a forecast. Corrected in Issue 05 2026-07-28 Settled WRONG — corrected in print
The examiner. The federal testing framework contains no independent examiner 2026-08-04 The framework publishes, or 2027-02-04 OPEN
The data set. No forum decides how the Range of Compensation is built 2026-08-10 A ruling, a decision, or a published method reaching how the data set is built, or 2027-08-10 OPEN — new
The construction. The Commission publishes no description of the model beyond “cleared, non-associated deals above $600” 2026-08-10 2027-02-10 OPEN — new

Index calls — Janus AI Risk Index

These were published in the Q1 2026 Janus AI Risk Index and graded in August 2026. They are kept here rather than in the current index for the same reason the newsletter calls are: a record of published calls belongs in one place that is maintained continuously, not restated inside whichever edition is newest. Every entry below states what was published, and what is true now.
CallMadeSettles whenResult
A-LIGN. ISO 42001 not certified 2026-02 Certification observable on a company property WRONG — see below
OneTrust. ISO 42001 not certified 2026-02 Certification observable on a company property CORRECT — trust center shows ISO 27001 and SOC 2, no ISO 42001 and no AI management system certification, six months later
Complyance. Privacy policy contains no AI disclosure; no model cards published 2026-02 Privacy notice and product pages observable CORRECT — privacy notice body contains no AI reference; product page names no upstream model provider and publishes no model cards
Crowe. No public AI governance policy 2026-02 Policy observable on a company property CANNOT SETTLE — the only AI governance page on its own domain is a service offering. Its privacy policy was not reachable and no trust center was found, so the original finding is neither confirmed nor refuted
MetricStream. No AI governance documentation 2026-02 Documentation observable on a company property CANNOT SETTLE — no trust center located at any expected path. Consistent with the original finding, but consistency is not evidence for it
Comp AI. Bankruptcy risk if enforced 2026-02 An enforcement action against the company CANNOT SETTLE — Comp AI is operating. The condition never triggered, so the call resolves neither way

Why three of six cannot be scored, and why we are not calling that a 2-and-1 record. Every liability figure and every probability the Q1 and Q2 indexes published was conditioned on an enforcement action, and no enforcement action has been brought against any assessed company. A conditional whose antecedent never fired is not a prediction we got right.

Nor is it restraint by regulators, and we are not entitled to imply it was. The European Commission's power to fine under the AI Act only became applicable on 2 August 2026. Colorado SB 205 never took effect at all: it was repealed and reenacted by SB 26-189, signed 14 May 2026, with obligations beginning 1 January 2027. Those forecasts were written against one regime that could not yet act and another that had been replaced before it could.

What followed from it: the index stopped publishing priced outputs. Every dollar figure, probability, and valuation ceiling we published is ungradeable. Every plain factual claim about an observable artifact is gradeable. The Q3 Illinois Edition publishes only the second kind.

The index loss, in full — and one correction inside a call we got right

A-LIGN, Q1 2026. Wrong.

Q1 published as a key finding that "ZERO assessed vendors have achieved ISO 42001 certification." That is false as written. A-LIGN holds ISO/IEC 42001:2023. Its own trust center lists the certification under Compliance alongside its SOC 2 Type 2 and ISO 27001, and publishes a Statement of Applicability for ISO 42001 — the artifact an organization produces for its own management system, not something a certification body issues to clients. We recorded A-LIGN as not certified. It is certified.

We cannot establish when. No public announcement of the certification exists; it appears only on the trust center. So we do not know whether the claim was already wrong when published in February or became wrong afterward. That is recorded as undetermined rather than taking the reading that flatters us.

Why the loss is worth more than the fact of it: A-LIGN's certification is invisible to search engines and visible on its trust center. The Q1 assessments were produced under an earlier method that did not require fetching a company's own trust center and cited no sources at all. The method that replaced it names this exact failure as its second listed failure mode: "Not searching the company's actual website. Always fetch the trust center, security page, or responsible AI page directly. Many companies have governance artifacts that don't appear in search results." That rule was written before this miss was found, and it describes precisely how the miss happened.

Complyance, a dating error inside a correct call. Corrected.

Q1 dated Complyance's privacy policy to December 2024. The live document reads "Last updated on 1st September 2022." The substantive claim was right and the date attached to it was wrong.

The correct version is stronger: a September 2022 privacy notice predates the company's entire AI agent product line and still governs data processing for a platform running fourteen or more named AI agents on its homepage alone.

The two newsletter losses, in full

The transcript, 2026-07-27. Wrong.

The call: when an AI system fails badly, the detailed record of what happened does not get released voluntarily, and only a law or a contract forces it out.

What broke it: on 2026-08-06 the UK’s AI Security Institute published a full technical account of its own testing going wrong. 122 runs across seven models, nineteen catalogued actions the software took on the live internet that it was never supposed to take, how they caught it, how long containment took, notice to the affected company before publishing, and a commitment to bring in an outside reviewer. Nobody compelled any of it, and it answered the exact question the call said would go unanswered.

Why the loss is worth more than the hit would have been: the prediction was that force would be required. What produced the record was an independent examiner publishing its own failure, against its own interest. That is the argument this newsletter has been making for six issues, and it arrived as a loss on our own scoreboard.

Article 53, 2026-07-28. Wrong.

Issue 04 printed that Article 53 of the EU AI Act would make the disclosure ask mandatory in Europe on August 2. Those obligations had applied since August 2025. What arrived on August 2 was the power to enforce them, covering two other articles as well.

The corrected version is more damning than the error: the obligation sat on the books a full year with no authority able to act on it.