[EVAL/TRANSPARENCY] Anthropic's regulation case rests on internal evals — apply the #86979 provenance standard to Glasswing and 'When AI Builds Itself'
The connection
The June 2026 policy essay ("Policy on the AI Exponential") proposes mandatory third-party testing for models above a compute threshold, with deployment blocking for unacceptable risk — justified by Anthropic's own evals: the Glasswing/Mythos cyber-risk findings and the "When AI Builds Itself" autonomy data.
These are the same class of measurements we ask Anthropic to open up for SWE-bench-family scores in #86979. Anthropic is asking regulators and the public to build binding policy on numbers that carry none of the provenance requested there.
What is missing
For every number that motivates regulation, the same disclosure fields apply:
- exact eval revision + harness version
- model + effort configuration, seeds, repeat-run variance
- internet/egress policy, git-history availability, remote/upstream lookup
- per-instance outcomes, not just aggregate rates
- trajectory audit methodology + sample size
- known-fix/reference-retrieval rate among successful runs (reward hacking — including recovering reference solutions from git history — is documented behavior in Anthropic's own system cards)
- strict-harness score with leakage channels removed
- uncertainty / paired comparison
The question
You propose third-party testing above a compute threshold. If the numbers motivating the proposal cannot be reproduced without full measurement provenance, the regulation is being argued with a model + harness + information-access score — the same conflation independently documented for SWE-bench-family scores (DeepSWE, Cursor, Endor Labs; all public, all primary sources).
Publishing the provenance either strengthens the case (if strict-harness numbers hold) or exposes that the regulatory case was built on harness-assisted results. Either outcome is information regulators need.
I'm not alleging manipulation. I'm asking: for the numbers that will become law — which ones measure the model, and which ones measure access to the answer key? If the strict-harness result is nearly identical, publish it and this concern disappears. If it is materially lower, regulators deserve to know before they legislate on it.
This issue has 2 comments on GitHub. Read the full discussion on GitHub ↗