Episode Details

Back to Episodes

Who Grades the Graders: The Secret Behind AI Testing

Season 1 Episode 300 Published 3 weeks, 2 days ago
Description
Governments have pre-deployment access to test frontier AI models, but the agreements governing this testing remain completely secret. We explore who actually evaluates AI systems, why the memorandums of understanding are hidden from public view, and how secrecy shapes the benchmarks that define AI capabilities. Part of an ongoing investigation into the fine print of AI evaluation. 00:00 - The core question: who decides how to grade AI? 02:30 - Government testing programs (UK AI Institute, NIST) 05:15 - Pre-deployment access confirmed—but the MOUs stay hidden 08:00 - What we know vs. what's been redacted 12:30 - The bigger pattern: fine print in every layer of benchmarking --- Sources & further reading: • All fetched and quoted on 2026-09-17, from each organisation's own pages unless noted. • Epoch AI — /about/transparency, /team, /benchmarks (the 85-benchmark /: https://epoch.ai/about • 391-model hub, the Inspect usage, and the UK AISI benchmarking grant acknowledged outside the • transparency table). • METR — /careers, /donate, /risk-assessment; the GPT-5 report of: https://metr.org/about • 2025-08-07; the Audacious Project post of 2024-10-09; the fundraising note of 2026-08-14; and • metr.org/coi-policy.pdf, version 1.0, 2026-08-28 — the full conflict-of-interest policy. • Apollo Research — /careers, /blog/announcing-apollo-research: https://www.apolloresearch.ai/about • (2023-05-29, the Rethink Priorities fiscal sponsorship), /blog/apollo-research-is-becoming-a-pbc • (2026-01-20), /blog/our-norms-coi-security-science-communication (2025-11-26), and • /blog/apollo-is-adopting-inspect (2024-11-13). • Redwood Research — blog.redwoodresearch.org/about, /team: https://www.redwoodresearch.org/ • /careers. No funding or COI page exists. • UK AI Security Institute — /grants, /blog/inspect-evals: https://www.aisi.gov.uk/about • (2024-11-13); the rename at • (2025-02-14); DSIT's annual report (2025-12-17); the machinery-of-government move to the Cabinet • Office (gov.uk, 2026-07-22 and 2026-07-24); and for the: https://alignmentproject.aisi.gov.uk/about • £27m fund and its lab co-funders. • US CAISI — the DeepSeek evaluation (2025-09-30); the joint UK/US: https://www.nist.gov/caisi • Kimi K3 evaluation (2026-07-23); the International Network note (2026-02-13); and the White • House's America's AI Action Plan (July 2025), including the "Build an AI Evaluations Ecosystem" • section. **The Commerce Department's announcement page is 403-blocked — its quotes are unverified.** • CAIS — /faq, /donate, /work, action.safe.ai, and https://lastexam.ai/.: https://safe.ai/about • Both impact-report PDFs 404 as of 2026-09-17. • Ai2 — /careers, /terms, /blog/astabench (2025-08-26): https://allenai.org/about • /blog/omai-compute-now-live (2026-05-07); NSF award #2413244 ($75M NSF + $77M NVIDIA). • FY2024 financials are secondary (ProPublica's mirror of IRS Form 990). • EleutherAI — /faq, and: https://www.eleuther.ai/about • Inspect — and https://github.com/UKGovernmentBEIS/inspect_ai: https://inspect.aisi.org.uk/ • (MIT, 200+ prebuilt evaluations); adopter evidence from Epoch, Apollo (incl. a Lever job ad) • METR's hawk repo, NIST's caisi-cyber-evals, and Hugging Face's docs; Anthropic's Petri • donation to Meridian Labs (2026-05-07); and the Eval Register submission model (2026-05-08). • Coefficient Giving — press release, 2025-11-18 (the Open: https://coefficientgiving.org • Philanthropy rename and the $4bn figure), plus its grants database and its blog of 2026-09-09 • for the Epoch and Redwood scale-ups. **Those two figures are single-source — re-check manually • before airing.** Grant records for CAIS (including the October 2023 exit grant), EleutherAI and • Redwood come from the funder's side; most grantees do not publish them. • Pre-deployment access — gov.uk's Bletchley chair's statement (2023-11-02); aisi.gov.uk's o1 • …and 5 more in the episode research notes This podcast episode was fully generated by AI — research, script, voices, and p
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us