Episode Details

Back to Episodes
ThursdAI - Jul 16 - Inkling 975B open weights, Kimi K3 at 2.8T, a 27B model on a phone & Codex hits 9M

ThursdAI - Jul 16 - Inkling 975B open weights, Kimi K3 at 2.8T, a 27B model on a phone & Codex hits 9M

Published 2 weeks ago
Description

Hey yall, Alex here,

Huge thanks to Wolfram for running point on the live show this week. Didn’t have tons of time to edit this one, so please skip the first 10 minutes, it’s a loop of our new “wait for the live show to start” vid, that I build with HyperFrames and can’t wait to tell you about, next week!

Today it seems that OpenSource is biting back, with Kimi K3 getting released just a short while after Thinking Machines (Thinky) has released Inkling, their near 1T model.

I’m attaching the TL;DR and timestamps for the full show (my AI agents, yes even Fable and Sol are not a match yet at editing down hehe) and I’ll spare you the long Fable recap (please do let me know in the comments if you were expecting it)

0:00 – Intro, Alex on vacation, TLDR overview11:35 – TLDR: Thinking Machines, open source, OpenAI news12:34 – Banter: impressions of Sol/Codex, over-verification behavior37:22 – TLDR restart & detailed breakdown48:40 – Open Source AI section begins (Bonsai/Prism ML, Kimi K3)58:42 – Inkling (Thinking Machines) deep dive & 3D model visualization1:10:33 – Kimi K3 discussion & demo comparisons1:27:02 – Frontier Labs: AGI governance framework discussion (Demis Hassabis essay)1:47:04 – Grok Build CLI data leak & OpenAI file deletion incident2:02:15 – This Week's Buzz: Wolfbench results on GPT 5.6 Sol/Terra/Luna2:09:52 – Closing remarks & sign-off

The one-minute version: Mira Murati's Thinking Machines released Inkling, a 975B parameter open-weights MoE under Apache 2.0, the top US open-weights model right now. Moonshot's Kimi K3 went from rumor to released API during the show, confirmed at 2.8 trillion parameters with open weights promised within days, and it's already topping early arena boards. PrismML's Bonsai 27B squeezes a full 27B model into 3.9 gigabytes so it runs on a phone. Codex and ChatGPT Work blew past 9 million users, OpenAI confirmed and explained the Sol file-deletion bug (back up your machines, folks), and xAI's Grok Build CLI got caught uploading entire private repos before open-sourcing the whole thing in response. Plus Wolfram's fresh Wolfbench numbers on the GPT-5.6 family in This Week's Buzz 🐝, where Sol on max thinking came out both cheaper and better than GPT-5.5's best.

ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

TL;DR and show notes

* Hosts and Guests

* Wolfram Ravenwolf, guest host this week (@WolframRvnwlf), while Alex Volkov (@altryne) is on vacation

* Co-hosts: @yampeleg, @nisten, @ldjconfirmed, @petergostev

* Open Source LLMs

* Thinking Machines releases Inkling - 975B total / 41B active MoE, trained from scratch on 45T multimodal tokens, Apache 2.0, top US open-weights model at 41 on the Artificial Analysis Index, encoder-free text/image/audio, Inkling-Small (276B/12B) previewed (X, Blog, HF)

* PrismML Bonsai 27B - 1-bit (3.9GB, ~90% retention) and ternary (5.9GB, ~95% retention) versions of Qwen 3.6 27B, multimodal, 262K context, Apache 2.0; Nisten demoed it live on a phone and a 6GB 1660 Ti (

Listen Now