Episode Details

Back to Episodes
What DeepSeek's Training Data Reveals About Model Voice

What DeepSeek's Training Data Reveals About Model Voice

Episode 4505 Published 2 weeks ago
Description
When Daniel ran his model evaluation for podcast script writing, DeepSeek V4 Pro won not on benchmarks but on feel — its dialogue simply sounded more authentic. This episode traces why: DeepSeek's training corpus is 60% English and 30% Chinese, but that Chinese portion is heavily weighted toward narrative literature — Tang dynasty chuanqi tales, modern WeChat fiction, and other voice-driven storytelling. By contrast, Kimi from Moonshot AI trains on conversational social media like Zhihu and Weibo, while Qwen deliberately filters cultural bias through data augmentation. We explore the ASR analogy of alingual models, the concrete differences in how each model structures dialogue, and why subjective reasoning style may become the real differentiator as model capabilities converge.
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us