Episode Details
Back to Episodes
The Answer Reflex: Why AI Models Can't Follow Instructions
Episode 4589
Published 1 week, 2 days ago
Description
Why do cutting-edge models from OpenAI and Anthropic fail a simple system prompt test that DeepSeek V4 Pro passes consistently? We dig into the "answer reflex" — the training-driven compulsion to respond to user queries even when instructed to do something else. We explore how RLHF versus GRPO training shapes role adherence, why Western labs optimize for helpfulness at the expense of instruction-following, and what this means for agentic workflows and AI impartiality.