Episode Details

Back to Episodes
Episode #574: Evals, Ontologies and the Unmappable World of Business

Episode #574: Evals, Ontologies and the Unmappable World of Business

Episode 574 Published 1 week, 1 day ago
Description

In this episode of the Crazy Wisdom Podcast, host Stewart Alsop sits down with Ryan Marsh of thestack.io to explore what it really takes to build production AI systems. They discuss how production AI has evolved from simple prompt-to-API demos into complex systems requiring evaluation suites, human-in-the-loop feedback mechanisms, and sophisticated approaches to handling context and data retrieval. The conversation covers the challenges of domain mapping, the fundamental difficulty of translating messy human business processes into structured systems, and the role of ontologies in AI development. Ryan and Stewart also examine the limitations of LLMs, the debate between specialization versus generalization in AI models, consciousness and cognition in system design, and the regulatory landscape facing AI companies. They touch on infrastructure constraints, the democratization of AI through open source models, and whether we'll eventually hit a ceiling where human intelligence can no longer distinguish between increasingly capable AI models.

Key Insights

1. Production AI systems today fundamentally differ from demos through their reliance on comprehensive evaluation suites that function like unit tests to measure and maintain quality, though they cannot be as deterministic. The key distinction is that production systems require clearly defined metrics for what good looks like, along with feedback mechanisms that allow the system to evolve over time. Without this foundational understanding of success metrics and continuous improvement processes, a system is not truly production-ready regardless of how many users it serves.

2. The fundamental challenge in building production AI systems is not the technology itself but rather mapping business domains into structured formats that models can work with effectively. This problem of translating messy, subjective human processes and language into precise specifications has plagued software engineering for decades. Different people within organizations use the same words to mean completely different things, and humans naturally operate with assumed context and imprecision that must be explicitly defined for AI systems to function reliably.

3. Large language models excel at generalization but struggle with specialization, which creates friction in production environments where specific outputs or styles are required. While they can code in any programming language, getting them to write code exactly the way a particular engineer wants remains extremely difficult. This explains why professional documentation and specialized coding tasks often require extensive prompting and fighting with the models, as they naturally gravitate toward their trained patterns rather than highly specific user preferences.

4. Modern production AI systems primarily solve classification problems wrapped in natural language interfaces rather than requiring true open-ended cognition. The models work best when they can leverage reasoning over provided information to make verifiable decisions, but they still lack common sense despite their vast knowledge. Success comes from teaching models everything about your specific domain and what good and bad outcomes look like, rather than relying solely on their general intelligence.

5. Context retrieval in production AI systems is fundamentally a data storage, search, and retrieval problem that has been solved many different ways throughout computing history. The appropriate solution depends entirely on the type of data being accessed, whether through vector databases, graph databases, relational databases, or even simple text search. The harnesses and frameworks for orchestrating AI agents have matured significantly, making the real challenge the quality and structure of the data being fed to these systems.

6. Human-in-the-loop feedback mechanisms are essential for production AI because models will inevit

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us