Episode Summary
Executive Summary: Diogo Almeida argues TypeSafe’s Jev represents a shift from language-model-centered AI to machine-native intelligence: models optimized for calibrated decisions inside software, not human-readable text. The conversation contrasts LLM demos and agents with more reliable, software-embedded decision APIs, emphasizing calibration, task design, and engineering control over speed or cost.
Main Topics: Machine-native intelligence vs. string generation (Priority: 5/5): Almeida frames Jev as a model class optimized for decisions and software execution, not for producing text that humans read. He argues current AI over-optimizes for strings and under-delivers on automation. Why LLMs feel smart but fail at workflows (Priority: 5/5): The discussion centers on the gap between impressive chat, coding, and research performance versus weak reliability in tasks like data entry, customer support, refunds, and underwriting. Optimization, calibration, and RLCD (Priority: 5/5): Almeida claims the key to usefulness is optimizing for the right task—calibrated decisions—via RLCD, which he positions as a broader framework than post-training or fine-tuning. Classifiers as an interface to intelligence (Priority: 4/5): He argues classifiers are not a minor ML technique but the natural interface for embedding intelligence into software, and that decision models should expose probabilities and thresholds to developers. Role of code, workflows, and agent architecture (Priority: 4/5): The interview explores how Jev fits into software systems, with Almeida preferring user-owned code and workflows over a generic AI harness, and suggesting future agent architectures may become more software-heavy. Reliability, calibration, and business deployment (Priority: 5/5): A major theme is that enterprise users want dependable background automation, not flashy demos. He emphasizes that reliability, calibration, and controllable thresholds matter more than novelty or raw model speed. OpenAI, RLHF history, and lessons from scaling (Priority: 3/5): Almeida reflects on his OpenAI background, the impact of RLHF on LLM success, and how different optimization targets produce different model behaviors and trade-offs.
Key Arguments: AI is broadly optimized for producing strings, but software needs calibrated decisions; optimizing for the wrong task explains much of the gap between chat success and workflow failure. Classifiers are not trivial—they are the interface through which intelligence becomes useful in software. Reliability is the real product value, not speed or low cost; users ultimately pay for systems they can trust in production. RLCD is intended to expose native engineering controls—thresholds, probabilities, and decision policies—rather than force users to prompt their way around model behavior. Current agents are constrained by brittle model behavior, KV-cache limits, and tool-call overuse; better decision models could enable more reliable agent architectures. Post-training on LLMs was easier because the target shape was still string-like; decision modeling requires a different pipeline and more fundamental changes. Generalization is not absent in AI, but it depends strongly on the optimization objective; RLHF generalized well because human-pleasing text is a broad subjective target. Engineering systems should combine ML with code and explicit logic, rather than expecting one model to replace the rest of the software stack.
Data Points: Time TypeSafe was in stealth: 2 years - Almeida says the company stayed in stealth for a long time before Jev’s public release. Time at OpenAI: about 4.5 years - He describes his tenure at OpenAI working on major projects including InstructGPT and RLHF. October 2016 to the prior conversation: nearly 10 years - The host notes they last spoke in October 2016, highlighting how long Almeida has been in the field. Typical human team equivalent: a dozen people - Almeida says Jev could give a random developer the equivalent of an MLE team of about twelve people. Customer support over-refusal example: sometimes refuses, sometimes not - He cites non-deterministic refusal behavior as evidence that blending language generation and decisions is bad engineering. Model development expectation: thought it would take 1 week; took 4 years - Almeida says he initially believed the project would take a week, but it became a multi-year effort. Product release timing reference: 3 weeks ago - The transcript opens by saying TypeSafe came out of stealth with Jev three weeks prior.
Pivotal Quotes: "Where the f is all the automation." — Diogo Almeida: His elevator pitch for Jev, expressing frustration that AI is intelligent but still not broadly useful for real work. "You get what you optimize for." — Diogo Almeida: The central thesis of the interview: model behavior follows the objective, so optimizing for strings yields different results than optimizing for software decisions. "I want AI to be so predictable that it's boring, boring like SQL." — Diogo Almeida: He explains his ideal end state: reliable, standardized intelligence that developers can trust like a database query.
Implications: If Almeida’s thesis is right, the next wave of AI won’t be bigger chatbots—it will be reliable decision systems embedded in software. That shifts value toward calibration, workflow design, and developer control, and away from demo-first agent hype.