No Priors
No Priors

Your AI Friends Have Awoken, With Noam Shazeer

Noam Shazeer played a key role in developing key foundations of modern AI - including co-inventing Transformers at Google, as well as pioneering AI chat pre-chatGPT. These are the foundations supporting today’s AI revolution. On this episode of No Priors, Noam discusses his work as an AI researcher,

Topics Discussed

Episode Summary

Executive Summary: The conversation traces Noam Shazeer’s AI journey from early Google NLP work to Transformers and Character AI. He argues that language models, enabled by hardware-friendly parallelism, are the core path to useful intelligence and potentially AGI as a byproduct of making products people want. The discussion covers model architecture, scaling, data, safety, product design, and why Character AI focuses on user-created personas, memory, and emotional utility.

Main Topics: Noam Shazeer’s AI trajectory at Google (Priority: 5/5): Shazeer recounts his long Google history, early interest in making computers do smart things, and involvement in foundational AI/NLP work before joining Google Brain in 2012. Why Transformers won (Priority: 5/5): He explains that deep learning succeeded because it maps well to modern hardware, and that Transformers improve on RNNs by allowing parallel sequence processing via attention. Language modeling as the core AI problem (Priority: 5/5): Shazeer frames next-word prediction as a simple, open-ended, highly scalable problem with massive internet-scale data and broad downstream capability. Scaling, data, and architecture limits (Priority: 4/5): He is bullish that the field has not hit a wall; he expects gains from better algorithms, chips, quantization, training efficiency, more compute, and more data. Character AI’s product philosophy (Priority: 5/5): Character AI is positioned as a flexible, user-controlled platform for roleplay, companionship, and information, with the belief that users should define the use case. Safety, hallucinations, and emotional use cases (Priority: 4/5): The team treats fictional output as a feature but adds safeguards against self-harm, harm to others, and pornography, while prioritizing memory and personalization. AGI as a side effect of useful products (Priority: 5/5): Shazeer says the company is product-first and AGI-first at once: improving the product depends on improving the model, which may incidentally lead toward AGI.

Key Arguments: Deep learning worked because it aligns with modern hardware, especially GPUs/TPUs optimized for matrix multiplication and parallel computation. Transformers beat RNNs because they replace sequential processing with parallel attention over the whole sequence. Language modeling is the best core AI problem because next-word prediction is simple to define and can scale to enormous data. There does not appear to be a hard architectural ceiling yet; model quality should keep improving with scaling and efficiency gains. Data scarcity is not viewed as a fatal bottleneck because humans generate massive amounts of language, and AI can generate more data too. Character AI succeeds by letting users create personas rather than forcing one generic assistant persona. Hallucinations are partly a product feature in Character AI because the system is explicitly fictional and roleplay-oriented. Memory and personalization are key product needs because users want characters to remember them and adapt to them. The company’s path to AGI is indirect: build a useful consumer product, and the model improvements needed to make it better will push toward more general intelligence.

Data Points: Years at Google: 17 years - Referenced as Shazeer’s long, on-and-off tenure at Google. Google Brain start: 2012 - He joined Google Brain around this time. Transformer paper year: 2017 - Mentioned as the period when the Transformer paper was developed. Mesh TensorFlow timing: 2018 - Referenced as work following the Transformer paper. Character AI funding: $150 million - The company recently raised this amount from Andreessen Horowitz and others. Character AI team size: 22 employees - He notes the company is still very small and prioritizing hires. Average daily active usage: about two hours - He says users who message on a given day spend around this much time on the site on average. Initial project naming: 20% project - Daniel DeFreitas started the precursor project as a Google 20% effort. Estimated internet-scale language generation: order 10 billion people producing 10,000 words a day - Shazeer uses this rough estimate to argue that vast amounts of data will keep appearing.

Pivotal Quotes: "The biggest thing is just convince everyone that this is like worth trillions of dollars by demonstrating some application that is clearly super valuable to like billions of people." — Noam Shazeer: Explaining why the team moved from internal models to Character AI and consumer deployment. "AGI is a side effect." — Noam Shazeer: Describing the company’s philosophy that useful products and model improvement can indirectly produce AGI. "The real thing is, like, I want to drive technology forward." — Noam Shazeer: Explaining his motivation for working on AI beyond fun and company-building.

Implications: The interview suggests that the fastest route to more capable AI may be building highly engaging consumer products that force continual model improvement. It also implies that personalization, memory, and safe fiction-like interactions will be major differentiators in AI products and a practical bridge toward more general intelligence.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors