Episode Summary
Executive Summary: Ahmad Mostak, CEO of Stability AI, explains how his personal experiences led him from finance into AI for social good, ultimately motivating Stable Diffusion and an open-source foundation-model strategy. The discussion covers the technology’s rapid compression, speed gains, creative and enterprise use cases, model governance, safety, and why he believes open access is essential for AI infrastructure.
Main Topics: Ahmad Mostak’s path into AI and Stability AI’s origin (Priority: 5/5): Mostak describes moving from math/computer science to hedge funds, then into AI after his son’s autism diagnosis and later humanitarian and public-health projects. This trajectory shaped his belief that powerful AI should be broadly available and socially beneficial. How Stable Diffusion emerged and why image models matter (Priority: 5/5): He recounts the evolution from CLIP/VQGAN experiments to latent diffusion and Stable Diffusion, arguing image generation was the right frontier because visual communication is central, yet still underdeveloped compared with text. Model efficiency, speed, and the compression of knowledge (Priority: 5/5): A major theme is that Stable Diffusion is small, fast, and increasingly efficient, challenging the assumption that bigger models are always better. He frames diffusion as compressing massive datasets into portable, useful systems. Creative, design, and media disruption (Priority: 4/5): Mostak highlights near-term disruption in art, concept design, video production, animation, industrial design, and presentation creation, where generative image systems can replace slow, expensive workflows. Open source, accessibility, and global equity (Priority: 5/5): He strongly argues that foundation models should not be controlled by a few corporations, claiming open source improves security, fosters innovation, and helps close the digital divide across countries and communities. Safety, bias, and opt-in/opt-out governance (Priority: 4/5): The conversation addresses safe-for-work filtering, dataset curation, attribution, and artist consent. Mostak says Stability aims for ethical use while still keeping models open and community-driven. Stability AI’s platform and enterprise strategy (Priority: 4/5): He defines Stability AI as a platform company building a foundation-model layer, with research, productization, enterprise deployment, and infrastructure support. Enterprises are framed as needing custom fine-tunes and consulting-style help rather than training from scratch.
Key Arguments: Foundation models are too powerful to be controlled by a single company; open source is both ethically preferable and more innovative. Image generation is a uniquely important AI frontier because visuals are central to human communication and still far from solved. Model size alone is not the key driver; better data, better conditioning, and efficiency matter more than simply scaling parameters. Stable Diffusion demonstrates that huge datasets can be compressed into small, fast models that run locally on consumer hardware. Enterprise customers should not rush to train foundation models from scratch because the architecture is still evolving quickly and training is expensive. Open access helps reduce digital inequity so people outside wealthy tech centers can benefit from AI. Governance should be community-based and democratic, with local and national participation rather than centralized control. Safety can be improved through dataset choices, filters, and opt-in/opt-out mechanisms without abandoning openness.
Data Points: Stable Diffusion launch date: August 23 - Mostak says Stable Diffusion was released on August 23rd. Stable Diffusion 1 generation time on A100: 5.8 seconds - He compares early generation speed at launch on an A100 GPU. Stable Diffusion generation time as of the interview: 0.86 seconds - He says performance had already improved dramatically by the time of the conversation. Projected speedup: 20x faster in two weeks - He claims a forthcoming model release would be substantially faster. Stable Diffusion 1 file size: 2 gigabytes - He describes the model as a compressed output of very large image-text datasets. Input data size: 100 terabytes / 100,000 gigabytes - He says the model was trained on roughly 100 TB of image-text pairs. Potential optimized model size: 400 megabytes - He speculates the model could be compressed further and run on an iPhone. Stable Diffusion parameter count: 890 million parameters - He contrasts Stable Diffusion with much larger text models. Stable Diffusion 2 parameter count: 900 million parameters - He notes the next version is around this scale. MacBook M2 generation time: 18 seconds - He says an M2 MacBook can generate an image in 18 seconds as of today. Stable Diffusion 2 release timing: 1 week old / last month - He states SD2 was released very recently, using informal timing during the interview. Hugging Face developer count: 380,000 developers - He cites this as the scale of developer adoption around Stable Diffusion. GitHub stars comparison: Overtook Ethereum and Bitcoin - He claims Stable Diffusion’s GitHub stars surpassed those projects by Monday. A100 cluster size: 4,000 to nearly 6,000 A100s - He discusses Stability’s cloud infrastructure scale. Training efficiency improvement: 103 to 163 teraflops per GPU - He credits AWS SageMaker optimizations for the GPT-NeoX training stack. GPT-NeoX downloads: 20 million downloads - He cites the popularity of the Eleuther AI language model family. UN COVID dataset size: 500,000 papers - He says the UN COVID initiative organized a freely available corpus of research. X Prize for Learning: $15 million - He mentions winning the prize for an offline literacy and numeracy app. NFT sale by his daughter: $3,500 - He recounts his daughter selling an AI-generated image as an NFT. Enterprise shoot replacement: $113,000 avoided cost / 3 hours / 2,000 shots - He says a fine-tuned model replaced a costly photo shoot workflow. Creative economy size: Hundreds of billions of dollars per year - He argues generative AI will disrupt creative industries at this scale. Video game industry size: $170 billion - Used as an example of a major sector affected by image generation. Movie industry size: $80 billion - Used as an example of a major sector affected by image generation. Open source language community size: 15,000 people - He references Eleuther AI as a large open community. Safe/unsafe data opt-in/out balance: 50-50 - He says artist signups for data participation have been split evenly between opting in and opting out.
Pivotal Quotes: "You can't have them controlled by any one company. It's bad business and it's not the correct thing ethically." — Ahmad Mostak: Explaining why Stability AI was built around open-source foundation models. "I think it'll be great to see these things proliferate so we can have a open discussion about it and also have the value created from just these brand new experiences." — Ahmad Mostak: Discussing the social value of widespread access to generative AI. "We view this community as big. We're creating millions, hundreds of millions of artists." — Ahmad Mostak: On why artist participation, consent, and governance matter in training data.
Implications: The interview frames generative AI as foundational infrastructure, not just a novelty. Expect faster, smaller, more customizable multimodal models, stronger enterprise demand, and continued conflict over openness, safety, consent, and governance.