The TWIML AI Podcast
The TWIML AI Podcast

This Week in ML & AI - 7/15/16: A Wingman AI for Pokémon Go and Wide & Deep Learning at Google

This Week in Machine Learning & AI brings you the week’s most interesting and important stories from the world of machine learning and artificial intelligence. This week's show features a conversation about public datasets, an AI-powered Pokémon Go Wingman, a new deep learning app for your

Topics Discussed

Episode Summary

Executive Summary: This episode surveys major AI/ML developments in mid-2016: the ethics of public data access and ownership in NHS/DeepMind research, a Pokemon Go Messenger bot as an example of natural-language UI, consumer and startup AI products like Prisma and Algorithmia, several funding/acquisition updates, and Google’s wide-and-deep recommender model. The recurring theme is that AI’s technical progress is outpacing policy, product design, and public understanding.

Main Topics: Publicly funded data, ownership, and value capture (Priority: 5/5): The host revisits the NHS/DeepMind eye-scan arrangement and argues that the central issue is not only access to public data, but who captures the economic value generated from it. He raises the possibility of new data licenses that force derivatives like trained models to remain open. Pokemon Go bot and natural-language interfaces (Priority: 3/5): A Facebook Messenger bot that answers Pokemon Go strategy questions is used to discuss whether natural-language interfaces are actually better than apps or web pages. The host concludes that a good LUI may be easier to build than a good GUI, while a bad LUI may still outperform a bad GUI. AI products going mainstream (Priority: 4/5): Prisma is highlighted as a standout consumer app bringing neural style transfer to everyday users, with the host noting that its success may come from hiding complexity rather than exposing it. Google Photos is mentioned as an even broader, more impactful example of consumer AI. AI business activity and “AI washing” (Priority: 4/5): The episode covers acquisitions and funding rounds involving SalesPredict, Amplero, and PhiveAI, while criticizing vague marketing language around machine learning products. Facebook’s data-center press tour is framed as publicity rather than real news, illustrating hype around AI branding. Cloud-based machine learning and framework ecosystem (Priority: 4/5): Algorithmia is presented as a platform for deploying trained models as services, with support for major deep learning frameworks. The host also notes framework popularity trends, benchmark work across CNNs and hardware, and tools for cheaper GPU experimentation. Google’s wide and deep recommender model (Priority: 5/5): The episode’s main technical deep-dive explains how Google combines wide (memorization-heavy linear models) and deep (generalization-heavy neural nets) approaches to improve recommendation quality and scalability, especially for the Google Play Store. Projects and open research culture (Priority: 3/5): The episode closes with several practical projects—cat-detecting sprinklers, a TensorFlow-powered robot, and NLP course project reports—showing how accessible deep learning had become for builders and students.

Key Arguments: Publicly funded datasets should not automatically become exclusive assets for one company; governments and public institutions should think harder about maintaining public ownership and broad benefit. A new licensing approach may be needed for data, one that extends copyleft-like obligations to models trained on the data and requires publication of model source code if the model is commercialized. Natural-language interfaces are valuable when they simplify access to information, but their real advantage is often in replacing bad interfaces rather than outperforming all other UI forms in absolute terms. Consumer AI succeeds when it feels ordinary; Prisma’s appeal comes from making advanced style-transfer look like a simple photo filter. Much of the current AI market activity is driven by hype and vague language, so readers should distinguish genuine technical innovation from “AI washing.” Google’s wide-and-deep approach works because linear models are good at memorization and neural networks are better at generalization, so jointly training both captures the strengths of each while keeping scalability. Operationally, deep learning systems are becoming practical at very large scale, as shown by Google Play’s huge training set and low-latency serving, along with cloud tools that make model deployment easier. Open-source implementations, benchmark projects, and student reports are accelerating the spread of applied machine learning knowledge across the community.

Data Points: Podcast date: Friday, July 15th, 2016 - Episode introduction NHS eye scan dataset size: 1 million images - Referenced in discussion of Google DeepMind access to Moorfields/NHS data Wrangle Conference date: July 28th, 2016 - Sponsor mention for Cloudera event in San Francisco Newsletter signup URL: twiml.ai/newsletter - Listener interest in a planned email newsletter SalesPredict acquisition: Announced by eBay - Business news on machine learning acquisitions Amplero Series A: $8 million - Funding round for marketing-focused ML startup PhiveAI seed round: $2.7 million - UK startup funding for autonomous vehicle platform TensorFlow ranking: Leading in every area but GitHub issues - Referenced from François Chollet’s framework popularity tweet Google Play recommendation scale: 1 million item app catalog - Wide-and-deep paper discussion Training examples: Over 500 billion - Google Play recommender system scale Serving latency: About 10 milliseconds - Google Play recommendation serving time Peak request load: 10 million app scoring requests per second - Google Play recommendation system capacity Framework release: DeepLearning4J version 4.0 - Java deep learning framework update GPU benchmark coverage: Multiple GPUs and Intel Xeon with and without CUDA - Justin Johnson’s CNN benchmark project NLP course project count: About 100 reports - Stanford CS224d project reports

Pivotal Quotes: "who's to say that if the data weren't public, another organization, such as a public university, wouldn't take up the research challenge?" — Sam Sherrington: Questioning the assumption that only a private company can unlock value from public data "the one with the slickest sales pitch." — Natasha Lomas (as paraphrased by the host): Critique of how public resources can be transferred to corporate interests "linear models are great. They're easy to use, easy to scale, and easy to understand." — Sam Sherrington: Explaining why the wide component matters in Google’s wide-and-deep recommender

Implications: The episode highlights a turning point where AI is becoming operational and consumer-facing, but governance, openness, and honest product claims lag behind. Expect more debate over data ownership, model licensing, and how to separate real ML value from hype.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast