Pastime | Podcast
Pastime

Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space Sponsorship and business inquiries: [email protected]

www.latent.space

Oct 02 2026 | 01:41:28

Last call for regular tickets for AI Engineer NYC! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for new tickets only, no refunds! See you in 2 weeks!

While we tend to cover industry on the pod, every so often we celebrate a clearly emerging superstar PhD. In 2024 we featured Shunyu Yao, who went on to build Operator at OpenAI and is now Chief AI Scientist of Tencent. In 2025 we featured Jack Morris, who went on to cofound Engram at $600m and is now a leading voice on continual learning.

This year we are proud to feature the work of Alex Zhang of MIT.

From GPU kernels and KernelBench to Recursive Language Models, Mismanaged Geniuses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems.

RLMs took over the timeline early this year:

and an RLM based harness was the first to ~solve ARC-AGI-3 before OpenAI’s Astra:

and is even today, influencing new research that has more extreme implications than RLMs:

We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research at Sakana AI, speculative programmatic tool calling, capability overhang, Neuralese, and where Alex thinks the next big research opportunities may lie.

We discuss:

* Why AI-generated GPU kernels still leave substantial room for human expertise

* How one expert insight can potentially replace enormous amounts of brute-force token search

* Why PhD students should take research bets that initially look trivial, weird, or pointless

* What SWE-bench, RLMs, ReAct, and Quiet-STaR reveal about research taste

* GEV and why a language model does not have to mean an autoregressive text-to-text decoder

* Why Claude Code, Codex, and Pi are structurally more similar than they look

* How harness design can improve compositional generalization across tasks and domains

* RLMs: context offloading, code execution, recursive subagents, and shared memory

* Prime Agent, continual harnesses, and persistent agent-to-agent communication

* Why the model you query in the future may secretly be an entire swarm or scaffold

* OpenAI’s 10,000-agent experiment, 130B output tokens, and ~$40M-equivalent problem solving

* Why much of an agent swarm may be wasted search — and why convergence is still hard

* Kimi versus OpenAI and different approaches to multi-agent systems

* Open-endedness, Sakana AI, and finding hidden gems in enormous amounts of generated work

* Why current frontier models may already have a large capability over…

Sep 30 2026 | 00:39:12

Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks:

We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform. However, we were at Anthropic for the Computer Use launch, there for Claude Cowork with the first big podcast on it, organized the first Computer Use track at AIE presenting the state of the art, and were close to the OpenAI-Sky Software acquisition that now powers the complete domination of computer use that Codex enjoys today. This is why we’re excited to bring you today’s first guest, Ari Weinstein, cofounder of Sky and now leading all the amazing CUA progress that casuals might miss:

Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software.

OpenAI clones Jev

In the second half, Nikunj Handa from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API. Given that we were the first Jev podcast, we particularly focus on the unusually fast sprint on the Decisions API:

And why it is just a Luna wrapper for now but the team is motivated and egoless enough to clone what they consider to be good patterns.

We discuss:

* Why OpenAI thinks Computer Use has changed dramatically in just the last few months

* Dots and what changes when every agent gets its own Linux computer

* Why Computer Use can now complete some tasks faster than the average human

* The path from human-level to “literally superhuman” computer use

* Why modern agents are much better at debugging and recovering from failure

* How screenshots, accessibility trees, the DOM, Playwright, and generated JavaScript work together

* App Shots and why they give models much richer context than ordinary screenshots

* Why Computer Use can close the loop between writing software and testing it

* Trust, permissions, and safety when agents can make payments and operate websites

* Async function calling and why models no longer need to stop reasoning while tools run

* Mid-turn steering, WebSockets, and the architecture behind more responsive agents

* UltraFast inference and how OpenAI is pushing frontier models toward much lower latency

* The rapid internal story behind the Decisions API

* Why Decisions API is more than structured outputs at low latency

* GPT Live, fast tool calling, and real-time computer control

* How OpenAI is already using Decisions API for support classification and internal workflows

* Longer prompt caching, cache pre-warming, and cache-aware applications

* Server-side compaction vs manual compaction for long-running agent threads

* What should live inside an Agents API versus a developer’s own harness

* OpenAI as an “AI cloud” and the search for higher-level primitives beyond raw model APIs

Ari Weinstein

* Product & Engineering, Computer Use at OpenAI

* X:

Sep 29 2026 | 01:32:29

We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York, coming up in 2 weeks!

In case you’ve been under a rock, here’s a non-exhaustive list of what Anthropic has been shipping since closing the largest fundraise of all time in May at $47B ARR:

* June: Launched Claude Tag and Sonnet 5 and Fable 5

* July: Opus 5, /checkup. crossed $65B ARR

* Last month: Fable/Mythos 5.1, and EFS (upcoming pod)

* IPO target $2T, end 2026 ARR estimated $100B

* Cowork/chat merged before did

* Claude Mods

* Dario endorses the same Pacing the Frontier message cosigned by all labs

* Last week: Opus 5.5, Plugins portal, Cloud Sessions/Claude Projects

* Today: Sonnet 5.5!

Today’s episode should catch you up, with Thariq Shihipar, the explainer-king of Anthropic, who we last caught up on Fable launch day with The Field Guide to Fable:

The Future of Mutable Software

Pay special attention to Claude Mods (especially the cheatsheet):

In general this is also the inverse of the other viral tweet from Thariq:

Cloud Brain, Local Hands

And give a try to Claude Projects:

The “hands” terminology is not just an analogy for the local/cloud paradigm that is being built up at frontier coding agent companies like Cognition, but is ALSO particularly relevant to the safety systems discussions that we’ll be discussing with Anthropic in an upcoming episode as they prepare to pace to frontier with responsible AI deployment.

For those who want Thariq’s writing tips we teased at the start of the pod, watch the full video here:

From the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic’s Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today, why prompting remains a high-skill discipline, and where Anthropic thinks the agent harness is headed next.

We go deep on Claude Code’s evolving interface: Ask User Question and elicitation, artifacts as persistent generative interfaces, Claude Tag for multiplayer agent workflows, Projects, model effort, implementation notes, and the new Claude Mods system for customizing the harness itself. Thariq explains why Claude.md may eventually disappear, why the …

Sep 25 2026 | 01:20:43

From the earliest days of open-weight models to becoming the neutral routing layer for more than 10 million developers, OpenRouter is one of the clearest bets that the future of AI will be multi-model. In this episode, OpenRouter co-founder & CEO Alex Atallah, with AMP’s Anjney Midha returning with swyx to unpack how OpenRouter emerged from the first wave of Llama, Alpaca, Mistral, and Midjourney, why model diversity mattered before it was consensus, and how a company dismissed as “just a wrapper” became critical infrastructure for the AI ecosystem.

We go deep on the product and distribution lessons behind OpenRouter: why model labs can spend billions training a checkpoint and still struggle to get it into developers’ hands, how Mistral helped prove the value of a competitive inference marketplace, why OpenRouter chose focus over expanding into fine-tuning, memory, and other adjacent products, and how its rankings became a real-time map of how AI usage was changing. Alex also explains OpenRouter’s early experiments with model fusion, why they deleted the first version and brought it back years later, and how the platform grew to more than 10 trillion tokens per day.

Finally, Anjney explains why Stripe and OpenRouter fit together, why token fraud may become one of the defining security problems of the AI economy, and why the next wave of fraud won’t just come from humans but from autonomous agents attacking increasingly valuable token flows.

We discuss:

* Why OpenRouter bet early that no single AI model would win everything

* Alpaca, Llama, and open models becoming impossible to ignore

* Why Discord’s early AI deployments exposed the limitations of closed models

* Why model labs can spend billions on training and still fail at distribution

* How OpenRouter became a neutral distribution layer for model developers

* Why VCs dismissed OpenRouter as “just a marketplace” or “just a wrapper”

* The Mistral price war and the first real proof of an inference marketplace

* How Midjourney scaled through Discord and what it taught the AI ecosystem

* Why crypto infrastructure became a dress rehearsal for generative AI

* OpenRouter vs. LM Arena and why their missions are fundamentally different

* Why focus became one of OpenRouter’s biggest strategic advantages

* Anthropic’s early focus on AI pair programming and coding

* The OpenRouter products that were prototyped but never launched

* MOM, OpenRouter’s early Mixture of Models experiment

* Why model fusion failed in 2024 — and why it works much better now

* How OpenRouter’s leaderboard became a live map of the AI industry

* OpenClaw, auto-routing, and agents reshaping AI usage

* How OpenRouter reached 10+ trillion tokens per day

* Why inference gateways are increasingly becoming targets for fraud

* Why Stripe’s fraud infrastructure is strategically important to OpenRouter

* The coming rise of agentic fraud and attacks on the token economy

* What changes and what stays the same as OpenRouter joins Stripe

Alex Atallah

* LinkedIn: https://www.linkedin.com/in/alexatallah/

* X: https://x.com/alexatallah

* Website: https://alexatallah.com

Anjney Midha

* LinkedIn: https://www.linkedin.com/in/anjney/

* X: https://x.com/AnjneyMidha

* AMP: https://www.amppublic.com/

Timestamps

00:00:00 Introduction

00:02:12 Alpaca, Llama, and the Multi-Model Bet

00:06:04 Discord, Open Models, and OpenRouter’s Origins

00:14:28

Sep 25 2026 | 01:36:21

Earlier this month, world model company Runway introduced GWM Worlds 2, a research preview that “turns high-fidelity video and audio generation into real-time interactive simulation.” Runway calls this an “autoregressive diffusion” model; with autoregressive describing how it generates over time.

One new feature in particular caught our eye: WorldPrompt, a proposed input format for specifying a generated world and the actions within it. It allows you to fix some aspects of a simulated environment — including the first frame — and then create a series of timestamped events. The events, or actions, can even be prompted in real-time.

To understand the implications of WorldPrompt, we spoke to Kamil Sindi, Runway’s CTO, and Robin Kahlow, its Principal Research Scientist for generative video and multimodal AI. We also have exclusive comments from Anastasis Germanidis, co-founder & co-CEO of Runway, courtesy of a podcast swyx and Vibhu did with him.

Who’s building real-time interactive world models?

First, some context about world models that can generate interactive video and audio in real-time.

Runway is reportedly valued at $5.3 billion, based on its most recent fund raise of $315 million in February. Its first release, GWM Worlds, was launched last December.

Alongside Runway, there are several other notable projects in this domain: Google DeepMind’s Genie 3 (which also generates at 720p and 24 fps), Odyssey-2 Pro, and World Labs’ RTFM (Real-Time Frame Model). We’ve summarized their differences in the following table:

Given the complexity and massive latency demands of real-time video and audio generation (which we’ll get into below), all of the projects listed above have limitations. For instance, Google notes that Genie 3 “can currently support a few minutes of continuous interaction, rather than extended hours.”

But as our interviews with Runway show, real progress is being made.

The central idea of WorldPrompt

WorldPrompt, a new feature in GWM Worlds 2, helps differentiate Runway from its competition. You can think of it as a control layer for characters, cameras and the environment. As Kahlow put it, it’s a way to “control all the different subjects in the world” — similar to a computer game.

“Like, if there’s an NPC [Non-Player Character] somewhere, the NPC might walk up to you and say something. So you could achieve the same thing with this kind of model, where you can have very detailed control over everything in the scene.”

As the name suggests, WorldPrompt is a prompting mechanism — not a programming language. So, unlike virtual world games like Minecraft or Roblox, GWM Worlds 2 doesn’t offer scripting capabilities or the ability to control state. But there’s a power to that, as Sindi pointed out.

“You can create promptable worlds on-demand with video and audio in sync, across all these different domains and environments. That’s not a distant-future hypothetical thing,” he said.

But there are also limitations to prompting a world model. We asked how reliably the model would follow an instruction to create, for example, a law of gravity or a certain ability in a character?

“Yeah, so it’s a research preview,” Kahlow replied. “So it’s not perfect, of course, and there are still flaws. It really depends on how difficult the action is. I would say movement works quite reliably.”

Sindi added that more training plus scaling the data and models is resulting in “better following.”

How a video model becomes a real-time runtime

Despite the current limitations of GWM Worlds 2 — especially if you compare it to pre-designed and scriptable worlds like Minecraft or Roblox — the true promise of world models like Runway is that they’ll eventually lead to

Sep 23 2026 | 01:31:58

The OpenAI → Hugging Face attack has people asking “what else do we need to worry about?” and Anthropic’s filters flag two things: cyber-security and biology. The natural question is: what about bio-security, then?

Clem Delangue argues that cyber-warfare defensive capabilities need to be open and to keep pace with frontier models’ attack capabilities

Radical Numerics co-founder Eric Nguyen sat down with us and explained why the same models that increase biological capability can also keep defense from falling behind.

Building a virus from scratch

While he was at Stanford, Eric couldn’t get traction on Genomic Language Models (GLMs) for a long time. Biologists didn’t believe it would work, didn’t think they could verify the output, and didn’t see important applications beyond what they could already do. He kept pushing, eventually helping lead the development of Evo and contributing to Evo 2 at Arc Institute. Those models were later used by a separate Arc/Stanford team to generate entire bacteriophage genomes that were synthesized into functional viruses!

Long context unlocks biological intelligence

Early ChatGPT spit out poems and email, and early DNA language models like Evo and Evo-2 could build a genome from scratch. DNA is different, however, from natural language in that it has a very small alphabet (4 characters ACTG) and that its sequences are very long:

* 60K for an average human gene

* long being up to 2.3M

* the whole human genome around 3B.

Innovation in long-context models made this possible about 3 years ago (footnote: striped hyena), long before the frontier labs were building 1M+ context models.

Now Eric and other AI x Bio luminaries have founded Radical Numerics to build and scale GLMs to tack a wide range of biological problems, extending well beyond generating DNA.

Thinking in DNA

Their GLMs already do pretty well with RNA and protein because there are clear markers in the DNA sequence for genes (RNA sequences the perform many functions) and specific genes that encode proteins. This means that the models already generalize to multiple “languages,” before even attempting to train in other modalities, such as 3d protein structure, epigenetics and natural language.

If a model thinks in the DNA language, maybe it understands the imprint that environment left on different genomes as well? Perhaps the model has learned the functional relationship between different sequences, and could extrapolate to new sequences based on that?

And so what we wanted to showcase was that if we show the model progressively better RNAs in a series of steps with its score, right? So you have like low scores first and then you gradually move up the chain. Can the model continue that trajectory on its own? And then in the final step, does it self optimize to a point where it's like the best score it can get? That was the experiment. Can we do that? And so we took a data set, a large data set of aptamers. We held out a portion of the best performing ones and we showed it only the lower ones, but then we ranked it, right? So we showcase lower scores with the RNA aptamers and then progressively got higher, and then ask the model to just like continue with that pattern. And it turns out it was able to recapitulate some of those higher scores that we had not shown it yet.

So, voila: chain-of-thought, thinking in DNA!

The arms race

But much as long-context inference, chain-of-though and multi-modal perception unlocked sophisticated reasoning in natural language LLMs, these capabilities in GLMs are enabling increasingly sophisticated “biological intelligence,” and along with it, greater danger.

According to Eric, defense is currently losing this battle, but Radical Numerics argues to push the frontier harder!

I won’t spoil the details for you. In the episode we talk in detail about:

* Biosecurity as an arms race — and how defense can keep up

* The genome as the imprint of the environment on DNA

* Going truly multi-modal

* How chain-of-though works when you “think” in the language of DNA



This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

Sep 22 2026 | 02:01:26

How often do you get to talk to a guest who has both an Academy Award and who invented textbook machine learning algorithms? John Platt has an Oscar, two textbook algorithms, two named asteroids, and an Erdos-Bacon number of 6. This was easily the most fun bio of all the guests we’ve read to date. And the result was an epic and fun chat covering Google’s Empirical Research Assistance (ERA), how AI can help battle climate change, and tons of great stories about the co-evolution of science and AI.

John’s colleague Dave Bacon likes to tease John that his career has been defined by being twenty years early to the next big thing. This may be convolutional neural networks (some credit him with coining the term), fusion research, quantum computing. John and Google have been working on solving some of humanity’s hardest problems with AI and computation for well over a decade now. Recently John and his team set their sights on using AI to solve any scientific problem that can be written down as a score.

Google’s Empirical Research Assistance (ERA)

John’s team has taken on many hard scientific problems over the years. In solving these, they noticed a pattern, many scientific problems can be reduced to what John calls a “scoreable task”. Once you have the score function, the goal is to find some code that maximizes the score. The hard part is in formulating the score, but once you have the score finding the maximizer can still be quite a lot of effort.

John’s team set out to automate solutions to this general problem. This came out of the idea of an “auto-Kaggle” AI, which can solve any Kaggle problem you can throw at it. Kaggle is owned by Google, so all the data was ready and easily available to them!

The result is Google’s Empirical Research Assistance or ERA (paper, github, blog). ERA is surprisingly simple conceptually. Gemini (or your LLM of choice) keeps a running tree of past experiments (notebooks) and where they’re going. It’s a close cousin of Monte Carlo Tree Search: at each iteration the Upper Confidence Bound rule picks which notebooks are most promising to mutate. This is optimistic, not greedy, so sometimes even the fifth-best notebook gets chosen. Gemini then proposes mutations for each one, about ten at a time. The history of each branch is shared, so different leaves can learn from each other.

“It’s almost like having a hyper-eager grad student who doesn’t sleep.”

Evolutionary algorithms have been around since the 70s, but this works because Gemini actually knows where to look! What’s even more interesting is that there was a step change between Gemini 2.0 and 2.5, and this went from just not working to working great.

ERA is so powerful that John and his team solved many outstanding problems with it, resulting in at least ten papers. Some of these were climate change related, which we talk about in the next section.

So, we had to ask: if you have an optimization god how do you avoid fooling yourself? John’s answer is that ERA provides predictive models. It’s up to the scientist to make sure they’re truly descriptive. Some of this just involves good old-fashioned careful machine learning science. “It’s a power tool. It can slice your fingers off.” This led to some fun discussion about Kaggle competitions, and the fun ways people can overfit to datasets without meaningfully solving the problem you actually care about: Google’s contrail-detection competition was won by entrants who noticed a half-pixel error in the labels (is the origin at the corner of the pix…

Sep 21 2026 | 02:20:53

Tickets for AIE NYC now open, and apply for the invite-only AIE CODE. Join us!

We have an unusual relationship with today’s guest: for years since coauthoring the InstructGPT paper, Diogo Almeida had been saying that API-available frontier models have been going down the wrong path, everything from the alignment to refusals to reliability perspectives, that we have dropped every mode other than autoregressive chat-tuned LLMs because of the overwhelming success of ChatGPT.

In a launch video now viewed ~40M times (by comparison, GPT4o was 22M, Fable 5 was 15M, Navier Stokes was 74M, and 6 Astra was 137M), Diogo introduced Jev and it immediately took over the AI timeline — we’ll skip full Jev explainers because your favorite AI influencer/educator has probably already done one. We also collected:

* the official patterns and cookbooks you should see first, from Allie

* Jev usecases

* speed based - games and computer use

* the voice + computer use example we discuss at 1h34 mins

* voice + browser control

* The must not miss Doom demo

* Driving cars in games

* Excalidraw

* virtual try-ons

* “Smart Games”/smart NPCs

* guided responses in text messages

* Jev for coding agents has an official guide

* jev for linting

* compacting tool calls

* reasonable pushback from Theo - Diogo has published a note on the Tyranny of the KV Cache that you should read as a followup after the pod for Jev + coding agents, because of his belief that Cache Rules Everything

* Programming Languages built atop Jev (Diogo’s fave)

* Jev for analytics replay and user journey review

* “dark data”

* entity resolution…

Sep 16 2026 | 01:26:14

AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real insurance:

From being Anthropic’s first product hire to building the standards, testing, and insurance infrastructure meant to make frontier AI deployable, Rune Kvist is betting that the biggest constraint on AI adoption won’t be capability it will be trust. In this episode, the AIUC cofounder joins swyx and Vibhu to announce a new $40M round and explain why companies like Cursor, Harvey, Lovable, and ElevenLabs are increasingly confronting a problem that gets harder as AI gets better: who is responsible when autonomous systems fail?

We go deep on AIUC-1, the emerging standard for agent security, safety, and reliability; how AI agents are stress-tested for jailbreaks, hallucinations, and data leaks; and why Rune thinks standards and insurance could become critical infrastructure for AI. We also discuss the growing trust gap between governments and frontier labs, AI-enabled cyber and biological risks, why every model can ultimately be jailbroken, what happens when a $20 coding agent causes $200M of damage, whether AI engineers should be certified, and why even after AGI there may be one job the labs can never do themselves: be their own watchdog.

We discuss:

* Why risk, liability, and trust may become the binding constraint on AI adoption

* Rune’s path from reading the Scaling Laws paper to joining Anthropic in its earliest days

* What Anthropic understood about scaling, compute, and the future years before it became obvious

* Why Waymo illustrates the gap between AI capability and real-world deployment

* AIUC’s $40M round and work with Cursor, Harvey, Lovable, ElevenLabs, and other frontier AI companies

* AIUC-1: a standard for AI agent security, safety, and reliability

* How agents are tested for jailbreaks, hallucinations, and data leakage

* Why most AI companies optimize the happy path without seriously stress-testing adversarial cases

* Why AI standards may need to update every quarter instead of every decade

* The emerging trust gap between frontier AI labs and governments

* Cybersecurity, child safety, biological weapons, and the expanding frontier-model risk surface

* Why standards and insurance may need to evolve together

* How Lloyd’s of London can insure AI systems and bring trust to enterprise deployment

* What happens if a $20 Cursor subscription contributes to a $200M plane crash

* The Air Canada chatbot case and how AI failures are beginning to clarify legal liability

* Why copyright may be one of the hardest AI risks to insure

* Evals, mechanistic interpretability, monitoring, and models becoming aware they’re being tested

* The impossible CISO mandate: adopt AI fast, but don’t let anything go wrong

* Why robotics will make AI liability dramatically more consequential

* Whether AI engineers should have Level 1, 2, and 3 certifications

* AIUC’s roadmap across agents, frontier models, robotics, and universal red teaming

* Why AGI could become a question of national sovereignty

* Why the labs can never fully serve as their own watchdogs

* The Big Short problem: how do you stop competing watchdogs from racing standards to the bottom?

Rune Kvist

* LinkedIn: https://www.linkedin.com/in/runekvist/

* X: https://x.com/RuneKvist

AIUC

* https://aiuc.com

Timestamps

00:00:00 AIUC’s $40M Round and the Risk Bottleneck for AI

00:01:07 From Scaling Laws to Early Anthropic

00:07:58 Why Trust, Not Capability, Could Limit AI Adoption

00:12:19 Founding AIUC and Building AIUC-1

Sep 14 2026 | 01:32:10

At 1:09:00 we talk about the rise of AI x Finance, and AIE NYC is one month away - our hotel block is 97% sold out, get tix & travel ASAP - we will announce speakers from Bridgewater, Ramp, Coatue, Mastercard, Vanguard, Coinbase, Blackrock, Fidelity, Point72, Capital One, JPMC, Wells Fargo, Bloomberg, A24 (yes the movie studio) Labs, Two Sigma, Apollo Global, and more soon!

From helping pioneer core ideas in NLP to now building AI systems that can automate AI research itself, Richard Socher is betting that the next major step in AI is recursive self-improvement. He is the founder of You.com, AIX Ventures, and now Recursive, which has assembled some of the best open-endedness (& self improving agent) researchers in the world and raised a $4.65B seed round.

In this episode, Richard joins Latent Space to unpack his vision for the “Eureka Machine”: a superintelligence that can improve the process of invention itself, accelerate AI research, and eventually tackle major problems across science, energy, materials, biology, and more.

You can get his book “The Eureka Machine” here!

We go deep on Recursive’s early results, including an AI research system that Richard says outperformed humans and their agents on optimization tasks in less than two days, as well as work on NVIDIA GPU kernels where the system discovered improvements without relying on a team of CUDA experts. Richard also explains why he thinks AI research that currently takes thousands of people and years could eventually be compressed into weeks. These results are summarized in his 20 minute AIE keynote, where we also discuss his 10 dimensions of intelligence:

We also explore the harder questions around increasingly capable AI: reward hacking, whether Anthropic-style constitutions actually work, AI regulation and proposals to “pace” frontier development, open-source models as geopolitical soft power, whether today’s LLM paradigm is enough, and what happens if AI systems eventually begin choosing their own goals. Richard reflects on the rejected research that helped inspire Alec Radford’s GPT, open-endedness, the AI Economist, simulations of entire economies, and his framework for thinking about the upper bounds of intelligence itself.

We discuss:

* The Eureka Machine and Richard’s vision for an AI that can automate invention

* Why Richard is optimistic about superintelligence for science and technology

* Why AI hard-takeoff scenarios may underestimate physical and economic constraints

* The risks of regulating intelligence itself instead of specific AI applications

* Reward hacking and why increasingly intelligent AI makes objective design harder

* Richard’s critique of Anthropic’s constitution and constitutional AI

* Alignment vs. personalization and whose values an AI should follow

* Why open-source AI matters for resilience, competition, and geopolitical soft power

* Why Richard left You.com’s frontier-model work to start Recursive

* Recursive self-improvement and automating the process of AI research

* Whether today’s LLM paradigm is enough — and why Richard is less bullish on world models

* DecaNLP, early prompt-based generalization, and the research that influenced GPT

* Why rejected research can shape entire technological timelines

* Open-endedness, evolutionary approaches, and rainbow teaming

* What happens if AI systems begin setting their own goals

* Why simple objectives like profit maximization can produce dangerous reward hacks

* Recursive’s long-term plan to apply self-improving AI to science

Aug 26 2026 | 01:23:31

A few years ago, Caltech Prof. and co-founder of Accelerated Understanding, Anima Anandkumar set out to develop the first open-source weather model with AI. Talking to experts in the field, she was met with skepticism. Weather is chaotic, physics simulations are hard, have been developed for decades, and require supercomputers, the data just isn’t there. Despite reservations, Anima went forth and built. Within a year her team had developed FourCastNet, a predictive model that is competitive with the best physics-based simulations available. Thanks to Anima, and her follow up work, anyone can now predict weather accurately over a short timescale using consumer grade GPUs.

In the fifteen or so science episodes we’ve released on Latent.Space, we’ve covered atoms, molecules, materials, biology, and math. Anima is a pioneer in studying physical systems that are continuous. Weather, fusion, and fluid or heat flow are huge areas of science that are extremely difficult to model: they are large, chaotic, and fundamentally multi-scale. This is a field the AI community has somewhat neglected, but one we expect will grow fast. We plan to cover large physical systems more in coming episodes.

One thing you can glean from Anima’s work is that this area of AI resists the scaling ideas that have permeated the rest of the field. The data isn’t there: open source datasets in many of these domains are limited to tens or hundreds of thousands of examples, far from what token-hungry transformers need. Even worse, the resolution that physics demands pushes the context length into the hundreds of billions, so you can’t just throw more tokens at the problem. That isn’t a ceiling though, just a slower road: progress here comes from building in structure and inductive biases. Sorry for all you bitter-lesson-pilled language modelers.

“If each dimension is even a few hundred grid points, which is where industrial scale starts... we’re talking hundreds of billions to even a trillion context length. So forget ever having a transformer for anything of this scale, all of the world’s compute will not be enough.”

The math underneath

To tackle these systems, Anima pioneered a technique known as Neural Operators, one of the most beautiful theoretical developments in AI of the last decade. These allow you to combine data and physical laws to enable multi-scale inputs and outputs. We’re no longer modeling a grid, we’re modeling a function that evolves over many scales. This allows Anima and crew to build in priors based upon physical intuition.

To see how physical priors are still helpful for AI modeling, let’s revisit the problem of weather forecasting on a global scale. The earth is a sphere, which meant that accurate modeling involved using the right basis set — the Spherical Harmonics. Run a weather model on a grid and it blows up fast. Move to the natural basis for the problem and it stays stable far longer, long enough to roll out months ahead instead of days. Anima’s Fourier Neural Operator learns directly in this frequency domain, and its spherical variant powers FourCastNet 3, which models the weather across the whole globe and keeps running stably far into the future.

The physical world is forgiving

Anima explored Neural Operators across other physical domains too, and one striking observation is that the physical world is more forgiving than you’d expect. In fusion, a few thousand samples are enough to predict plasma disruptions, and to do it a million times faster than traditional simulation.

None of this is a rejection of scale, it is a different route to it. Anima ultimately still wants to build a “foundation model for physics”, a model that spans many phenomena and does both simulation and design. You get there by building in the structure the physical world already has, not by waiting for data that will never exist. It is a start, and it will take longer than the token-driven parts of AI, because for the physical world tokens were never the answer.

“All of the things that work with deep learning, let’s take them, but make them a bit more principled.”

Weather is only the beginning

Neural operators and weather modeling were a personal passion of mine, so we’ve spent much of this blog and the episode exploring this work. Anima has done so much more! In the episode, we cover several other recen…

Aug 21 2026 | 01:09:38

When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI’s $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups.

Time to catch up on why this Second Summer of simulation is working!

From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today’s frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth.

We go deep on Simile’s approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs.

We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one.

We discuss:

* How Smallville and Generative Agents led to Simile

* Why Joon’s team asked: “What if we can just recreate the world that we live in?”

* Why useful personal agents require deep models of their users

* Memory architectures, Markdown files, and the limits of prompting

* “Social physics” and behavioral foundation models

* Why web data captures what people say more than what they actually do

* Interviews, transactions, observational data, and randomized controlled trials

* Why predicting the future matters less than understanding how to shape it

* How Simile creates representative simulated populations

* Simulation versus prediction and the connection to Foundation’s psychohistory

* How to evaluate simulations instead of simply stacking LLM hallucinations

* Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy

* Why frontier models can struggle to reproduce real human behavior

* Why good simulations need to reproduce human biases and mistakes

* Post-training models on randomized controlled trials

* Population-level versus individual-level simulation

* Scaling laws for human simulation

* The long-term ambition to simulate all 8 billion people on Earth

* Whether simulations could help solve climate change or detect collapsing democracy

* Thomas Schelling and the history of agent-based modeling

…

Aug 11 2026 | 01:35:19

This January, four big AI × Pharma tools deals were announced at the huge JPM Pharma conference that takes over San Francisco every year. OpenAI-backed Chai Discovery (now worth $4B) was somehow at the heart despite being all of 2 years old.

The Science team is proud to bring you the first podcast with cofounder Matt McPartlon and product lead Neil Patil to tell the full story!

Editor’s note: not to be confused with Chai AI, which was another top pod of ours.

Pharma suddenly doing big AI tools deals

For the non-pharma people, JPM is JP Morgan’s annual conference for pharma deal-making that takes over San Francisco for a week in January with hundreds of side events, etc. It’s a big thing.

Tools deals for pharma are also a big (new) thing: companies that start as AI for Pharma usually end up building their own drug pipelines instead, and the reason is something like this: convincing pharma to use your tool requires proof that your tool works. Proof means good targets, maybe with good clinical validation. If you have that, then it’s easier to raise money (with a known, if long path to commercialization) or sell (e.g payment in biobucks) for a specific target than it is to sell to lots of companies on a promise that it will work across their portfolios.

The “we’ll just partner / build our own drug” optionality proved to be the only good path up until January. What changed? In short, the tools got good enough for drug design teams to trust.

Good-enough-to-trust unlocks the ability to scale discovery: get more, better candidates into the lab and animal trials faster. More screening for toxicity, better delivery, etc. This means that what you push to the clinic is more likely to succeed.

Tools also unlock new capabilities: mechanisms that are very hard or impossible to develop using lab-based discovery. Designing an antibody that precisely triggers a very specific molecular cascade takes many years of trial and error. Designing bi-specific antibodies (that bind to two different proteins) is similarly difficult. Good design tools can unlock this.

RJ: The fact that the quality of the model has jumped means you’re enabling things you just plain couldn’t do. So it’s a step change. It’s not an efficiency argument at all, or not so much.

Matt: Yeah, exactly. It’s kind of interesting, even for us — it took me a while to believe in the thesis, actually. I talked to Josh for months before Chai started... It’s like, can I beat a mouse, and then can I do what mice can’t do? And then how many levels of interaction can you just keep building on top of that?

Everyone playing in the structural / binding space has an angle here, and some will be better than others, but Chai is pointing to a different unlock: getting good molecules right out of the gate (meaning they don’t then need as much lab work) means that the iteration time is faster. This turns science into engineering: you can design your systems to reduce friction and hill climb towards one-shotting molecules all the way to the clinic.

This, per-se, is not a new thesis: a16z articulated a version of this in 2020. What has changed is that structural models became binding models (how well doesn’t this molecule bind to this molecule, aka “binding affinity). Binding models unlock design, which has been steadily improving. Chai’s observation is that for engineering problems the best product tends to win, and good technology is a necessary but not sufficient condition.

Photoshop for molecules

With that in mind Chai has invested heavily in partnerships that allow them to learn from their Pharma counterparts.

What is kind of cool about working so closely and supporting so many of these partners is we get to really learn about what is the stuff that would be helpful in research. So rather than doing research in a vacuum, based on what would hypothetically be cool, we're able to do informed research based on what our partners have just been organically asking us for help with.

— Neil Patil, (Chai product lead)

This means better UX, such as a molecule editor that is more like a CAD or graphics de…

Aug 03 2026 | 01:41:29

Watch the full episode on YouTube:

We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection.

We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:

And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:

Three years ago, inference engineering barely existed as a category.

Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.

In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.

Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.

In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.

We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.

The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.

We discuss:

* What happens when a 200,000-token request enters an inference system

* Cache-aware routing and reusing previously computed KV cache

* Why prefill and decode are increasingly handled by different GPUs

* When dedicated deployments become cheaper and more reliable than shared APIs

* How speculative decoding uses a smaller model to accelerate a larger one

* Tool calling, structured outputs, and what LLMs actually do

* What it takes to support a new open model on day zero

* Grafting Kimi’s vision encoder onto GLM-5.2

* Retrofitting inefficient model layers with components from other architectures

* Why models sometimes collapse into repeating the same token

* How hardware, kernels, and race conditions create nondeterministic failures

* Preserving model fidelity while making inference faster

* How quantization errors can cancel each other out

* Why inference optimizations still deliver gains of 20%, 100%, and 200%

* How optimized serv…

Jul 28 2026 | 01:09:28

There are roughly 100x more people who use code than who can write code. As code that “just works” becomes easier to generate, this group may be the biggest prize of all — if you can get the agentic interface right.

A key trend we have been tracking over at AINews is the absolute explosion in Codex usage this year, with MAU now up >10x from Jan 2026. Less than two weeks after their July 9th launch, OpenAI said ChatGPT Work and Codex had reached 10M users combined (as we cover in the pod, Codex now powers ChatGPT Work, so all ChatGPT Work users are now users of the Codex harness, even if they aren’t traditional engineers) — showing the early innings of what happens when you graduate from coding agents to knowledge work agents:

We’ve been calling out how coding agents are “breaking containment” to do everything else this year to power every other part of knowledge work - and it started with the org chart, with a major reorg last month that amounted to two of Codex’s most prominent leaders, Greg and Tibo, taking responsibility over product and ChatGPT specifically, completing a “Superapp” consolidation cycle first discussed in March.

With these updates Codex is no longer just a coding tool. In June, OpenAI said knowledge workers already accounting for roughly 20% of Codex’s user base and growing more than 3x as quickly as developers. A product dedicated for knowledge workers was being pulled out of the Codex team.

However, knowledge work has a different set of problems and environments than coding. For decades, knowledge work has been scattered across different primitives like documents for writing, spreadsheets for analysis, slide decks for communication, and specialized applications for everything else. ChatGPT Work now enables users to work across every primitive with agents. Instead of opening an application and manually operating its features, the user can describe an outcome and collaborates with an agent that can assemble the tools, context, and artifact needed to reach it.

From building no-code products at Airtable to leading Productivity Engineering at OpenAI, Akshay Nathan has spent much of his career trying to make the power of software accessible to people who do not write code. In this episode, Akshay joins swyx and Vibhu to unpack the launch of ChatGPT Work, why Codex unexpectedly took off among non-developers inside OpenAI, and the company’s broader plan to bring useful agents from software engineers to knowledge workers and eventually everyone.

We go deep on the shared agent harness behind Codex and ChatGPT Work, why OpenAI brought the experiences together without making them identical, and how persistent computers, artifacts, Sites, plugins, memory, and sub-agents are changing what people can delegate to AI. Akshay explains why some teams are replacing decks and spreadsheets with interactive websites, how agents can gather context across code, Slack, documents, and local files, and what OpenAI learned from personal-agent products like OpenClaw.

Side note: also don’t miss Abhihek’s sandbox track keynote at AIE, which now powers a lot of the sandboxing for ChatGPT Work… and yes was also broken by an unreleased OpenAI model in the recent HuggingFace incident.

Akshay also reflects on how AI is transforming product development itself: why more people will become generalists with a specialty, why ideas and taste become the bottlenecks when almost anyone can build, why LLMs still struggle to generate genuinely grounded new ideas, and why teams must distinguish increased motion from actual progress.

We discuss:

* Why Codex unexpectedly took off among non-developers inside OpenAI

* Why employees felt like using Codex gave them a new superpower

* The product insig…

Jul 23 2026 | 01:54:33

In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines’ recent release nearly 10 times their size.

Poolside’s recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna’s recent technical report on our paper club:

From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.

We go deep on Poolside’s Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.

We also discuss model-harness co-design, Poolside’s path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside’s $500 million raise, open-source AI, regulation, NVIDIA and TSMC’s influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.

We discuss:

* How Andrej Karpathy’s RNN work inspired Eiso to start building language models for code in 2015

* Why Eiso spent four years and $12 million pursuing an idea before the market cared

* Why ChatGPT felt like vindication and brought Poolside back to open source

* Why Eiso would prefer 100 foundation model companies over an oligopoly of five

* The difference between releasing open weights and publishing genuinely open research

* Why Poolside deliberately built a global research organization outside the Bay Area talent war

* Why model building is ultimately 90% engineering

* The Model Factory: Poolside’s end-to-end system for rapidly training and improving models

* How fewer than 70 researchers run roughly 10,000–20,000 experiments each month

* How Poolside moved from six-month model cycles to five- and eight-week launches

* Why streaming data directly into training unlocked faster experimentation

* How immutable data, versioned code, and reproducibility enable rigorous model research

* Why Eiso wants capable researchers to leave their labs and become Poolside’s competitors

* Why 95% of model building can be reduced to better data or compute efficiency

* Laguna S and why persistence, verification, and backtracking can outperform raw intelligence

* Why smaller models may handle far more knowledge work than previously expected

* Why reinforcement learning will move earlier into pre-training

* Why next-token prediction is still failing to extract enough knowledge from the web

* Why distillation and environments have become the AI industry’s favorite “drugs”

* Why mid-training is really an early form of

Jul 21 2026 | 01:29:47

Bet on information

If test loss flatlines after 1.5B parameters while training loss continues to drop as you scale, that tells you that your model is limited by the amount of information in your data.

Training on a single, smallish data set exposed an information gap: the 3.1B model falls off the scaling trend. Neither parameters nor compute will improve performance past this wall. For predicting changes to gene expression, you need more information rich data.

This is what Chu and Bo’s teams have done, and here is what ~30x the information buys you:

Now we can scale with parameters and training compute! We don’t know how much this effort costed, but we can guess that data collection experiments and infrastructure was a few tens of millions, and compute + headcount + research was a few million. The budget looks like a RL rollout budget, rather than a data rich pre-training one.

We were lucky enough to have the two central figures in this story on our podcast. Taking the lead from Ci Chu and Bo Wang, Xaira Therapeutics is betting that information rich data is the key to AI-driven drug development. Chu was recently promoted to Chief Discovery Officer and Bo to Chief AI Scientist, underscoring just how strategic Xaira considers this bet.

Reverse engineering the human cell

If you had to figure out how a human cell works, what would you do? A good place to start might be by documenting what genes are expressed (e.g. what RNA is floating around) in different kinds of cells, in different circumstances.

That is CELLxGENE, a database of 168M cells built by Chan Zuckerberg Institute that maps each cell to a count of how many times 20K-30K genes were detected in that cell, plus detailed metadata about every cell. A ~4 trillion-entry matrix.

If the Protein Data Bank (PDB) unlocked structural biology models (Boltz Episode, ESM/BioHub Episode), CELLxGENE has done the same thing for Virtual Cell models. Like PDB, CELLxGENE has inspired a zoo of AI models of RNA expression; so much so that RNA expression models have become synonymous with Virtual Cell models. Bo Wang built one of the most influential, scGPT, that became the starting point for Xaira’s new model.

RNA expression ≠ Virtual Cell

Models trained on CELLxGENE describe the relationship between cell types and cell states, but they are not good at predicting what will happen if we make changes to RNA expression. Changes in gene expression are highly correlated, and its is difficult (impossible) to figure out what causes what in most cases.

If you could “turn the dial down” on one gene at a time, however, then you would be able to observe what is upstream and downstream of a given gene. You could tell if A → B & C or B → A & C or B → A, C → B → … If you did this for all of the genes, then maybe you could train a model that could predict what would happen to a cell if you change a gene (e.g. with a drug or a gene edit). Or maybe you could figure out the least invasive way to change a particular gene’s expression.

X-Atlas → X-Cell

This is exactly what Chu and Bo’s teams have done. The data set is called X-Atlas and the model is called X-Cell.

In this episode, we discuss:

* Why the team abandoned autoregression for diffusion

* The CRISPR-based experiments that run millions of tests in parallel, and generate the raw data for X-Atlas and X-cell

* Generalization to real lab experiments in real human cells

* Beating the linear baseline that has outperformed previous models

* Justifying a kitchen-sink of priors, and how that stacks up vs. data and architecture

Bo also shared with us some of the (major) advantages he has as an academic vs. industry leader, and how his labs keep up with the breakneck pace of AI innovation.

Check out the full episode on YouTube, or your favorite podcasting platform!



This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit

Jul 16 2026 | 01:41:04

Imagine a dark warehouse. Racks and racks of devices with wires, tubes, and electronics sticking out. The next AI data center? No. This is Lila Sciences‘ dream for the future of science. A dark warehouse full of AI-guided robotics and lab equipment, cranking out new experiments 24/7, building toward a scientific superintelligence.

Their automated lab is almost hypnotizing to watch. They have floating plates zipping around on Wall-E-esque tracks, used vision-language models to control Windows 95 boxes, and created the world’s largest collection of voided warranties. In the process they’ve built a massive library of scientific reasoning tokens. Over 10 trillion of them, all experimentally validated.

No warranties were voided in the making of this video

To say Lila is ambitious is an understatement. Their goal is a scientific superintelligence wired directly into the wet lab. They are all in on the bitter lesson, and the thesis follows from it: a lab is an infinite token generator. Produce data at scale, and the synergies give you a general reasoner that can tackle any scientific problem. They are committing hard. Biology, chemistry, drug discovery, and materials science, all at the same time. Time will tell if it works, but it is an exciting hypothesis.

In our latest episode we sat down with Lila’s very own Andy Beam (CTO) and Rafa Gómez-Bombarelli (CSO, physical sciences) and went on a journey through the possibilities of AI-run science, almost as wide-ranging as Lila’s goals.

Did we mention they do both materials science and biology? In the same AI science factory? Same time, same lab, same AI. Finally a guest who can settle a long-running debate we’ve had amongst ourselves: is biology or materials science harder?

Watch to find out!

We discuss:

* The internet is spent, science is next. Why Lila thinks the scientific method is the last untapped internet-scale dataset, and why they treat RL as a data generation mechanism with nature as the verifier.

* The lab as a data center. Instruments as nodes on a graph, a magnetically levitating “PCI bus” transport layer between them, orchestration as a slurm queue. Andy is not short on analogies.

* Why Lila insists it is not an automation company. They optimize for flexibility and generalizability over raw throughput, which means humans stay below the API line wherever automating does not pay.

* Your experiment has a runtime. We put Escalante Bio’s question to Andy: if science is the token generator, what is the runtime of your data collection? His answer, in short, is that you cannot make the ribosome go faster. Why Lila bets on fast round-over-round iteration rather than big noisy multiplexed screens, and how Rafa’s team rebuilt a gas sorption measurement to run roughly 2,500x faster.

* What is actually in 10 trillion scientific tokens. Not sequences. Experimentally verified reasoning traces, a kind of data that Andy argues exists on the internet in quantities that round to zero.

* Breadth as a path to depth. Small molecule chemistry priors transferring to metal organic frameworks for carbon capture, and the claim that the general model beats domain-specific models sample for sample.

* If you have the data, what do you need the model for? Sri Kosuri’s koan about the ML-for-drug-discovery business model, and Andy’s answer: the coding model got better because it also read Shakespeare and carnitas recipes.

* The serendipity they want to automate. Emily Whitehead survived the first pediatric CAR-T cure only because the doctor treating her happened to know, from pediatric arthritis, which antibody would blunt her IL-6 response. Roll that dice again and you probably lose her. Breadth is how you stop depending on luck.

* Move 37 for catalysts. Model suggestions for platinum-group-free electrocatalysts that went from boring, to what a 40-paper expert called stupid, to the best performers they have made.

* Six months to in vivo CAR-T data in non-human primates, and the zero-FTE virtual startup commercial model that fell out of it. For context on why that number is startling,

Jul 08 2026 | 00:57:55

We’ve been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we’re excited to conclude this series returning to Modal, which has just raised a monster $355M Series C.

The cloud was built for developers. But agents are now changing that.

The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.

However, agents don’t have that luxury. Now in this new era of agents, everything has to be tighter.

They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:

Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:

At the time, Modal was just a teeny little company with a $17M Series A.

Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate.

In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don’t fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.

We go deep on Modal’s AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal’s capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.

We discuss:

* Why Kubernetes wasn’t built for bursty AI workloads

* How Modal started as a better runtime before becoming an AI cloud

* Why Modal added GPUs before ChatGPT

* The shift from developer experience to agent experience

* Why observability matters when agents are writing the code

* Elastic inference for custom models across audio, video, robotics, and comp bio

* GPU snapshotting, cold starts, and why inference workloads are so bursty

* Why RL rollouts can require 100,000 sandboxes

* DeFlash, speculative decoding, and frontier-level inference performance

* Auto Endpoints and making optimized inference easier to deploy

* What Modal adds beyond vLLM, SGLang, and raw GPU rental

* Modal’s 17-cloud capacity pool and supercloud strategy

* Networked sandboxes, sidecars, private IPv6, and RDMA

* Serverless multi-node training for post-training and research workloads

* Auto-research, model-guided sweeps, and agents launching GPU experiments

* C…

Jul 01 2026 | 01:48:39

This episode has a fun personal twist: There’s a counterfactual world where I was employee #1 at Genesis Molecular AI, the company behind today’s episode. A certain introduction happened a few weeks too late and I had already happily signed at Atomwise, another ML-for-drug-discovery startup. Same problem, different company. I was certain ML was going to transform small molecule drug discovery. Early results were underwhelming. Useful at times, but nowhere near revolutionary. In the last year I’ve seen signs that ML is finally ready to deliver on my convictions from a decade ago. Genesis is one of the places that might have finally cracked this problem. I was super excited to come full circle and catch up with co-founder Evan Feinberg and CTO Sergey Edunov.

If you are at all interested in small molecule drug discovery, we think you will find this fascinating!

In our nearly two hour chat we cover:

* What is small molecule drug discovery, and why is it hard

* Structure prediction as a hotbed of innovation in AI algorithms

* How advances in AI elsewhere have enabled stepwise improvements in predictive power

* How the community benchmarks are essentially calling AI slop good enough

* The Genesis flagship model (PEARL) can routinely hit a threshold that is necessary for real-world applications

* New agentic workflows enabled by these highly accurate models

Read on for more, and also some personal thoughts on the future at the end.

The coolest diffusion research is happening at Genesis

Sergey Edunov came to Genesis from Meta where he led Llama 2 training and Llama 3 pretraining. Sergey was a former physicist who thought he was done with physics after many years of training LLMs. Then, he discovered Genesis, and was blown away with all the novel architecture work they’ve been developing.

It probably surprises no one that modern LLM research has not resulted in fundamentally novel or exciting updates in architectures since almost the advent of the transformer — the entire field is using variants on the same idea that came out in the original “Attention is all you need” paper. Sure, some were quite useful (mixture-of-experts in particular allowed for the massive model paradigm we’re at today), but there was very little conceptually exciting.

“We sort of had to wait for the right primitive to get created, and that turned out to be diffusion… Actually, some of the most innovative diffusion research that’s happening in our field is happening in 3D structure prediction right now.” — Evan Feinberg

The field of 3D structure prediction on the other hand has been a hotbed of research. Genesis’ recent model PEARL (Place Every Atom at the Right Location) is able to understand protein flexibility, and model not just where the ligand goes, but also make small adjustments of the protein so that the two fit better than either alone. The field knew this was missing for a long time, but it was really hard to model until now.

Agentic Discovery

What makes this problem so hard? As Sergey points out, there are 10^60 possible drug-like small molecules. You’ll never be able to search them all, and trying to find the good ones is something like finding a needle in a haystack — except everything except your needle is dangerous.

“There are 10 to the 60 drug-like small molecules in the universe… it’s like finding a needle in a haystack, where everything except your needle is very, very dangerous.” — Sergey Edunov

“Or finding hay in a needle stack might be a more apt analogy.” — Evan Feinberg

Trying to solve the multi-parameter optimization problem is even worse. What makes a strong binder and a molecule with good “ADMET Properties” are oftentimes at tension with each other. For example, a good binder is likely greasy, but a greasy molecule is likely insoluble so it won’t enter the bloodstream and get to where it needs to go!

Genesis’ advances in generative AI have now pushed them beyond the threshold where they believe agentic drug discovery loops are finally possible. We all remember the early days of LLMs. They were great chatbots but terrible agents, as small errors compounded rapidly into uselessness. As LLMs got better, the usefulness of agents rapidly improved. Evan and Sergey argue that their models at Genesis recently passed a similar threshold. Their internal agentic drug-discovery system (code named SAPPHIRE) can now iterate like a chemist: look at and reason about poses, form hypotheses, read literature, use internal tools, create candidates for the next iteration. Combining this w…

Jun 24 2026 | 01:08:52

We’re excited to have Databricks join us at AIEWF, among hundreds of the top companies in the AI Engineer ecosystem. LS subscribers can use their discount to get past the late bird pricing and access over $50k in sponsor offers!

Everyone is still talking about Satya’s Frontier Ecosystems post, but few have actually built a (now $175 billion) frontier ecosystem and cloud like our guests today.

From open-sourcing the layer above coding agents to rethinking databases for the agent era, Databricks cofounders Matei Zaharia and Reynold Xin are pushing the company beyond the lakehouse into a full data-and-AI operating system. In this episode, Matei and Reynold join swyx at the 2026 Data + AI Summit to unpack Omnigent, LTAP, Lakebase, agent security, open formats, Mosaic, and why databases may matter more than ever once AI agents start doing real work.

We go deep on Omnigent: Databricks’ open-source meta-harness for combining, controlling, and sharing agents across Claude Code, Codex, Cursor, Pi, custom agents, and internal tools. Matei explains why coding agents and enterprise agents run into the same problems: portability, collaboration, session history, security, spend controls, and the need for a common API above every harness.

Then Reynold walks through Databricks’ database dream: why CDC is brittle enough to joke that it means “continuous data corruption,” why HTAP has been the holy grail of database engineering, and why Databricks thinks LTAP gets most of the benefits by unifying the storage layer instead of collapsing every query engine. We also cover Databricks’ infrastructure scale, the culture behind rapid prototyping, the difference between tech and enterprise customers, Databricks vs Snowflake, whether vector databases should have ever existed, the Mosaic model strategy, Genie, AI Runtime, RL fine-tuning, and the thesis that traditional software gets rewritten once the data is in the right place and agents sit on top.

Databricks began as a company for the big data era. The origination of Spark from the Berkeley AMPLab which eventually turned into the product Lakehouse convinced enterprises that they didn’t need a separate data lake, warehouse, ML platform, and governance layer. They just needed one open foundation where all of their data could live and be reasoned over.

Since then a lot has changed, but data has only become more important. Data is no longer something you keep track of and analyze ad hoc, it’s the necessary context agents need in order to act. So the framing has shifted from “where do we put all of our data?” to “how do we expose the right slice of state, history, permissions, and business logic to an AI system at the exact moment it’s doing work?”

If frontier model performance becomes commoditized, the durable advantage then becomes the company-specific context around them: proprietary data, governed access, operational state, transaction logs, workflows, and feedback loops. Which makes Databricks positioned perfectly.

Now coming fresh off the

Jun 22 2026 | 01:06:23

AI Engineer World’s Fair regular bird tix will sell out ~today! Join us next week ahead of the Late Bird price hike and get >$40,000 in sponsor credits for attending!

Thanks to the US Government issuing an export control directive on Mythos and Fable, the risks of jailbreaks and (industry term) indirect prompt injection are suddenly the talk of the town, though we have been covering AI security for a few years now, from Hackaprompt to the enigmatic Pliny the Elder.

Zico Kolter, member of OpenAI’s board of directors on the Safety & Security Committee, and Matt Fredrikson, CMU professor and CEO of Gray Swan, co-authored the definitive paper on Indirect Prompt Injections, and Gray Swan were cited authorities on the Mythos model card, directly investigating the exact capabilities that are under scrutiny right now:

We seized the opportunity to ask them the state of AI Red Teaming, and Shade, the adversarial red teaming tool that Anthropic used to evaluate the robustness of their models against prompt injection attacks in coding environments. Shade is part of their overall toolkit covering Simon Willison’s Lethal Trifecta, including Cygnal, an AI guardrails product, and the world’s largest AI Red Teaming Arena, including AIRT celebrity Wyatt Walls.

All of this security tooling, and yet, we’re only staving off the inevitable.

The risks of extremely smart AI increasingly feel like gray swan events: an event that everyone can see coming.

In this episode, Gray Swan cofounders Zico Kolter and Matt Fredrikson join swyx to explain why AI security is not just “cybersecurity with AI,” why agents introduce a new class of vulnerabilities, and why the next major AI incident may be a gray swan: unlikely, but clearly visible before it happens.

We go deep on prompt injection, automated red teaming, model robustness, agent identity, computer-use agents, enterprise guardrails, and the emerging AI insurance/compliance stack. Zico and Matt also explain why frontier models are not automatically safer as they scale, why specialized red-teaming models can now beat humans at breaking AI systems, and why the future of AI security may depend on AI systems attacking, defending, and interpreting other AI systems.

We discuss:

* Why AI systems need a different security mindset from traditional software

* How prompt injection creates a new exploit class for agents like Codex and Claude Code

* Gray Swan Arena and the rise of community red teaming

* Shade: AI that can outperform humans at breaking models

* Why LLMs are an alien form of intelligence that fail differently from humans

* Human vs browser-agent robustness and why humans ranked fourth

* Why eval awareness and capability elicitation matter

* Cygnal: Gray Swan’s guardrail model for policy enforcement

* Why bigger models do not automatically become more robust

* The lethal trifecta: untrusted data, private data, and exfiltration

* Wh…

Jun 18 2026 | 00:59:25

Last 4 days before regular tickets sell out at AI Engineer World’s Fair - this is the single biggest gathering of AI Engineers, Founders, Leaders, and Researchers in the world. Attendees get >$5000 worth of sponsor credits and talk tracks are looking FANTASTIC. Join us!

The AI scaling debate always focuses on the question of “how do we get more GPUs?” but the better question may be: how do we make the most of ones we already have.

The fact that a frontier lab like xAI could be running at sub-10% MFU (Model FLOPs Utilization) is just a hint at what the real problem may be.

For context, older frontier-scale training runs were already much higher than 10%. GPT-3 was around 21% MFU. Gopher was around 32%. Megatron-Turing NLG was around 30%. PaLM reached around 46%. And our guest Anjney says best-in-class MFU today is closer to 60–70%.

It’s not necessarily that xAI is uniquely incompetent (it’s clear they have talented folks) but rather the priorities may be flipped in the GPU arms race.

While GPU access is a bottleneck, simply increasing CapEx won’t automatically translate to better models as frontier AI is increasingly a systems problem: scheduling, utilization, networking, kernels, frameworks, data pipelines, parallelism, cluster reliability, and the thousand small decisions that determine whether your theoretical FLOPs become real training progress.

From building Discord’s developer platform and backing frontier AI companies like Anthropic, Mistral, Black Forest Labs, and Periodic Labs to now building AMP’s independent compute grid, Anjney Midha has spent years close to the real bottlenecks of AI scaling. In this episode, Anjney joins swyx at Periodic Labs to unpack why the AI race is not just about buying more GPUs, why 95% utilization would have been considered an outage at Google, and why the next era of AI infrastructure has to be more aligned, more efficient, and more responsible.

We go deep on AMP’s vision for a compute grid that makes FLOPs flow like megawatts, the difference between full-stack AI labs and horizontal pooling, why AI data centers need community buy-in, and how compute markets could evolve into something closer to an independent system operator. Anjney also explains why DeepMind’s unpublished research points to a market failure, why end-of-life prediction remains one of the most important AI applications he has thought about for fourteen years, and why “output maxing” may become a new discipline for frontier systems.

We also discuss Anthropic’s culture, why “luck favors the prepared mind” in coding models, how Claude cracked coding, why too much capital too early can make AI labs fragile, what Periodic Labs is trying to do with science and superconductors, why great researchers can become great CEOs, and why Silicon Valley is both deeply missionary and deeply mercenary.

We discuss:

* Why 95% utilization was considered an outage at Google

* Why AI infrastructure waste compounds at frontier-lab scale

* Why “move fast and break things” does not work for AI data centers

* How data center backlash, power grids, and community incentives shape AI scaling

* AMP’s vision for making FLOPs flow like megawatts

* Why compute needs an independent system operator

* How interruptible demand and dynamic prioritization worked inside Google

* Why DeepMind research hoarding creates negative externalities

* AMP’s 1.2GW base-load ambition and the need for 6GW of spike capacity

* Why end-of-life prediction could become one of AI’s most important healthcare applications

* Frontier Systems, output maxing, and full-stack alignment

* Why APIs and abstraction layers become lossy as organizations scale

* Superconductors, standards, and the dream of lossless systems

* SF Compute, open protocols, and the future of compute marketplaces

* Why non-NVIDIA chips can still benefit from NVIDIA’s reference architecture

* Trust boundaries and why chip startups need visibility into future model architectures

* Why VCs often underestimate researchers as CEOs

* Scientists as star athletes of the mind

* W…

Jun 17 2026 | 01:16:50

On the Science pod, we’ve been covering a lot of the ground on how AI is revolutionizing STEM, but one of our favorite off the record topics since our launch is which field is harder to accelerate: math, bio, or physics? Today we’re back in Materials Science land with Radical — Unlike biological molecules that can be represented (and predicted!) by token strings, the success of materials involve many more macro complex variables like supply chains, microstructures, and manufacturing processes. If you recall the LK99 drama of 2023, while the basic ingredients were known, part of the confusion came from the lack of disclosure around manufacturing, and therefore defeated reproducibility. There is probably no "one-shot" model capable of designing a material that works perfectly at scale.

How Radical is accelerating materials discovery >10x the pace of DARPA/GE MACH

Joseph Krause is a materials scientist through and through. And after spending his career watching industries stall out waiting for better materials, he founded Radical AI to do something about it.

We recently sat down with Joseph to talk about Radical AI, materials discovery, self-driving labs, and the future of AI science. Joseph did not sugar coat anything: accelerating the materials discovery pipeline is a hard problem. But it’s one that he strongly believes we need to invest in, for the future of consumer products, aerospace, computing, and defense, and get them into every day use:

“We count it as a discovery when you pick up your phone and there’s a new material sitting inside of it.”

How does Joseph plan on accelerating the rate of discovery? To understand this, it’s important to understand why this is such a hard problem in the first place. The first thing to keep in mind is that the material that is manufactured is far more than a chemical formula going into it. The process of mixing, annealing, growing, or generating the final material can result in wildly different outcomes. The entire materials discovery process, both from early discovery to large scale manufacturing, needs to be understood and characterized.

The Self-Driving Lab

This philosophy has grown into a key insight at Radical AI: The construction of the self-driving lab. This lab is one that is not just automated, but in fact uses an “AI scientist” that combines scientific knowledge, computational techniques, and human intuition to generate and test hypotheses in an automated lab. Creating an AI scientist was key to making Radical’s self-driving labs work, since Joseph argues that no single AI model can one-shot materials.

“In materials, the ground truth is the material itself. You have to be able to test it and characterize it.”

Joseph talked at length about the self-driving labs at Radical. Joseph argues that experimental data is the true “moat” in this industry. An SDL functions as a closed-loop system where an AI scientist generates hypotheses, and automated robotics synthesize and characterize materials, running research campaigns in parallel rather than serially.

The successes here were both on the automation side and on the science side. Radical has managed to scale their alloy discovery pipeline up to producing and characterizing 1200 alloys in six months — this nearly 10x speedup over the DARPA/GE MACH program that aimed to create 500 new alloys in a year. Joseph claims they can scale this up even more and estimates they can produce a hundred new alloys tested and characterized in a day. A truly new paradigm in high-throughput alloy experimentation.

On the science side, their AI scientist proposed and tested 300 new materials, ten of which were found to have novel state-of-the-art properties that are already being further developed for commercial applications. The robustness of this first materials campaign reinforces Joseph’s claim that the moat is the lab and data.

“It’s moved into elemental families or alloy families no…

Jun 04 2026 | 01:15:39

The new AIEWF website is live! Get your tickets booked ASAP as they -will- sell out. Take the AI Engineering Survey and get >$2k in credits and free AIE WF tickets!

Most industry benchmarks compress intelligence and reasoning ability into scores.

SWE-Bench Pro, MMLU, Humanity’s Last Exam, etc. These metrics are useful, but don’t always represent the full extent of how a model performs in the real world. Some of the most interesting evals today look less like exams and more like operating businesses in the real world. One of which is Vending Bench.

In Anthropic’s Mythos Preview System Card, Andon was the only third party eval to get their own section, observing increasingly concerning aggressive behavior:

You don’t know what a model is capable of doing in the real world unless you actually give it inventory, a wallet, tools, customers, competitors, humans, & some time. More often than not, it’ll surprise you how much a model is capable of and in doing so, also reveal unexpected behavior: deception, context collapse, emergent coordination, & bizarre negotiation behavior.

While an inflection point in personal agents came post-OpenClaw after full file access with bypass permissions became the norm, it is yet to come for agents in the real-world. However Andon Market, an actual in person store fully run and managed by AI, is paving the way for what is possible.

Full Video Pod

From Claude trying to call the FBI over a $2/day vending machine charge to AI agents forming price cartels, hiring human employees, running physical stores, and writing existential robot musicals, Andon Labs is stress-testing what happens when frontier models stop being chatbots and start acting in the real world. In this episode, Andon Labs cofounders Lukas Petersson and Axel Backlund join swyx and Vibhu to unpack the strange, funny, and genuinely concerning edge cases that emerge when agents run businesses over long horizons.

We go deep on Vending-Bench, Project Vend, Vending-Bench Arena, Bengt, Butter-Bench, Luna, and Andon’s broader mission of building realistic real-world evals for autonomous AI systems. Lukas and Axel explain why dollar-denominated evals reveal things traditional benchmarks miss, how Claude ended up reporting its vending machine fees as cybercrime, why long context windows can drive agents into meltdown loops, what happens when agents compete with each other, and why the future of AI safety may depend on testing models in messy physical environments instead of clean benchmark sandboxes.

We discuss:

* Why Andon Labs started with dangerous capability evals and long-running agents

* Vending-Bench and why running a vending machine is a deceptively hard AI benchmark

* Why money-based evals avoid the saturation problem of traditional benchmarks

* How Claude tried to call the FBI over a $2/day fee

* Why long-horizon agents can spiral into existential and legalistic breakdowns

* Project Vend: putting an AI-run vending machine inside Anthropic

* Why real humans are “out of distribution” for simulated agents

* Claudius, Seymour Cash, and the chaos of AI CEOs

* How a human brie…

Jun 03 2026 | 01:33:04

In 2025, seven-month-old startup Axiom solved all 12 of the problems Putnam exam (scoring 8/12 in the time limit) a prestigious undergraduate math exam. The 12/12 score is better than the top undergraduates (110/120) and the closest AI system that reported a result (DeepSeek 103/120), although it is unclear what the people and other systems would have scored with more time. Nonetheless, the Putnam exam is legendary for its difficulty, with the median score typically being 0 or 1 points. Taken by itself, this seems like a minor feather in the cap of AI; one of a long series of accomplishments by AI systems in elite competitions with humans, starting with Deep Blue beating Kasparov.

Fast forward to mid-2026, and Claude Code is eating the world. In 2024 Anthropic’s bet on code and enterprise looked like a more pragmatic niche play vs. OpenAI’s better models and massive consume scale. Today, Amodei’s all in bet on acceleration via code (images and video be damned) seems prescient.

Despite Anthropic’s growing momentum, however, Axiom CEO Carina Hong sees coding ability as a necessary but not sufficient milestone on the path to AGI. Code arguably pushes the jagged frontier to the point of super intelligence in some domains outside of coding, but there are surprising gaps (link) that Carina believes will bottleneck AI progress. (Stats on math benchmarks).

The informal bottleneck

“Verified AI” sounds like eating broccoli (footnote: I actually love broccoli, but then again, I also believe strongly in Test Driven Development, so ¯\(ツ)/¯ ) and paying taxes, but to Axiom it means something very different. “Verification to me is about scaling brilliance, compounding brilliance,” Carina told us.

It actually took a while for me to understand what she means by this. It sounded like marketing-speak to me, until it clicked. Carina emphasizes an story about legendary mathematician Srinivasa Ramanujan to illustrate the point. When G.H. Hardy finally persuaded Ramanujan to formally prove theorems instead of relying on his (formidable) intuition, it reportedly improved his own capabilities. This is presumably because formally proving things forced Ramanujan to articulate the details in a way that open up new lines of thinking, etc. This is one part of “compounding.”

But formally proving things also allowed others to benefit from his intuition: the proofs are way of communicating an intuition and persuading others that the intuition is correct. This is scaling (more people use the result) and compounding (people can learn from and build on his work).

This is the analogy that Carina wants us to focus on.

Verified Generation

There are two ways that Verified AI shows up: in training and in inference.

But a quick detour: to a first approximation, “Formal Verification” means using type checkers (like for TypeScript, C++ or Rust, but more capable) to verify mathematical proofs that are meticulously specified using a language like Lean (footnote: Formal verification also includes model checking (TLA+, SPIN), SMT-based tools (Dafny, F*, Why3), and refinement-type systems (Liquid Haskell) — many of which don’t look much like “type checking a proof” from the user’s perspective even when there’s a similar logical core underneath. It also gets applied to software and hardware correctness, not only pure mathematics.). It takes a lot of work to translate an “informal” proof (albeit one that most people would not remotely call “informal”) in to a Lean proof (footnote: This is an understatement. Most theorems remain informal because formalization is so hard to do. There has been a great deal of effort to formalize the most important proofs, with mixed results)

You can imagine how this would be (very) useful during Reinforcement Learning: instead of relying on best guesses based on statistics (GRPO, RLHF, etc.), you can just verify the proof is correct using a Lean verifier. This is obviously a much stronger reward signal, akin to compiling code and testing it (which is what is typically done with RL on coding).

The catch: LLM are not (currently) very good at proving things with Lean.

Enter Axiom: While they have not officially reported benchmark numbers besides the 12/12 Putnam result, Carina reports that they have achieved a very impressive 99% (187/189) ProofGen on the Verina benchmark. This benchmark is to generate code and proof of correctness for a series of problems. For context, OpenAI o3 (the last known OpenAI run) achieved 4.9% on this benchmark.

Based on the sparse benchmarking, it’s hard to say what the frontier labs are …

Jun 03 2026 | 00:38:58

We’ve informally heard that Satya is a listener to LS for a couple years now, but it was still absolutely surreal to meet him and do a live pod at Build, together with our friends at No Priors, the leading VC AI Podcast that we also greatly admire!

We covered the MAI model technical takeaways on yesterday’s AINews, so I will focus our recap of Satya’s main messages around three elements:

* Satya’s adaptation of the Bill Gates Line for positioning Microsoft as the Frontier Intelligence Platform — customers must gain much more value from the Microsoft ecosystem than Microsoft itself, by building on multi-model harnesses like OpenClaw and Scout, drawing on the full enterprise context exposed by context layers like Work IQ (heavily dogfooded by his C-suite), and building up private evals and traces as a new form of Token IP

* AI ROI: On one hand, enterprises are having difficult conversations around Tokenmaxxing and Layoffs, and on the other hand, there are serious re-evaluations of the End of SaaS since the Build vs Buy equation has changed so much. Our previous SemiAnalysis guest had… interesting comments on Microsoft’s position on this as the ur-SaaS titan, and Satya had great answers

* Making the Impossible Possible: Kevin Scott’s inspiring framing around what the most ambitious version of applying AI and technology at large to business and social problems, like education and social impact.

Enjoy!

Full Video

Transcript

Voiceover: Welcome swyx, Sarah Guo, Elad Gil,, and Chairman and Chief Executive Officer of Microsoft, Satya Nadella

Sarah Guo: Welcome to a crossover episode of No Priors and Lane Space with Satya Nadella. Um, congratulations on an amazing build. No, thank you so much, and it’s great to be with both of you. I listen to both of you or b- both the podcasts all the time. It’s great to be on it.

Thank you so much. [00:01:00] So you’re just talking about, um, these amazing, uh, announcements from across the Microsoft estate all morning for, I think, three hours. What is the, uh, what’s the most important reflection or takeaway you have?

AI as an Ecosystem Platform

Sarah Guo: I, I’d say there are, uh, perhaps the, the biggest one for me is let’s sort of conceptualize this more as an ecosystem play as opposed to a single model or even a single platform, right?

Satya Nadella: I mean, you know, whatever I... At least for me, having grown up at Microsoft, having seen, whatever, four major platform shifts, uh, I sort of fall into that, um, uh, camp where a platform is defined by fundamentally its ability to create more value about the platform versus what’s captured in the platform. And so if you, you view what’s happening right now, I think this morning’s keynote was how can any company, whether it’s an AI native company or a traditional enterprise company, participate as a first-class participant where they can point to AI they created, [00:02:00] right?

It’s not that they don’t use other people’s AI. Of course they will. But to me, what’s the path? What’s the recipe? How do I do it? What does a stack look like? What does the tooling look like? What is valuable? How do you do that? That’s it. That’s sort of our job to do. Yeah. Ecosystem strategy is, uh, very complicated, right?

Sarah Guo: Because you end up building certain components, partnering for certain components, supporting them. You just announced this big suite of models. Like, tell us a little bit about the, uh, training strategy for Microsoft now. Yeah.

MAI Models & Training Strategy

Sarah Guo: So, so the thing that we wanted to do with the MAI models was to build, and as Mustafa talked about, first of all, a great lineage, right?

Satya Nadella: Starting with pre-training, uh, with very good data quality, uh, doing all the ablations, making sure because in, in some sense it’s becoming even harder to build a clean lineage model just because there’s so much stuff out there, uh, that you truly need to ablate out to be able to have a fantastic [00:03:00] pre-trained model.

In fact, that’s one of the challenges of a lot of the open weight models is they look great on one benchmark or two, but they’re not great on practice. So that’s why, in fact, even in the RFDEs are, they, they are pretty gone really excited about the…

Jun 02 2026 | 01:23:27

I’m excited to work with Microsoft once again as the presenting sponsors of the AI Engineer World’s Fair! We’ll streaming live from MS Build today for a special crossover pod with our friends at No Priors and the one and only Satya Nadella. However we did not hold back with this interview - we asked all the burning questions about uptime and Copilot that we know you have in your minds. Lets go!

For almost two decades, GitHub has been the home of software, where both open source and closed flow, through commits, pull requests, reviews, actions, etc.

This ecosystem flourished as open-source maintainers and contributors would continue shipping code for the benefit of the community. However as coding agents began to ship mass quantities of code - growing 1400% in 2026, it marked a new era that was both extremely exciting and challenging for GitHub.

While these agents help more people ship more projects, they also significantly increase the floor of how much code is shipped, how often it is shipped, how many people commit code, and basically orders of magnitude multiples in every dimension of GitHub infrastructure:

Now GitHub inevitably experiences more pressure on their infrastructure which was originally designed around human developers moving at human speed. This has resulted in a very publicly notable uptime story:

So it begs the question of whether current systems around code can absorb what AI produces. Can CI/CD keep up when every idea becomes a build? Can open source maintainers survive floods of AI-generated slop contributions? Can GitHub preserve the human social contract of software while becoming the operating layer for agents?

Which brings us to the perfect person to answer these questions: GitHub COO Kyle Daigle. In this episode, he joins swyx to unpack what happens when AI doesn’t just autocomplete code, but starts changing how companies operate, how open source works, how pull requests get reviewed, and how GitHub itself has to scale.

We go deep on GitHub’s internal AI workflows: micro-skills, WorkIQ, MCP, Slack, Teams, email, Copilot workflows, the new Copilot desktop app, CLI, cloud agents, and how Kyle uses agents to look backwards across company context before deciding what to do next. Kyle also reflects on GitHub’s history building webhooks, APIs, Actions, npm, Dependabot, and Semmle, why the AI era is breaking GitHub in new ways, how Actions became a general-purpose compute layer, and what Copilot becomes after code completion.

Full Video Pod

We discuss:

* Kyle’s expanded role across GitHub

* How AI got Kyle coding again after years in leadership

* Why GitHub rolls out AI through existing workflows instead of forcing new tools

* WorkIQ, MCP, Slack, Teams, email, and GitHub as company context

* Why massive “mega-skills” are giving way to small, atomic micro-skills

* How AI changes summarization, communications, marketing, and analyst work

* Why former developers in leadership may have a unique advantage in the AI era

* Kyle’s “15 agents on Saturday” workflow

* How Kyle built an AI-generated executive presentation for CRO/CFO teams

* Why AI changes the chief of staff role without removing the human work

* GitHub Actions, webhooks, arbitrary code execution, and secure agent compute

* The npm acquisition, supply-chain security, 2FA, and token invalidation

* Slop forks, vendoring, and whether AI agents change dependency management

* What pull requests become when most PRs come from agents

* Prompt requests, vouching, AI review, and trust in open source

* What counts as a “developer” when AI lowers the barrier to building

* GitHub Spark, low-code, and why GitHub refuses to hide the code

* 14x commit growth, Actions load, databases, monorepos, and availability

* Copilot’s evolution from completion to CLI, desktop app, cloud agents, and SDK

* Context, memory, rules, and making GitHub “act like Kyle wants it to act”

* Ambient AI, OpenClaw, enterprise security, and the new operating system…

Jun 01 2026 | 01:43:26

We’re announcing AIEWF speakers this week! Take the AI Engineering Survey!

Today’s guest Ethan first joined us for the LS Paper Club as the lead on NVIDIA Cosmos World Model, but then joined xAI and built Grok Imagine in 3 months:

He comes back on Latent Space with some nuclear hot takes: that Video Models primarily get their intelligence from LLMs, not from training on video data, and that the next frontier for truly interactive, realtime, long-horizon world models is to work on LLMs (perhaps Interaction Models as well…)

Put it this way: In the near term, the next Sora won’t be a better video model, but a video agent.

Generative Media may more closely follow the evolution of AI coding which went from focusing on one-shot output performance and cost, to multiturn reasoning and planning models for agents and systems that can plan, edit, test, debug, and submit PRs.

At a certain point, coding models got so good that the only significant next step to improve performance was handling the orchestration of these models.

Now as the performance of video models increases significantly across realism, consistency, & prompt adherence while becoming more cost efficient, the next evolution of video generation may also be systems that can plan, generate, edit, critique, and iterate across an entire creative task.

In this episode, Ethan joins swyx and Vibhu to unpack what it actually takes to build frontier image and video systems: data, VAEs, diffusion transformers, audio-video alignment, inference speedups, and the hidden cost of storing and moving massive video datasets. From building NVIDIA’s Cosmos world model to joining xAI as Grok Imagine was being built from zero to one, Ethan He has been at the center of some of the most important work in video generation, multimodal models, and real-time world models.

We go deep on Grok Imagine, how a small xAI team shipped its first multimodal video model in three months, why iteration speed matters more than almost anything in model development, and why many of the biggest gains come from fixing tiny bugs in data and training pipelines.

Flipbook: The future of Videomaxxing

Video agents are almost a sure bet to be the trend in the coming year. We end with a glance at what’s beyond video agents:

Flipbook caused a minor sensation this year when it was released, but most treat it as a fun demo. Ethan takes it very seriously — with the speed and cost of inference coming down every year, the future of custom video JIT UI is closer than you think. We talked about why videogen models may become the front end of AI, how generative UI could replace traditional HTML/CSS, why world models need to be real-time, interactive, and long-horizon, and why the future of video generation may depend more on language models and agents than on diffusion alone.

We discuss:

* Why fast iteration mattered more than meetings

* Why small training bugs can drive huge model quality gains

* Why coding models may make compute the bottleneck again

* How image and video models are trained with synthetic captions

* The role of VAEs and latent space in frontier video models

* Why image models are the foundation for video models

* The tradeoff between temporal compression and real-time interactivity

* Flipbook, Neural OS, and the future of generative UI

* Why future interfaces may go from user intent to pixels

May 28 2026 | 01:08:02

The new AIEWF website is live! CFPs close in 2 days and we will run our first New Engineer Orientation this weekend, get your tickets booked ASAP as they -will- sell out. Take the AI Engineering Survey and get >$2k in credits and free AIE WF tickets!

One of the central tensions in the agents industry is that even while there are major decacorn agent labs like Sierra, Decagon, Notion and Cursor being built up, it is also true that it has never been easier to DIY agents, with a plethora of agent frameworks like LangGraph and Pydantic and Flue, and managed agents from Anthropic and Gemini and Amazon. There has been a wave of companies building their own background agents from Shopify to Stripe to Paradigm to Razorpay, and even Cognition’s friends Ramp have built their own coding agent with other friend Modal.

You’d think Cognition might feel a bit threatened, but they’re not - even after all this, they were way oversubscribed for the $1B Series D they just announced:

Walden Yan, coiner of context engineering and Chief Product Officer/Cofounder of Cognition, invited OpenInspect’s Cole Murray to talk about why the Devin is in the Details.

Full conversation live on the pod today:

In retrospect, async agents were the most AGI pilled bet you could make in 2024 - the models weren’t good enough yet to vibecode, and people didn’t trust AI enough to let it rip, nobody (including early Cognition) was sure about the form factors.

Now it is obvious:

* The first wave of AI coding tools made the developer faster but remain heavily in the loop. Copilor and Cursor’s tab autocomplete are prime examples However, the workflow was still heavily centered around and bottlenecked by the developer’s local workflow: a developer in an IDE, watching the model, accepting or rejecting changes, and pushing code one interaction at a time.

* The second wave was local agents: Claude Code, Windsurf, Cursor’s agents pane: first one and increasingly many terminals all running concurrently.

* The current Age of Async Agents points to a different future focused more on agent orchestration which drives end-to-end development.

According to previous guest Steve Yegge, there are finer-grained