Episodes
-
65: Dream-RSI: The Safety Property They Built and Never Claimed
Dream-RSI (arXiv 2609.14858) is a genuinely elegant idea from seventeen authors at Google, Google DeepMind, the University of Maryland and the University of Virginia: a finished discovery run is a tree in which every node already carries its outcome, so an alt...
-
66: Equally Unproven, Unequally Inspectable: Jev versus Needle 3
Two companies shipped a calibrated typed-decision product within seventy-two hours of each other, and the obvious episode about them is not available: there is no shared benchmark between TypeSafe's Jev and Cactus Compute's Needle 3, and their ground truths ar...
-
64: RSIAgent: The Thing That Improves Is Not the Thing That Does the Improving
Four days after we read a seventy-five-page roadmap asking what genuine recursive self-improvement would architecturally require, six authors from Aether AI, UC San Diego and the University of Illinois Chicago shipped a paper claiming a working instance of it ...
-
63: Jev and the System One Model: A New Calling Convention
TypeSafe AI came out of two years of stealth on 15 September 2026 with a model called Jev and a category name they coined for it, the System One Model, and the interesting thing they shipped is not a new kind of intelligence but a new way of calling one. Inste...
-
62: The Last AI Built by Humans: Reading the RSI Roadmap Before Reading the Title
Thirty-three authors from Shanghai Jiao Tong University, Theseus Labs, Tsinghua, ByteDance, ModelBest, Xiaohongshu, Humanlaya, Agent-Native Research Lab and Shanghai AI Lab published a 75-page survey titled "The Last AI Built by Humans: Toward Genuine Recursiv...
-
61: State of Play: What a Takeover Would Actually Require
The Los Angeles Times asked its readers on Friday whether there is really a ten percent chance AI could kill us all. This third episode on the week of the Coxon resignation answers with the numbers that exist and the parts list that does not. Metaculus puts al...
-
60: China Is Not Waiting: Open Weights, the DeepSeek Harness, and the Three-Month Moat
While Washington spent the week of September 8 arguing about a resignation thread, a repository in Hangzhou crossed 220,000 GitHub stars. DeepSeek Harness is an MIT-licensed, plugin-based agent harness in developer preview that shipped four releases in four da...
-
59: The Coxon Moment: What Is Verified, What Is Reported, What Nobody Knows
On September 8, pretraining researcher Jacob Coxon resigned from Anthropic and posted seven messages on X: the labs are "racing straight to self-improving superintelligence and gambling with our lives," and "the people building AI earnestly believe that it cou...
-
58: Astra, Day Four: The Threshold Was the Mouse
Four days after GPT-6 Astra shipped, the aggregate benchmarks say almost nothing moved. Artificial Analysis scores it 61, identical to GPT-5.6 Sol. Epoch calls its record score "within the uncertainty range" of the existing trend. The people using it say somet...
-
57: GPT-6 Astra: The Capability Jump That Shipped With a Blind Spot
OpenAI released GPT-6 Astra on September 3, 2026, and two things happened at once that had never happened together. A lab rated its own model Critical for offensive cyber capability, built an unprecedented monitoring stack around it, and shipped it. And in the...
-
56: Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems
AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate through multi-agent systems by induc...
-
54: Silent Reasoning
We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query through iterative computation in a high...
-
49: Qwen-Audio-3.0-TTS — Open to Closed, Six Months of Improvement
On July 20, 2026, Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a hosted text-to-speech model in two tiers — Flash for real-time voice cloning, Plus for high-quality synthesis — currently ranked #1 on the Artificial Analysis TTS Arena with a quality Elo of...
-
48: Kimi K3
Moonshot AI has released Kimi K3, a 2.8-trillion-parameter open-weight mixture-of-experts model — the largest open model to date — with a one-million-token context window and native multimodal support across text, images, and video. The model activates only 16...
-
47: Trillion-Parameter Agentic Intelligence
Ant Group's inclusionAI team has released "Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale", describing two open-weight model families — Ling 2.6 for low-latency generation and Ring 2.6 for deep agenti...
-
46: The Global Workspace
Anthropic's interpretability team has published "Verbalizable Representations Form a Global Workspace in Language Models", introducing the Jacobian lens (J-lens) — a new technique for reading what a language model is internally representing at any point during...
-
45: arXiv @ 35: The Quick Hack That Swallowed Science
Today, July 1, 2026, arXiv officially spun out from Cornell University to become an independent nonprofit. We cover the full arc — from Paul Ginsparg's NeXTstation in 1991, through the preprint revolution that broke academic publishing, through the AI explosio...
-
44: Qwen-AgentWorld: Language World Models for General Agents
A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning. In this work, we investigate how world modeling based on language models can further push the boundaries of...
-
42: The Board Has Been Terminated
On April 24, 2026, the White House fired all twenty-four members of the National Science Board by email — the independent governing body of the National Science Foundation, the agency that funded the public internet, the graphical web browser, 3D printing, the...
-
40: Rough Consensus and Running Scared
Between October 2025 and April 2026, cryptographer Daniel Bernstein published a seven-part blog series titled "NSA and IETF" alleging that intelligence agencies are using the IETF standards process to weaken the next generation of internet encryption. The disp...
-
39: Symbols Strike Back
A controlled experiment pits a neuro-symbolic system against a vision-language-action foundation model on the same robotic manipulation task, same robot, same simulation, same evaluation protocol — and the results are devastating for the foundation model. The ...
-
38: The Numbers Changed
Two papers published days apart have reduced the estimated physical qubit count needed to break widely deployed public-key cryptography by roughly two orders of magnitude — from around one million to as few as ten thousand. Together, they compress the timeline...
-
35: The Theorem Machine
Recent advances in foundational models have yielded reasoning systems capable of achieving a gold-medal standard at the International Mathematical Olympiad. We introduce Aletheia, a math research agent that iteratively generates, verifies, and revises solution...
-
34: Spinning to Zero
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate closes a gap that has been open since Claude Shannon defined the theoretical floor for lossy compression in 1948. For nearly eighty years, practical vector quantization methods fell expon...
-
32: The Green Gambit
Nvidia committed $26 billion over five years to building open-weight AI models. This episode examines the strategy: open weights as hardware lock-in, the Nemotron Coalition, NemoClaw agent runtime, the Vera Rubin and Feynman hardware roadmaps, and what it mean...
-
30: The Megatron Problem
Every competitive frontier model going forward is sparse — a Mixture-of-Experts architecture where each token activates only a fraction of the total parameters. That decoupling of parameter count from per-token compute sounds like a free lunch. The engineering...
-
29: In Lockstep
Every LLM-based text-to-speech system shipping today carries a structural flaw: text tokens and audio frames move at incompatible speeds inside the same model, forcing engineers to choose between reliability, quality, and inference cost. Hume AI's TADA: A Gene...
-
27: The Bitter Lesson
Rich Sutton published a 1200-word essay in 2019 arguing that 70 years of AI research proved one thing: general methods leveraging computation always beat human-curated knowledge in the long run. Most researchers disagreed. Then the last five years happened. No...
-
25: The Window
The economics of vulnerability discovery just broke. In twenty minutes, Claude Opus 4.6 found a novel use-after-free memory bug in Firefox — one of the most audited codebases on the internet, backed by millions of CPU hours of continuous fuzzing. That single r...
-
24: From Shadows to Worlds
Language models can quote the manual on a bicycle and still miss a broken chain. Beyond Language Modeling: An Exploration of Multimodal Pretraining argues that this is structural, not incidental: text is a lossy compression of reality, and models trained only ...
-
23: Saguaro: The Algorithm That Doesn't Wait
Speculative decoding already beats autoregressive generation — but it still has a sequential bottleneck: verification must finish before drafting restarts. Saguaro (Speculative Speculative Decoding) breaks that dependency by pre-speculating for likely verifica...
-
22: Qwen's Best Day Was Its Last
On the night Alibaba shipped Qwen3.5 — a 397-billion-parameter sparse mixture-of-experts model with 17B active parameters, a 1M-token context window, and a small-model family the open-source community had been waiting for — they fired the person who built it. ...
-
21: dLLM: Diffusion Gets a Framework
Every major language model in production today — GPT, Claude, Gemini, Llama — generates text the same way: left to right, one token at a time. That sequential assumption has been so productive for so long that most researchers treat it as fixed. A team at UC B...
-
20: DualPath: Breaking the Storage Wall
As AI agents run for hundreds of turns with ninety-five percent KV-cache hit rates, the bottleneck shifts from compute to storage I/O. DualPath from Peking University, Tsinghua, and DeepSeek exploits idle decode-engine storage NICs to load KV-cache via RDMA, a...
-
19: Agents of Chaos
Someone finally ran a proper pentest on autonomous AI agents. Natalie Shapira, David Bau, and thirty researchers deployed LLM agents with persistent memory, email, Discord, and shell access then spent two weeks red-teaming them. Eleven failure modes, every one...
-
18: The $20K Arms That Changed Robotics
The most important robotics breakthrough of the last three years was not a new algorithm or a bigger model. It was making the hardware cheap enough to collect enough data. We trace the ALOHA lineage from a twenty thousand dollar bimanual teleoperation rig in a...
-
17: The Math That Proves You're Human
World ID's proof-of-personhood system went from a centralized iris database to a quantum-secure, open-source cryptographic protocol where no single entity holds biometric data. We walk through the Daugman iris code, Shamir Secret Sharing, Secure Multi-Party Co...
-
16: H-Neurons: The Neurons That Make AI Lie
A team at Tsinghua University claims to have identified the specific neurons that predict when a large language model is about to hallucinate. Less than 0.1% of MLP neurons, identified via sparse logistic regression, generalize across domains and even detect f...
-
15: The Age Reversal Trial: Sinclair, Hype, and the Eye of the Storm
The FDA has cleared the first-ever human trial of a therapy designed to partially reverse cellular aging. Life Biosciences' ER-100, an epigenetic reprogramming treatment using a subset of Yamanaka factors delivered via AAV vector, will be injected into the eye...
-
14: Writing Data in Glass — Microsoft Project Silica and the 10,000-Year Storage Problem
Microsoft Research published a complete system for writing data into borosilicate glass using femtosecond lasers. A palm-sized square holds nearly 5TB and survives for over 10,000 years. This episode traces the 30-year journey from Eric Mazur to Project Silica...
-
13: Fast KV Compaction via Attention Matching
MIT researchers propose compressing LLM context in latent space rather than token space. Using closed-form linear algebra instead of gradient descent, Attention Matching achieves 50x KV cache compression in seconds — dramatically outperforming summarization on...
-
12: Kolmogorov Complexity — Sunday Greatest Hits
The only full textbook on Ilya Sutskever's famous reading list. Why did a deep learning pioneer tell John Carmack to study algorithmic randomness? Because compression is intelligence — and this book is the mathematical foundation for that claim. We cover Kolmo...
-
11: DreamZero — World Action Models are Zero-shot Policies
NVIDIA introduces DreamZero, a 14-billion parameter World Action Model that jointly predicts future video and robot actions from a video diffusion backbone. Unlike Vision-Language-Action models that fail on physically novel tasks, DreamZero achieves over 2x im...
-
10: DeepMind Dispatch #1: From Autonomous Mathematicians to AI Musicians
Our first DeepMind Dispatch covers three papers: Aletheia — a system that generates and verifies mathematical proofs autonomously; advances in Hutter optimization for large-scale model training; and Lyria 3, DeepMind's latest music generation model. We break d...
-
9: BitDance: Scaling Autoregressive Generative Models with Binary Tokens
We present BitDance, a scalable autoregressive (AR) image generator that predicts binary visual tokens instead of codebook indices. With high-entropy binary latents, BitDance lets each token represent up to 2^256 states, yielding a compact yet highly expressiv...
-
8: SkillRL: Don't Give Agents Memories, Give Them Skills
SkillRL from UNC Chapel Hill achieves 89.9% on ALFWorld with a 7B model — beating GPT-4o by 41.9 points. The secret: distilling raw experience into compact, reusable skills instead of storing verbose trajectory memories.
-
7: ΔBelief-RL: Rethinking How AI Learns to Act
We explore a bold new framework that rethinks reinforcement learning from the ground up — replacing reward maximization with belief updating, and asking whether AI agents should learn the way scientists do.
-
6: Building a Robot Mind in the Open
Alibaba DAMO Academy built a complete embodied AI system in six months — eyes, hands, imagination, unified brain — and open-sourced everything. Seven model checkpoints, Apache 2.0, zero gating. This is the story of RynnBrain.
-
5: From Blood Sacrifice to Universal Translator
In July 2024, a French nonprofit's open-source voice AI went viral for demanding human sacrifice mid-conversation. Seven months later, the same team used the same architecture to build a real-time speech translator that runs on your phone. This is the story of...
-
4: The Week China Open-Sourced The Frontier
In a 48-hour span, three Chinese AI labs independently released frontier-class open-weight models. Step 3.5 Flash from StepFun delivers frontier intelligence with just 11 billion active parameters. MiniMax M2.5 offers comparable performance at one-twentieth th...
-
3: DreamDojo — Teaching Robots to Dream
Researchers from UC Berkeley, NVIDIA, and UT Austin introduce DreamDojo, a framework that teaches robots physical skills by learning from large-scale human videos. Instead of expensive robot-specific data, DreamDojo distills 5 years of human video into a gener...
-
2: Generative Modeling via Drifting — One-Step Image Generation
Researchers from MIT and Harvard propose Drifting Models, a new paradigm for generative modeling that achieves state-of-the-art image generation in a single forward pass. Instead of iterating at inference time like diffusion models, Drifting Models evolve the ...
-
1: Attention Is All You Need — The Paper That Changed Everything
In our inaugural episode, we dive deep into Attention Is All You Need — the 15-page paper from June 2017 that introduced the Transformer architecture and reshaped all of artificial intelligence. We break down how it works, why the title is a Beatles joke, and ...