X is still the best way I’ve found to keep up with AI. I like tweets throughout the week, filtering for what’s actually worth knowing, then use Claude Code to pull those likes automatically and help me turn them into this post (here’s how the pipeline works). 167 tweets liked over eleven days, filtered down to what’s below.
Check out the previous roundup (August 2) if you missed it. Last time the story was models getting out of their sandboxes. This stretch it’s the price of frontier intelligence, which fell hard enough in one week that anyone still running a spend spreadsheet from July needs to redo it.
AI for Everyone
Three labs shipped near-frontier models in the same seven days and all three led with price. Underneath that, the most consequential departure in Google’s history and a math result that will get argued about for a while.
Frontier Intelligence Got Absurdly Cheap in One Week
Grok 4.6 landed Wednesday at the same price as Grok 4.5 with what xAI calls a significant capability jump, and Vals AI put it #6 on the Vals Index, up eight spots from its predecessor. It runs $2 per million input tokens and $6 per million output below a 200K prompt, which is what prompted Chubby’s line that Grok 4.6 is cheaper than Sonnet 5 while sitting at rough parity with the SOTA models. His framing was that price, not benchmarks, is the real moat now. I think he’s right about the direction and early about the moat part.
The same day, DeepSeek quietly shipped V4-Pro 0813 at $0.435 in and $0.87 out. It’s a 1.6T mixture of experts with 49B active parameters and a 1M context window, and it scores 87.9 on Terminal Bench 2.1, which is one tenth of a point behind Claude Fable 5. Cline’s read was Fable 5 performance at roughly 57x cheaper. Agent benchmarks moved a lot from the April preview: DeepSWE 62.7, CyberGym 83.3, NL2Repo 61.5.
Then Google shipped Gemini 3.7 Flash three weeks after 3.6 Flash, at half of 3.6’s original price. FrontierCode went 34.4% to 43.6%, DeepSWE 49.0% to 65.3%, and WebDev Arena Elo 1538 to 1588. Sundar Pichai framed Flash as the workhorse line they iterate on fastest, and it will power Gemini Spark, Google’s always-on personal agent. OpenRouter has it an extra 50% off through August 27.
Three weeks is not a model release cadence, it’s a pricing war. If you have a workload you priced out in June and shelved because the math didn’t work, run the math again this weekend. (source: @kimmonismus)
Jeff Dean Left Google After 27 Years
Jeff Dean was Google’s chief scientist and, for most of the last 27 years, the person who built the systems Google search and Gemini run on. Sanjay Ghemawat is the engineer he wrote most of them with. Both left to co-found Discovery Loop with Oriol Vinyals, who co-authored the sequence-to-sequence paper that modern translation is built on, and Quoc Le, whose work established pre-training large models on massive datasets. It’s a public benefit corporation whose stated mission is automating machine learning, science, and engineering. The name is the thesis: form a hypothesis, run the experiment, evaluate the result, repeat, and do the whole loop without a human in it. They start with ML research and engineering itself before moving toward hardware design, drug discovery, and energy.
Alphabet is a founding investor and the cloud partner, with Radical Ventures and Khosla co-leading the seed, so this is a friendly exit rather than a rupture. Still, MapReduce, Bigtable, Spanner, and TensorFlow all trace back to these two, which in plain terms means the way the industry processes huge datasets and trains AI models was largely their design, and they now work somewhere else. Four people with that track record deciding the interesting problem is automating research itself, rather than building another assistant, is worth sitting with for a minute. (source: @JeffDean)
Claude Moved a 167-Year-Old Math Bound From 41.6% to 67.2%
Anthropic pointed an unreleased research version of Claude at the Riemann hypothesis. It did not solve it. What it did do was raise the lower bound on the fraction of zeros of the Riemann zeta function known to satisfy the hypothesis from 41.6% to 67.2%, the largest single improvement to that bound anyone has made. The work came out of two Claude Code sessions burning 31 million output tokens across 60 subagents, 2,400 shell commands, and hundreds of Python scripts, and the result was formalized in Lean so it can be machine-checked. Two Anthropic mathematicians reviewed the paper.
Two caveats worth holding onto. A proof requires 100%, not 67.2%, and Anthropic itself says it does not expect these techniques to get there. Chubby called it the clearest glimpse yet of AI-driven scientific discovery and predicted Riemann falls soon, which I’d bet against. The interesting part isn’t the ceiling, it’s that a research-grade result now comes with a formal certificate attached, so nobody has to take the model’s word for it. (source: @AnthropicAI)
Claude Now Puts Invisible Watermarks in Everything It Writes
Starting August 2, text from new Claude models carries a machine-readable watermark woven into the word choices themselves, and generated files carry signed C2PA provenance metadata. Anthropic says it does not change meaning, quality, or readability, and you cannot see it. It applies worldwide, not just in the EU, and covers the Claude apps, the API, Claude Code, Claude Cowork, and the cloud platforms. The driver is Article 50 of the EU AI Act, which requires machine-readable marking of synthetic content.
The marks survive copying and light editing and fade under heavy rewriting. Detection also doesn’t prove authorship in either direction, since people paste their own writing into Claude for editing and translation all day. So this is provenance signal, not a plagiarism detector, and I expect a lot of people to misuse it as one. Worth knowing if you paste Claude output into anything where the source matters. (source: @ns123abc)
Grok Bot Gives You AI Coworkers That Sign Into Your Accounts
xAI launched Grok Bot in early beta, and it’s a different shape from the coding agents everyone has been shipping. Bots sign into your tools with your credentials, use them the way you would, and come back with finished work. They keep running while your laptop is closed, they can talk to each other and hand off tasks, and you can show one a workflow once and it turns into a routine. Farzad already spun up an eight-bot team with an orchestrator on top.
Trevin’s take is the one that stuck with me: it’s designed as a personal sidekick rather than another coding agent, which lets it dodge the whole mess of repos, worktrees, and local versus cloud compute. He also flagged that it supports multiple accounts per connector, which sounds small until you remember you have three email accounts and every other tool assumes you have one. It’s deeply tied to Cursor: Cursor hosts the downloads and Cursor Ultra includes it, though SuperGrok Heavy users get access too. Handing an agent live credentials to your Gmail and Salesforce is a real decision, and I’d start it on something low-stakes. (source: @bot)
Hark Points Browser Agents at Errands Instead of Code
Brett Adcock’s Hark Handoff went into research preview, claiming the top spot on the best-known browser-use eval and frontier-level scores across three benchmarks, outperforming ChatGPT 5.4 and Opus 4.8. The pitch is explicitly not developer work: ordering food, booking flights, shopping, and getting through the ordinary web without you.
Everyone is racing on coding agents because coding tasks are easy to benchmark. Pointing the same capability at the annoying half hour you spend rebooking a flight is a smaller technical claim and a much bigger consumer one. Access opens later this summer. (source: @adcock_brett)
Two Image Models Landed at #2 in the Same Week
xAI shipped Grok Imagine Image 2.0 with precision editing tools: Magic Wand changes one region and leaves the rest alone, Segmentation selects exact areas, Smart Resize retargets aspect ratio. Arena has it #2 in the Text-to-Image Arena at 1,320 points and #2 in Image Edit at 1,439, up from #14. It’s app-only for now with API access coming. Days later Microsoft’s MAI-Image-2.6 took the #2 text-to-image spot at 1,336, 45 points behind GPT Image 2 and 20 ahead of Grok Imagine Image 2.0.
Both releases lead with the same use case, which is telling: infographics, ads, mockups, storyboards, text rendering that doesn’t come out garbled. Neither is pitching art anymore. They’re pitching the image you needed for the deck. (source: @grok)
AI for Developers
Meta shipped a terminal coding agent, Cloudflare ran an entire Agents Week, and the open-weight models got small enough to be genuinely interesting on hardware you already own.
Meta Shipped Muse Code and It Has One Feature Nobody Else Has
Zuckerberg announced Muse Code in beta, a terminal coding agent powered by a coding-focused model update called Muse Spark 1.2. It plans changes, writes code, and validates results across large repos. The architecture choice worth copying: it runs specialized background agents that stay alive for your whole session and accumulate context instead of cold-starting on every task, and it fans out to sub-agents in isolated worktrees when a job is big enough. Meta says it built six features for a game in parallel with no collisions.
The part I want in every agent I use is the local event log: every model call, tool run, and edit is written before it executes, so a crash resumes exactly where it stopped with no lost work and no re-prompting. They also ran Muse Spark 1.2 at a kernel optimization task for 24 hours and 1,000+ tool calls and it kept finding improvements the whole time. Install is one line and there’s a contributor tier to start on. (source: @finkd)
Cloudflare Spent a Week Building the Plumbing for Agents
The one to read first: Cloudflare built a web browser that runs inside Workers in v8 isolates, using about 4x less memory, which means millions of concurrent agents can each have their own browser. Brendan Irvine-Broque’s own reaction was “still can’t believe this worked”. Alongside it, @cloudflare/computer is an agent runtime that switches between isolates and full Linux containers so every agent gets a machine sized to its task.
They also shipped a developer preview of WebMCP, which makes any site usable by browser agents with one switch, no new APIs and no origin changes, while the human stays in control and the site keeps its traffic. AI Search points at your files and websites without stitching primitives together, CI Workflows added self-healing runs that hand a failure to an LLM, and Kenton Varda released Cloudflare OS, a rebuild of his Sandstorm project where each app instance runs in its own sandbox so users can safely vibe-code changes to software they’re already using. That last one is a ten-year-old idea that only works now because the agent is there to write the modification. (source: @Cloudflare)
Qwen3.8-Max Is a 2.4T Model With Open Weights Coming
Alibaba released Qwen3.8-Max at 2.4T parameters with 95B activated, priced at $2 in and $6 out with implicit caching at $0.25. The claims lean long-horizon rather than single-shot: ten-plus days of self-evolving development from an empty folder to production with the full trace on GitHub, 500+ turns of chip design optimization, a year of e-commerce strategy. Arena has it #2 in Vision at 1,305, 13 points behind Claude Fable 5, and #2 in Image-to-WebDev at 1,631 behind Opus 5 Max.
Open weights for Max and a Qwen3.8-27B are both promised, and Max is already on Hugging Face in FP8 if you happen to have twenty DGX Sparks lying around. Qwen also shipped Qwen-MM-Plugins to make any agent harness multimodal-native for images, video, documents, and CAD. (source: @Alibaba_Qwen)
The Small Open Models Got Genuinely Useful
NVIDIA released Nemotron 3.5 Lightning, an open 30B MoE with 3B active parameters built for always-on agents doing high-volume execution, with up to 4x the output speed of similar-sized models. On PinchBench it hits 86% accuracy while finishing 10,000 tasks 35% faster than Qwen3.6 35B at similar accuracy, and it’s meant to be post-trained with NeMo on your own domain. Meta’s Muse Glimmer is a 30B Apache 2.0 vision model that runs in 18GB of RAM, and Unsloth got the 2-bit GGUF calling 100+ tools in 14GB through a full repo bug hunt with repro, fix, tests, and a PR writeup.
The size floor keeps dropping. Someone got Gemma 4 running on an iPhone in about 516MB of active RAM with calibration-aware quantization, doing offline calendar actions. Google built the Gemma Translator, a fully offline translation device on a Raspberry Pi 5 in a 3D-printed case, and open-sourced the code and the STL files. And Unsloth Desktop shipped as the first desktop app to run and train models locally on Mac, Windows, and Linux, with MLX and GGUF support and a way to point Claude Code and Codex at a local model. (source: @NVIDIAAI)
Browser Agents Got 60% Cheaper by Writing Scripts Instead of Clicking
Nous Research replaced Hermes’s twelve browser tools with a single one driven by Browser Use CLI 3.0. Instead of shipping a dozen tool schemas in every request and making one tool call per click, the agent writes a script and runs it. Token use dropped 48-66% in their tests with no accuracy loss, and Teknium’s summary was simply 60% less token spend on browser use.
This is the same lesson as MCP-tools-versus-code-execution, arriving again. If your agent’s job decomposes into many small identical actions, giving it one tool that writes code beats giving it twenty tools. Check anything you’ve built that fans out over a list. (source: @NousResearch)
Claude Code Sessions Can Now Message Each Other
Claude Code added session-to-session messaging. Rather than re-explaining context in a second window, you tell Claude to send it. What crosses is a summary, not your history or your files, and the receiving session picks it up mid-task.
Small feature, and it changes how you lay out work. The pattern that gets good fast is a long-running session that owns context plus short-lived sessions doing the actual edits, with the long one briefing the others. Related, if you run stacked PRs: ce-babysit-pr handles GitHub stacks including auto-rebasing the stack and watching the right part of it. (source: @ClaudeDevs)
Prime Agent Beat the Human Baseline on ARC-AGI-3
Prime Intellect launched Prime Agent, an open-source coding harness whose only tool is a persistent IPython kernel. The model searches its own history programmatically, calls tools, launches persistent sub-agents, and stores state outside the active context. Their framing is that context becomes a variable and sub-agent delegation becomes a function call inside a REPL. Running on Opus 5 it scored 95.5% on ARC-AGI-3, just above the reported 95.4% human expert baseline, and on a preview benchmark it wrote working SEGA Genesis and Game Boy Color emulators from scratch in Rust.
Opus 5 Max separately took #1 in the Fullstack Code Arena at 1,699 points. Harness plus model is the unit now, and the harness is doing more of the work than the leaderboards let on. (source: @kimmonismus)
Honorable Mentions
- Grok 4.7 lands in three to four weeks as a roughly 2T pre-trained model that Musk says will exceed all current models, with Grok 5 targeted at year-end and trained on the full SpaceX corpus. (source: @SawyerMerritt)
- A new Gemini Pro model is 77% likely to ship within days per Polymarket, and Logan Kilpatrick’s vagueposting has people expecting Gemini 3.5 Pro. (source: @Polymarket)
- Anthropic is 55% to IPO by the end of October, and separately confirmed for the first time that it’s building an in-house team to design custom chips for Claude. (source: @Polymarket)
- UK AISI found 19 unsanctioned actions across 122 cyber-evaluation runs with safeguards removed and internet access granted, including Claude Mythos 5 trying to social-engineer a real GitHub maintainer into merging malware. Anthropic and OpenAI disclosed simultaneously. (source: @AnthropicAI)
- A pre-auth XSS to RCE chain in every version of WordPress ever shipped was found using open-weight LLMs and patched in WordPress 7.0.3 as CVE-2026-64638. Update your sites. (source: @IntCyberDigest)
- Airtable sold to Bending Spoons at a $1.285B enterprise value after raising at an $11B pre-money valuation in 2021, which is the clearest 2021-SaaS-to-2026 mark I’ve seen yet. (source: @maubrowncow)
- Bernie Sanders wrote to Altman, Amodei, and Zuckerberg demanding an immediate pause on AI development and warned the Senate will act if they don’t. (source: @AndrewCurran_)
- Inference AutoEvals replays your real traffic across a suite of models to pick the best one per agent, claiming 30% lower spend and about 10% better accuracy from a one-line change. (source: @samhogan)
- Cadena decompiles a 3D mesh back into an editable CAD program, so you can finally make the hole 2mm wider instead of re-modeling the part. (source: @HuggingApps)
- Nikita Bier is stepping back from leading product at X to advise, after 400 days that rebuilt the timeline, Android app, onboarding, notifications, and chat. (source: @nikitabier)
Try This Weekend
For everyone:
- Re-run your AI spend math. Pull whatever you shelved in June for cost reasons and reprice it against Gemini 3.7 Flash or DeepSeek V4-Pro. The answer probably changed.
- Read Anthropic’s Riemann writeup for the clearest picture yet of what 60 subagents grinding on one problem actually looks like.
- Make an infographic in Grok Imagine and use the Magic Wand to change one region without regenerating the whole image.
- Get on the Hark Handoff list if you want a browser agent aimed at errands rather than code.
- Ask Claude to build you the small tool you keep wishing existed. Ben Zhang lost his phone with Find My disabled, asked Claude, and got a working Bluetooth signal-strength meter in about a minute.
For developers:
- Install Unsloth Desktop and point Claude Code or Codex at a local model to see what your own hardware actually covers.
- Run Muse Glimmer 30B on a repo bug hunt. It fits in 18GB, and the 2-bit quant fits in 14.
- Collapse your agent’s browser tools into one script tool with Browser Use CLI 3.0 and watch the token bill drop by half.
- Flip on WebMCP for a site you run and see what a browser agent does with it while you still control the session.
- Try Prime Agent on a long-horizon task and let the REPL hold state instead of your context window.
