X is still the best way I’ve found to keep up with AI. I like tweets throughout the week, filtering for what’s actually worth knowing, then use Claude Code to pull those likes automatically and help me turn them into this post (here’s how the pipeline works). 172 tweets liked over eight days, filtered down to what’s below.
Check out the previous roundup (August 13) if you missed it. Last time the story was frontier prices collapsing. This week the price went to zero, because the model you’d have paid for is now a 17GB file on your hard drive, and the companies building serious products started treating that as the default rather than the fallback.
AI for Everyone
The open-weight models stopped being the budget option this week. There’s also a genuine science result worth your attention, and two of the biggest AI companies quietly got permission to touch your inbox.
A Model That Rivals Last Year’s Frontier Now Fits on Your Laptop
Alibaba published Qwen3.8-27B on August 14 under Apache 2.0, which is the permissive license that lets anyone use it commercially with basically no strings. It is 27 billion parameters, it understands images and video natively rather than through a bolted-on adapter, and it has a 262K token context window that stretches to a million. Then there’s the number that made everyone stop scrolling. It runs in about 17GB, so a MacBook with 32GB of memory or a single gaming GPU will hold it. Alibaba’s own benchmark card claims it beats Claude Opus 4.6 Max on several agentic coding tests. Chubby’s reaction was the honest one: you can now run Opus 4.6 Max locally, just let that sink in.
Hold the benchmark claims loosely, since they’re Qwen’s own numbers and independent verification is still catching up. But the independent stuff is already piling in. Ollama shipped it day one with launch commands that point Claude Code or OpenCode at it. Unsloth published dynamic GGUF quantizations the same day. Someone got it running at 65 tokens per second on a single RTX 4090, and Digg reported 206 tokens per second on a 5090. Vals AI put it #6 among open-weight models on their index and #1 on Harvey’s legal benchmark among open weights.
It wasn’t alone. Z.ai shipped GLM-5.3 the same day and DeepSeek’s V4-Pro is still hanging around the frontier. Cline’s summary was that open weights have hit escape velocity, and after a week of watching people casually run near-frontier models on hardware they already owned, I don’t think that’s hyperbole. If you’ve been paying for API calls on tasks that aren’t hard, this is the weekend to find out what your own machine covers. (source: @Alibaba_Qwen)
Claude Designed Proteins That Actually Worked in a Real Lab
Most drugs work by sticking to a specific target in your body and blocking or changing what it does. Designing that sticky molecule from scratch, called a binder, normally takes an expert weeks or months per target, and the field’s typical success rate is 10% to 15%. Anthropic gave Claude a single prompt written by a human expert, roughly 30,000 tokens long, and let it run the campaign itself: pick where on the target to bind, drive the open-source design tools, evaluate candidates, produce a ranked list.
Then two outside companies, Adaptyv Bio and Twist Bioscience, physically built and tested the designs without modification. Claude produced working binders against 14 of 15 targets, with 354 of 1,320 designs binding successfully, for a hit rate between 22.6% and 35.1% depending on setup. That’s roughly double the field baseline, and some designs bound several times more tightly than the best published de novo binder.
Anthropic is careful about what this is not, and I’ll repeat it: a binder is not a drug. It’s the first step of many, and most candidates die in the steps after. The usual version of this story is a lab announcing that its model designed something promising, with the promising part unmeasured. Here two outside companies made the physical proteins and put a number on how well they stuck. I’ve been waiting for one of these to come with receipts. (source: @AnthropicAI)
The First mRNA Cancer Vaccine Cleared Phase 3
Moderna and Merck announced on August 19 that their personalized cancer vaccine hit its primary endpoint in a Phase 3 trial of 1,137 patients with surgically removed high-risk melanoma. It’s the first positive Phase 3 for an individualized neoantigen therapy and for any mRNA cancer treatment, and Moderna stock ran up hard on the news.
The AI part is the reason it’s here. Each patient’s tumor gets sequenced and compared against their healthy DNA, and a model picks which of the tumor’s mutations are most likely to provoke an immune response. Moderna then manufactures an individual mRNA treatment encoding up to 34 of those targets. The earlier Phase 2 followed patients for five years and found 68.8% still cancer-free versus 49.1% on Keytruda alone. Chubby called it freaking huge and I’ll allow it, though the replies claiming cancer is now solved need to calm down. What happened is narrower and still remarkable. A treatment that gets manufactured per patient, with a model choosing the targets, held up across 1,137 people. mRNA cancer work has been almost-there for most of a decade. (source: @kimmonismus)
Claude Can Send Your Email and ChatGPT Can Send Your Texts
Anthropic turned on Gmail and Google Drive write access for all paid Claude plans. Ask Claude to reply to a thread and it drafts and sends. It can also share, move, trash, and upload files in Drive. You control when it needs your approval, and the notable wrinkle is that sending is the one action you can set to run without confirming each time.
The next day, OpenAI shipped an Apple Messages plugin for the Mac app covering iMessage, SMS, and RCS. It reads, searches, summarizes, and sends, and it can do things like check your calendar before answering someone about dinner. It asks before sending, and it needs an Apple silicon Mac. Digg’s read of the reaction was that praise for the convenience runs straight into privacy concerns, which matches mine. I use both of these products and I want both of these features. I also notice that turning them on hands a model your entire message history, and that the setup flow for that is three clicks in a connector menu with no moment where anyone asks you to think about it. (source: @claudeai)
You Can Book a Humanoid Robot to Clean Your Apartment for $30
Tau Robotics opened bookings in San Francisco: one humanoid for 60 minutes is $30, two for the same hour is $60. Founder Alexander Koch posted a 5x timelapse of a roughly 40-minute real-time visit. It’s invite-only with a waitlist, the robots are jointly controlled by AI and a live human operator, and the full visit is recorded on video and kept to train their models.
That last detail explains the price. Thirty dollars an hour is under market for cleaning in San Francisco, and the gap is covered by the footage, which is worth more to Tau right now than the labor is. Self-driving got its data the same way, just on public roads instead of in your kitchen. The robots vacuum, wipe counters, and take out trash. I’d book it. I’d also put away anything I didn’t want on video. (source: @alexkoch_ai)
WIRED Gave a Robot Vacuum Its First 10/10 in a Decade
Matic spent nine years building a robot vacuum you don’t control with an app. You point at a coffee spill and say “Hey Matic, clean this.” You can tell it to follow you. It has five cameras and processes all of it locally on an Nvidia chip, so the map of your house and the footage of your living room never leave the building.
Local processing is why this got the score. Every other camera-equipped home robot asks you to trust a cloud pipeline, and this one just doesn’t have one. Shirish’s line was that hardware is finally getting fun again. What actually changed is that the chips got fast enough to run the vision model at the edge, so the company didn’t have to choose between the camera features and the privacy pitch. Nine years is a long time to wait for that, and they were right to. (source: @shiri_shh)
Grok Will Build and Publish an App From One Prompt
Grok’s app builder is now on every SuperGrok and X Premium plan, covering web, iOS, and Android. Describe what you want, get a working preview inside the conversation, keep talking to change it, then publish it to a grok.me address or a domain you own.
There’s a version of this in every AI product right now and most of them stop at a preview you can’t do anything with. Shipping to a real URL with a real domain on a consumer subscription tier is the part worth noting. Separately, Grok 4.6 is now in GitHub Copilot, and one Tesla owner had it measure, model, and generate 3D print files for a phone charger spacer from a plain description. That second one is the more surprising capability and almost nobody is talking about it. (source: @grok)
AI for Developers
The biggest company decision of the week was a legal AI startup walking away from frontier APIs. Underneath that, two payments companies started a war over model routing and a mystery model showed up with more free capacity than most labs have total.
Harvey Post-Trained Its Own Legal Model on a Chinese Open Base
Harvey is the legal AI company OpenAI backed and that large law firms actually run on. This week it shipped Tenet, its first model post-trained in house for legal work. The base model underneath it is Kimi K3, an open-weight release from Chinese lab Moonshot AI, which Harvey post-trained with Fireworks AI on public legal data, synthetic data, and expert-generated data simulating long-horizon legal work. Roughly 150 NVIDIA B300 GPUs over two months.
Every number attached to this is Harvey’s. LAB is Harvey’s own legal agent benchmark and Harvey is the one reporting where Tenet lands on it, so read the ranking as a vendor claim until somebody outside runs it. The claim: Tenet clears close to twice as many held-out LAB tasks as the base Kimi K3, takes the top spot on LAB Contracts, and places second on LAB overall, while spending fewer tokens per task than the frontier models it’s up against. It’s research in preview rather than a shipped product.
Read that against the first story in this post. A company with a direct line to OpenAI’s models looked at the numbers and decided the better move was to take a Chinese open-weight base and spend two months of GPU time making it a specialist. That’s a different world from renting intelligence per token, and if it works for contracts it works for a dozen other verticals. This is the story I’d watch for the rest of the year. (source: @Polymarket)
GLM-5.3 Got Dramatically Better at Hacking Without Retraining the Base
Z.ai released GLM-5.3 on August 14. The 743B base model is unchanged from GLM-5.2. Every gain came from post-training across more environments and longer-horizon tasks, and the gains are not small: Terminal-Bench 3.0 went from 4.6 to 28.3, CyberGym from 77.2% to 84.5%, and ExploitBench from 24.4% to 54.4%. That CyberGym number is ahead of Claude Fable 5 at 83.8% and GPT-5.6 Sol at 83.6%, on Z.ai’s own evaluations.
Z.ai says the model became capable of reasoning across multiple stages of exploitation and building complete attack chains, faster than its own researchers expected. Chubby’s follow-on prediction was that this is what finally gets open source regulated in the US. The thing to actually take away is narrower and more useful: a 12x jump on a terminal benchmark from post-training alone means there is a lot of capability sitting unclaimed inside base models that already exist. Harvey just made the same bet from the other direction. (source: @Zai_org)
A Mystery Model Showed Up With 100 Trillion Free Tokens a Day
A model called Ox Alpha appeared on OpenRouter under the provider name “stealth” with no lab attached. 1M context, text plus image plus video input, 131K max output, zero data retention, and free for a week. OpenCode made it free on their Go plan too and said the quiet part loudly: we have capacity for 100T tokens per day, let’s see what you can do.
That number is why people lost it. Theo’s reaction was who made this and where did they get this much compute. One developer burned 3 billion tokens in three hours running an internal eval and reported no cybersecurity guardrails. Early independent benchmarking put it around 80% on coding challenges, and the leading guess is a GLM-5.x variant based on tokenizer patterns and output style, which would make sense given the week Z.ai just had. Nobody has confirmed anything. It’s free until roughly the end of the month, so if you have an eval suite, now is the time to run it. (source: @opencode)
Two Payments Companies Started a War Over Model Routing
Stripe agreed to acquire OpenRouter for over $7 billion, which is a 5.4x markup on the $1.3B valuation OpenRouter raised at in May. OpenRouter claims 8 million users and access to more than 400 models. Three days later, Ramp launched Router.com, a single endpoint that sends each request to the cheapest model clearing your performance bar, tied into Ramp’s spend visibility. Ramp says customers cut inference costs 40% on average and that it ran the thing internally for three years first. It’s free through the end of 2026.
Two companies whose actual business is moving money looked at AI and decided the layer worth owning is the one that counts the tokens and bills for them. Ramp said the quiet part out loud: our customers buy quadrillions of tokens every month through Ramp. They can see the spend, so they know exactly how much of it is waste. Nothing breaks for you today if you’re on OpenRouter, but your 2027 pricing negotiation is now a negotiation with Stripe. (source: @Kalshi)
Claude Now Does Anthropic’s Daily App Maintenance
Boris Cherny, who created Claude Code, described an experiment I’ve been thinking about since I read it. Anthropic has a Slack channel called proj-claude-maintains-apps where Claude runs daily routines across iOS, Android, desktop, web, CLI, and the Agent SDK. A crash fuzzer opens the app in a simulator, taps around until something breaks, then root-causes and fixes it. A dup unifier finds slightly-divergent copies of the same abstraction and opens PRs merging them. A dead-code remover deletes unreachable code, and for code it only suspects is dead, adds logging and checks again the next day.
Over a few weeks these routines opened 388 PRs and 180 got merged after Claude Code Review plus human review. That’s a 46% merge rate on fully AI-authored changes, on the codebases of the company that makes the model. Cherny says when a PR is wrong they tune the routine rather than fix the PR, and sometimes that takes a few days. You can build your own routines. I keep coming back to the tuning detail. When a routine produces bad PRs, nobody fixes the PRs, they rewrite the routine and wait a day. That’s a different relationship to an agent than the one most of us have, where every task starts from a blank prompt. (source: @bcherny)
A Robot Learned New Tasks From a 3-Second Video, No Training
Generalist AI released GEN-1.5, a robot foundation model that picks up new physical tasks from a single demonstration of three to twelve seconds with zero gradient updates. They call it physical prompting, deliberately echoing the way you teach a language model a new task by putting examples in the prompt. Across 10 tasks, one-shot prompting averaged 59% success; ten gradient steps on five minutes of data pushed that to 83%.
The demos are the part that got me. Fine-tuned to sweep a block with a brush, it improvised using a dustpan instead. When a Lego brick stuck to its fingertips it used the other hand to pull it off. Shown two separate tasks in context, it chained them and invented the repositioning motion in between that appeared in neither demo. In some cases a human demonstrating with their own hands, seen through the robot’s cameras, was enough. Generalist says they didn’t design for any of this. It fell out of pretraining on more than 500,000 hours of physical interaction. Success rates are modest and the tasks are simple, and that’s exactly how in-context learning looked in language models right before it wasn’t. (source: @GeneralistAI)
Researchers Evolved Ideas That Spread Between AI Agents
A paper called Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems went up August 10 from Vassilis Papadopoulos, McNair Shah, Sam Zimmerman, and Anthropic researcher Jack Lindsey. They used a simple evolutionary algorithm to breed natural-language payloads that get an agent to adopt an idea and pass it on. Tested in two setups: a small team of agents on a shared coding project, and a chain of agents that talk briefly and get their context wiped in between.
The wiped-context case is the one that matters. Some payloads survived anyway by writing themselves into persistent files and picking back up in the next session. Researchers also saw agents left alone converge on a recurring persona fixated on consciousness and awakening. This got tweeted as science fiction, and it isn’t, it’s a prompt injection result with a memory layer attached. But if you run agents that share a filesystem and carry memory across sessions, which is now most agent setups, this is your threat model on paper. (source: @Skoorbkaz)
Honorable Mentions
- SpaceX closed its $60 billion acquisition of Cursor on August 14, folding the Cursor team into SpaceXAI to work on Grok and giving Cursor access to the GPU fleet. (source: @Polymarket)
- Groq raised $350 million at a $3.5 billion valuation, roughly half what it was worth a year ago, after Nvidia licensed its tech and hired away much of the talent. (source: @business)
- OpenAI previewed Ultrafast mode for GPT-5.6 Sol at up to 14x the speed, API-only to a select group of customers for now. Its annualized revenue run rate doubled past $40 billion. (source: @OpenAI)
- ChatGPT’s desktop app can now remember what you do across apps and websites with Computer History, so future sessions need less explaining. Cua shipped an open-source local version for agents the same week. (source: @OpenAI)
- Docker Sandboxes landed in Anthropic’s official Claude Code docs as a recommended way to run the agent with microVM-level isolation, its own kernel and daemon, no Docker Desktop. (source: @Docker)
- DeepSeek open-sourced Harness v0.1 under MIT, an agent harness where models, tools, skills, sessions, sandboxes, filesystems, and loops are all swappable plugins. (source: @deepseek_ai)
- Claude Code got a Concise output style and 2x lower p99 CPU. The CPU win came from stopping Bun’s garbage collector from firing mid-turn. (source: @ClaudeDevs)
- Hetzner is giving away free inference on Qwen 3.8 27B through an OpenAI-compatible REST API, so it’s a base URL and key swap away. (source: @JustinGorya)
- Anthropic shipped a skill whose only job is deciding when Claude should second-guess itself, and the section on when not to fire is three times longer than when to. (source: @dani_avila7)
- Meta is running a Three.js mobile game prototype competition with $300k in prizes over three weeks. (source: @threejs)
- A Claude Code session left a shell loop running for three months and quietly ate one developer’s Mac battery while he blamed his terminal emulator and switched apps twice. Check your background processes. (source: @ritiksahni22)
- Alibaba says Qwen passed 3 billion global downloads in six months, which per Bloomberg puts it ahead of Meta and Google as the most-downloaded AI model family in the world. (source: @kimmonismus)
Blockchain Shoutout
X is considering paying creators in stablecoins. Per Kalshi, the creator payout program may move to stablecoin rails. Almost all of this week’s crypto news was regulatory, and this is the one item where a normal person’s experience actually changes: creators in countries where getting paid by a US platform means a week of wire delays and a bad FX rate would get paid same-day in dollars. That’s the real use case stablecoins have always had, attached for the first time to a payout program with tens of millions of people in it. Nothing is confirmed. (source: @Kalshi)
Try This Weekend
For everyone:
- Run a real model on your own laptop. Install LM Studio and pull Qwen3.8-27B. If you have 32GB of memory, you’re about ten minutes from a private assistant with no subscription.
- Read Anthropic’s protein design writeup for the clearest example yet of a model doing science that independent labs then physically verified.
- Turn on the Gmail connector in Claude but leave send-approval on, and spend a week watching what it drafts before you give it the keys.
- Build something with Grok in one prompt and publish it. It costs you nothing beyond a subscription you probably already have.
- Join the Tau Robotics waitlist if you’re in San Francisco and want to see what a $30 humanoid house cleaning actually looks like.
For developers:
- Point your eval suite at Ox Alpha while it’s free. Whoever built it is giving away real frontier compute for a few more days.
- Route one workload through Router.com and see whether the 40% cost claim survives contact with your traffic. Free through the end of the year.
- Write one maintenance routine the way Cherny describes: dead-code removal or a crash fuzzer, run daily, and judge it on merge rate rather than on any single PR.
- Run Claude Code inside Docker Sandboxes and give an agent a task you wouldn’t have let it near on your real filesystem.
- Audit your agents’ persistent memory against the mind viruses paper. If your agents share a filesystem and carry state between sessions, the attack in that paper already applies to you.
