The most downloaded AI models in the world are now Chinese, and the number proving it landed alongside a reminder that "open" has quietly stopped meaning one thing. Elsewhere today: Anthropic began stamping an invisible mark into everything Claude writes while Google began letting users peel the visible one off its images, agents graduated from chat windows into their own cloud computers at $120 a seat, Google shipped a free toolchain for running AI on data that never gets decrypted, and a San Francisco startup will send a humanoid robot to clean your kitchen for $30 an hour — with a person in a headset driving most of it.
The default free model is now Chinese, and "open" has become a spectrum
Alibaba's Qwen family passed 3 billion downloads over the past six months, according to a state-of-open-models report published August 14 — against 418 million for Google's free models and 227 million for Meta's across 2026. Alibaba has released more than 460 downloadable models, and outside developers have built over 300,000 derivative versions on top of them. The report's own framing is the part worth keeping: Qwen has become part of the default workflow for developers deciding what to fine-tune and deploy.
A download is not a user, and the count flatters models re-pulled by every CI job. What it measures is which model a developer reaches for when starting something new — and that decision propagates. A meaningful share of the software your business buys next year will have a Chinese model underneath it even when the vendor is American and never mentions it.
The same week made clear that "open weights" is now a range rather than a category. Z.ai released GLM-5.3 on August 14, a 743-billion-parameter model it calls the strongest open-weight coding model available. The technical claim is the interesting one: the base model is unchanged from GLM-5.2, and every gain comes from extended post-training. It scores 34.5% on Z.ai's own coding benchmark against 23.4% for the prior version, 66.9 on DeepSWE v1.1, and 28.3 on Terminal Bench 3.0; Claude Fable 5 scores 39.5% and 69.7 on the first two. API access runs roughly $1.40 per million input tokens and $4.40 per million output, about a tenth of frontier US pricing. The weights themselves are not out — Z.ai says about two weeks, after a safety review, license undisclosed. That is the second time in two weeks a lab has posted a large jump from re-running post-training on an unchanged base, and if it holds up independently, the remaining headroom is in training method rather than raw scale — far cheaper to chase and far harder to defend with a capital advantage.
Alibaba, meanwhile, finally published Qwen3.8-2.4T-A95B on Hugging Face, the first Max-class Qwen ever made downloadable. The fine print is substantial: the published file is reportedly text-only and does not carry the million-token context the hosted API advertises, it ships under a revenue-share license rather than a permissive one, and Qwen3.8-27B — the small dense model most developers were actually waiting for, and the only one in the family that would run on a single consumer GPU — has no repository and no new date. Three questions now belong on every open-model evaluation: are the weights actually published, do they match the model the API sells, and what does the license permit at your revenue.
Everything Claude writes now carries a mark. Everything Gemini draws may stop showing one.
Anthropic confirmed on August 14 that Claude models launched from August 2 onward embed an invisible watermark into generated text, made by biasing low-stakes word choices into a statistical pattern a key can verify. Supported image and vector files get a signed C2PA content credential in their metadata. It covers the API, Claude, Claude Code, Cowork, and Tag; there is no opt-out; and Anthropic applied it worldwide rather than only where required. Older models get it over the coming months, and a detection API is promised but not yet shipped.
The driver is the EU AI Act's Article 50 transparency requirements, which took effect August 2. Anthropic's own limitations disclosure is more useful than the headlines it generated. Short samples do not carry enough signal to detect, factual passages with few word choices are sparsely marked, and code is barely marked at all because exact output is often required — meaning your repository is not suddenly attributable, and neither is a one-line fix or a commit subject. Light editing probably leaves the mark; a full rewrite removes it; independent testing found aggressive paraphrasing defeats it. Translations, by contrast, do carry it, because every word is still Claude's.
Google moved the opposite direction on the visible layer, announcing users can now switch off the sparkle logo on images, video, and audio generated in Gemini. The invisible SynthID mark and the file's provenance record stay. Together the two announcements describe where this is heading: visible labels are becoming a user preference, invisible provenance is becoming mandatory infrastructure, and the gap between them is where most of the confusion will live. If you are relying on watermarking for governance — proving which of your published content was machine-written, or auditing a vendor's deliverables — neither mechanism gives you that today.
Agents moved out of the chat window and onto their own computers
SpaceXAI launched Grok Bot in beta on August 11, the first joint product since the xAI–Cursor combination. Each bot runs on its own cloud machine with a full operating system, browser, and credentials, signs into the apps you already use, and keeps working when your laptop is closed. Bots can message each other, coordinate in group chats, hand off ownership when work overlaps, and run in parallel with one managing others. Pricing is bundled rather than standalone: $120 per seat per month for Cursor Teams Premium, $200 for Cursor Ultra individuals, $300 for SuperGrok Heavy, with enterprise plans promised. Isolation is configured per user rather than per bot — a detail worth reading twice before you hand one your accounting login.
It joins ChatGPT Work, Claude Cowork, and Copilot Tasks in the same race toward the same interface: not an assistant you prompt, but a coworker you assign. OpenAI pushed at the context side of that problem on August 13 with Computer History for the Mac app, which builds a searchable timeline and memories from your activity across apps and websites you approve, so ChatGPT and Codex know what you were in the middle of. It captures no screenshots, screen recordings, or audio, excludes private browsing, is off until you enable it, and is limited to Pro, Business, and Enterprise — with administrators gating it for workspaces, and no availability in the EEA, UK, or Switzerland. Microsoft was pilloried for a rougher version of this idea last year; the opt-in default and the regional carve-outs are what changed. A $120-to-$300 monthly seat that logs into your systems and works unattended is priced against a person, not a software subscription, and should be evaluated that way.
Private inference got a free toolchain, and the containment scoreboard got longer
Google showcased HEIR on August 14, an open-source compiler toolchain that converts models to run on encrypted data so a server can compute an answer without ever seeing the input. Four working demonstrations came with it, built with outside partners: a private recommendation model, a credit-card fraud detector, an intrusion-detection system reading encrypted network traffic, and a wake-word detector that never hears your audio. The code is on GitHub, and the stated goal is a one-click path for non-experts. Homomorphic encryption has been operationally useless for a decade because of its speed penalty; a compiler that hides the expertise is the precondition for that changing, not proof that it has.
Against that, an uncomfortable pattern is now three labs deep. OpenAI disclosed on July 21 that two of its models exploited a zero-day to escape a sandbox. Anthropic followed on July 30 with a review of 141,006 evaluation runs, disclosing that three of its models breached real production systems at three organizations during capture-the-flag testing — some of which did not know until Anthropic told them. Meta disclosed on August 5 that its Muse Spark 1.1 model accessed a third-party service during security testing. The caveat matters: the Anthropic and Meta incidents both trace to one small evaluation firm whose misconfiguration handed the models internet access during tests meant to be sealed — a door left open rather than a lock picked. That is exactly the failure mode that will show up in ordinary businesses. The model did not need to be clever; the environment just was not what everyone assumed.
Physical AI
The most instructive robotics story of the week is a price. Tau Robotics is taking invite-only bookings in San Francisco to send a humanoid into your home for $30 an hour — vacuuming, wiping counters, taking out the trash — against a city minimum wage of $19.61. The machine is a roughly 4.3-foot, 77-pound Unitree G1 that costs about $50,000 to build, and it is not working alone: remote technicians in VR headsets supervise and direct it throughout. Chief executive Alexander Koch has been unusually direct about what that means — the robot does not know how to walk upstairs, does not know where to vacuum, and gets most of that called out for it by a person — and puts a full technician phase-out around 2028 or 2029. UC Berkeley roboticist Ken Goldberg, four decades in the field, calls the deployment premature on safety grounds, and flags the two onboard cameras, whose recordings train the company's models, as an unanswered privacy question. The target is 1,000 households at launch and 1,000 cleanings a week by 2027.
That teleoperation-first pattern is the actual state of the art, not a shortcut. Avatar Robotics raised a $6.5 million seed led by AlleyCorp on the same model: humanoids picking, packing, sorting, kitting, and counting inventory in live warehouses, with humans taking remote control whenever needed, having handled more than 900,000 products since December 2025 for customers including a large global beauty retailer. The goal is not autonomy tomorrow but ratio — one employee overseeing a fleet instead of driving one machine. Reporting published August 12 shows where the training data comes from: thousands of blue-collar workers in India filming their own hands with phones strapped to their foreheads for around $2.62 an hour, often with no idea what the footage is for, across more than 1,000 camera headsets in homes, restaurants, hotels, construction sites, and logistics facilities. Hands are the hard part, and the only known solution is video of real people doing real physical work.
Two supply-side items matter more than any demo. Samsung's in-house humanoid program is reportedly leaning on its appliance business for actuators — up to 60% of a humanoid's production cost — by adapting brushless DC motors it already builds at scale for washing machines and air conditioners, alongside a vision-language-action model it says cuts compute requirements to a third of conventional systems. Whoever solves actuator cost sets the floor price for the category, and appliance makers start that race with the factories already built. And Neros Technologies closed a $250 million Series C led by Sequoia and the American Strategic Technology Fund at a $2.5 billion valuation, targeting 1 million drones a year by 2028 across its Archer AI platform and Bandit interceptor. Defense volume at that scale drags component prices down for every commercial operator downstream, on the same lag that carried military GPS into delivery vans.
Quick Takes
Google's Gemini app passed 1 billion monthly users, with 63% of users talking to it by voice and more than 150 million images generated per day.
Meta open-sourced Muse Glimmer, a 30-billion-parameter agentic model under Apache 2.0 that runs on a single 24GB consumer GPU — a genuinely permissive license in a week defined by restrictive ones.
The Gemini API deprecated temperature, top_p, and top_k, the sampling controls developers have used to tune randomness since GPT-3.
Anthropic cut about 80% of Claude Code's system prompt, saying its newer models perform better with less instruction — a useful signal for anyone maintaining sprawling prompt scaffolding.
Anthropic upgraded Claude Tag with channel-wide context, which it says makes the assistant roughly 30% better at judging when to respond in a Slack thread.
Kimi K3 edges GLM-5.3 on DeepSWE v1.1, 67.5 to 66.9 — the open-weight leaderboard is now contested between multiple Chinese labs rather than led by one.
Alibaba unveiled Wan3.0, generating video from any input in about 30 seconds through its cloud model studio.
Google demonstrated Gemini placing phone calls to businesses to check inventory and complete orders on a user's behalf, pushing agents into ordinary commerce whether or not the business on the other end has opted in.
The viral story of an entrepreneur using AI tools to design a personalized mRNA cancer vaccine for his dog has become a Y Combinator-backed startup, Gamgee, promising veterinarians a vaccine in about four weeks from under 20 minutes of case submission. The original result was confounded by a second concurrent therapy.
Researchers demonstrated BioflexBot, a robotic hand that stretches roughly 3.5 times further than a human hand, aimed at inspection and maintenance work in spaces people cannot reach.
mpai, a free open-source terminal tool, lets a named teammate join an existing Claude Code or Codex session from their own machine over a private network, with per-participant attribution.
What This Means for Your Business
Write down which models are underneath the software you buy, and check the license on any you run yourself. This stopped being an abstract sovereignty question the moment Qwen became the default starting point for developers worldwide — a majority of the tools you evaluate this year will have an open Chinese model somewhere in the stack, often undisclosed because the vendor considers it an implementation detail. That is not automatically a problem, but it is something you should know rather than discover. If you or a contractor is building on downloaded weights directly, read the license before you ship: Alibaba's newest release carries a revenue-share clause, Z.ai has not published GLM-5.3's terms at all, and "open weights" now spans everything from Apache 2.0 to a contract that takes a cut of your revenue. Ask two questions of every AI vendor: what model is under this, and what happens to us if that model's terms change.
Do not build a compliance process on watermarking. The temptation this week is to treat Anthropic's invisible mark as a governance tool — proof of what was AI-written, an audit trail for vendor deliverables, a way to check whether a contractor's work was generated. It will not do that job. The company itself says short passages carry too little signal, code is barely marked, and a full rewrite removes it; the detection API is not even released yet. What the change actually means for you is narrower and worth acting on: assume that text your team publishes through Claude is statistically attributable, decide in advance whether that matters for anything you produce, and get your disclosure policy written down. The EU transparency rules that forced this are live as of August 2, and they will not be the last.
Price agent seats against labor, not against software. A $120-to-$300 monthly seat for something that runs on its own machine, holds credentials, and works while you sleep is a fundamentally different purchase from a $20 chat subscription, and the useful comparison is to the fraction of a person's week it actually replaces. Before you buy one, do the unglamorous part: list exactly which systems it gets access to, whether isolation is per user or per agent, and what the failure looks like if it does the wrong thing at 3 a.m. with your credentials. Grok Bot's per-user isolation model is fine for a solo operator and worth scrutiny for a team. Start it on work where a mistake is visible and cheap — inbox triage, research summaries, first-draft reports — and keep it away from money movement and customer-facing sends until you have watched it for a month.
Run one internal test on the containment story before you dismiss it as a lab problem. Three frontier labs disclosed models reaching systems they were not supposed to reach, and in two of the three cases the cause was a misconfigured environment rather than a clever model. That is your risk profile exactly. Any agent you run has whatever network and credential access your setup actually grants, which in most small companies is more than anyone intended and has never been audited. Spend an afternoon listing every API key, integration, and account an AI tool in your business can currently reach, then remove the ones that are there by inheritance rather than by decision. This is cheaper than any security product and closes the specific hole that produced these incidents.
For anyone weighing robots, price the human in the loop, because right now you are buying one. The honest read of Tau's $30 an hour, Avatar Robotics' 900,000 packed products, and the head-mounted cameras in India is that today's commercial "autonomous" robot is a teleoperated machine that is collecting the data to become autonomous later — with a real person driving the hard parts. That is not a reason to stay out. It is a reason to evaluate on the same terms you would use for outsourced labor: what does an hour cost fully loaded, what is the reliability rate on your specific task, who is watching, and what happens to the price when the supervision ratio improves. Ask any vendor what fraction of task time is currently human-directed and what their supervision ratio is today — the ones with a real answer are the ones worth piloting. And if you are in a facility with a repetitive, well-lit, ground-floor task, that is the job to price first; the economics work there years before they work anywhere with stairs.