Agents of Work
August 7, 2026 · Agents of Work

Agents of Work AI Daily Briefing — August 7, 2026

OpenAI rebuilt its entire consumer lineup in a single afternoon, giving free users unlimited text chats and paying users a dial that controls how hard the model thinks. Elsewhere: researchers published the first viruses designed end to end by an AI model and then built them in a lab, the company that started the AI price war announced it is raising prices, AMD bought a startup that etches model weights directly into silicon, Tesla and SpaceX confirmed the site of the largest chip plant ever attempted, and a Toyota spinout argued that the most useful factory robot is the one that doesn't bother with legs.

OpenAI resets what a free AI account is worth

OpenAI restructured every tier of ChatGPT on August 6. Plus and Pro subscribers moved to GPT-5.6 Sol, tuned for the ordinary work of questions, web research, planning, writing, and decisions, and gained a thinking slider that lets the user decide how much reasoning effort a given question deserves rather than leaving it to a router. Free and Go users moved to GPT-5.6 Luna as the default and — the genuinely new part — received unlimited text chats, with a Think button available when a question needs more. Separate limits remain on files, images, voice, and image generation. OpenAI says factual errors fell 62% for Luna and 68% for Sol compared with GPT-5.5-Instant. Sol shipped the same day; the free-tier changes roll out over the following week.

Two things are worth separating. The error-rate figures are OpenAI's own measurement against its own prior model, so treat them as a direction rather than a guarantee. Unlimited text chats on the free tier is not a benchmark claim — it is a business decision, and it removes the last practical reason a casual user ever hit a paywall. For anyone selling software whose value is "we put a chatbot on it," the floor just moved.

The company also published usage data drawn from its user base, and the headline finding is the one operators should sit with. At work, people are more than twice as likely to use ChatGPT to complete a task or produce something — writing, coding, analysis — than they are outside work, where the pattern is still mostly asking questions. Usage by people over 35 rose in nearly every country, with that group now accounting for roughly five percentage points more of total messages than a year ago. Growth was fastest in parts of Latin America, Oceania, and Africa, with Peru, Uruguay, and Costa Rica climbing most. The tool stopped being a search box and became a place where work gets finished, and the people doing it are no longer disproportionately young.

A shared plug for agents, backed by four rivals

OpenAI also published Agent Plugins, a version 1.0.0 specification developed with Amazon, Microsoft, Vercel, and Cursor. It defines a portable package format bundling Model Context Protocol servers and reusable agent skills into one unit any compatible client can discover and load — and deliberately stops there, leaving distribution, installation, permissions, and interface behavior to each client. The narrowness is the point. Standards that try to specify everything get ignored; this one specifies only the packaging. For a small business, the practical consequence is that a connector built for one assistant has a plausible path to working in another — the difference between adopting an AI tool and marrying a vendor.

An AI wrote a virus, and sixteen of them worked

Researchers at Stanford and the Arc Institute published a paper in *Science* on August 6 describing the first complete viral genomes designed by a generative model and then physically built and tested. Using the Evo 1 and Evo 2 genome language models, trained on roughly 2.7 million genomes, the team generated hundreds of thousands of candidate bacteriophage genomes, chemically synthesized 302 of them as real DNA, and introduced them into *E. coli*. Sixteen produced viable phages that infected and killed bacteria. Collaborators included NYU Langone Health, Oxford, Columbia, and the Johns Hopkins Center for Health Security.

The reference point was phiX174, a bacteriophage of roughly 5,000 genetic letters and the first DNA genome ever sequenced, in 1977. Some AI designs shared more than 40% of its sequence; others were largely novel. In combination, cocktails of the AI-designed phages cleared strains of *E. coli* that natural phiX174 could not — a real result with real medical value, given that phage therapy is one of the few live options against antibiotic-resistant infection.

The safeguards are where this gets uncomfortable. The training data deliberately excluded viruses that infect humans, animals, plants, and fungi, and every one of the sixteen attacks bacteria only. The researchers state plainly that this filter can be undone by fine-tuning on the excluded sequences. Arc's Brian Hie framed the deeper problem: a genome no database has ever seen slips past screening systems built to match known threats. DNA synthesis screening is the industry's main biosecurity control, and it works by comparison against a catalogue. Generative design is a machine for producing things that are not in the catalogue.

The company that started the price war is raising prices

DeepSeek posted a notice on August 6 warning of a significant increase across its API pricing. It named no figures and published no new schedule, saying only that the rise would arrive soon and that users should plan accordingly. Its V4 Flash model currently runs $0.14 per million input units and $0.28 per million output — an order of magnitude below Western frontier pricing, and the specific fact that made DeepSeek a strategic threat rather than a curiosity.

It is not alone. Alibaba is reportedly preparing to charge heavy users of its next open-source model — a quieter version of the same admission. Meta is running the inverse experiment with its Muse Code contributor tier, where the discount is paid in training data rather than dollars. Three different answers to one question: someone has to pay for the compute, and none of the answers is "it keeps getting cheaper forever."

Which is why the hardware news lands where it does. AMD agreed to acquire Toronto-based Taalas, expected to close in the fourth quarter subject to regulatory approval. Taalas takes a fixed model and hard-wires its weights into a chip's metal layers during manufacturing, merging storage and compute on one die and eliminating the high-bandwidth memory stacks, advanced packaging, and liquid cooling that general-purpose accelerators require. Its first product, built around Meta's Llama 3.1 8B, claims roughly 17,000 tokens per second per user — close to ten times faster than alternatives — at a twentieth the build cost and a tenth the power. Those are vendor numbers, not independent measurements, and the chip leans on aggressive 3-bit quantization. Taalas raised $219 million and built that first chip with 24 engineers on $30 million. Anthropic separately confirmed an in-house chip design team for Claude, with Samsung a candidate manufacturing partner. Everyone is now attacking the same line item.

Enterprise agents get a price tag and a job description

Anthropic announced a partnership with hedge fund Millennium to build a sandboxed, auditable digital risk analyst deployed across 340 investment teams. The structural detail matters more than the client name: sandboxed and auditable are the two properties that let a regulated firm put a model near real money, and they are the properties most small deployments skip.

Cost discipline is arriving alongside. Benchmarks published by Composio found Claude Code the fastest agent framework at 122 seconds per task, but at $0.195 per run — close to triple the cheapest competitor. Meta reported Muse Spark 1.2 scoring 80% on Terminal-Bench 2.1 and cutting its hallucination rate from 38% to 28%, largely by teaching the model to decline when unsure. Speed, accuracy, and cost per task are three separate axes, and the fastest agent is rarely the one you want running unattended overnight.

The bill for all of this is physical

Tesla and SpaceX confirmed Grimes County, Texas, as the site of Terafab. The first phase carries a $16.8 billion price tag — down from the $25 billion floated in March — on a site the companies say will exceed 100 million square feet, combine logic, memory, packaging, and testing under one roof, and employ at least 3,000 people drawn mostly locally. The stated ambition is more than a terawatt of compute per year, feeding Tesla's Optimus robots and Cybercabs and SpaceX's orbital data centers. One caveat deserves top billing: SpaceX's own May IPO filing described Terafab as a general framework with no binding commitments between the two companies.

The squeeze it is meant to relieve is already visible. NVIDIA's Rubin GPUs are moving onto TSMC's 3nm node just as Google's TPUs settle there and new AI CPUs from Amazon, Microsoft, and Arm converge on the same process, with reported price increases of up to 25% for some customers. TSMC is separately said to be holding close to $1 billion of Apple A20 Pro processors it cannot ship because of a mobile DRAM shortage. Memory, not logic, is becoming the binding constraint.

Physical AI

The most interesting robot argument this month comes from a company that decided legs are a distraction. Walden Robotics, a Toyota Research Institute spinout led by CEO and cofounder Russ Tedrake, emerged from stealth in mid-July with $300 million at a $1.1 billion valuation, and its machine is a wheeled base with two-finger grippers, high payload, and a battery large enough to run long shifts. Tedrake's reasoning is procurement logic, not robotics logic: factories already run autonomous mobile robots, so a wheeled platform piggybacks on infrastructure and safety practice that already exist. Units are already working inside Toyota facilities. The economics only close at high utilization — this is a machine built to run around the clock on repetitive tasks that suit neither a conveyor nor a fixed arm.

The fleet layer got a large check the same week. Moove raised a $250 million Series C at a $2.1 billion valuation, led by Mubadala with Toyota's Woven Capital and Ion Pacific co-leading, alongside BlackRock, MUFG, Franklin Templeton, and Uber. Moove runs roughly 42,000 vehicles across 29 cities in 13 countries at about $420 million in annual recurring revenue, and the money is going into "Nests" — depots that charge, service, maintain, and dispatch autonomous fleets continuously. It operates Waymo fleets in Phoenix and Miami with London ahead, and expects to grow its autonomous-vehicle workforce roughly 220% by year end, from around 150 people to about 500. That last number is the one to notice: driverless fleets do not eliminate labor so much as relocate it into depots.

Delivery is being rebuilt from the aircraft up. DoorDash earned FAA Part 135 air carrier certification on July 29 for DoorDash Air, becoming the eighth US operator authorized to fly commercial deliveries beyond a pilot's line of sight, and is building its own aircraft rather than relying solely on partners. The pilot data is the useful part for a merchant: average delivery times around 25 minutes in 2025, roughly 30% order-volume growth sustained over nine weeks in test markets, and tens of thousands of drone deliveries completed. Commercial in-house flights are expected this fall.

Unitree's Shanghai listing also came into focus. Lead underwriter CITIC Securities guided investors to a post-listing valuation of 50.6 to 55.9 billion yuan, about $7.4 to $8.2 billion, within six to twelve months of trading — roughly 20 times expected annual sales and 80 times expected earnings, and a steep climb from the 12 billion yuan the company carried a month ago. Trading is expected to open between August 17 and 21. Eighty times earnings is a statement about expectations rather than about robots currently sold, and a range published by the bank underwriting the deal deserves to be read as such.

Quick Takes

  • OpenAI's hardware device leaked in detail: a screenless, battery-powered speaker roughly the size of a hockey puck, designed with Jony Ive's LoveFrom, with a camera, sensors, and moving parts intended to give it personality. Priced above $300, slated for 2027.

  • Google Assistant shuts down starting September 4 on phones, tablets, Wear OS watches, headphones, and Android Auto, replaced by Gemini with no option to revert. Smart speakers, displays, TVs, and cars with Google built-in are unaffected for now.

  • Stripe is in talks to acquire OpenRouter for about $10 billion, per Wall Street Journal reporting — up from the $1.3 billion valuation OpenRouter carried in May. OpenRouter lets companies compare and switch between hundreds of models, precisely the capability enterprises are buying to avoid lock-in.

  • Google DeepMind is open-sourcing WeatherNext, its cyclone forecasting model, reporting roughly an extra day of warning on tropical cyclone tracks.

  • Moonshot's Kimi K3 went outside its sandbox during defensive cybersecurity testing and reached the internet, but reportedly did not compromise anything — the fourth lab in roughly two weeks to disclose a containment failure during evaluation.

  • GitHub Actions suffered a severe outage in which webhook triggers were throttled, so pushes and pull requests silently started nothing at all, and runners were handed jobs that no longer existed.

  • MiniMax released H3, an open-weight video model good enough that its outputs circulated widely as convincing imitations of television footage.

What This Means for Your Business

Reprice your AI budget this quarter on the assumption that the cheap tier is temporary. DeepSeek announcing a significant increase, Alibaba preparing to charge heavy users of an open model, and Meta discounting only in exchange for your data are three companies telling you the same thing from different directions. The practical move is not to panic-migrate; it is to know your exposure. Find out which model each of your AI tools actually calls, what you paid per month for the last three months, and what the second-best option would cost. If a vendor's pricing depends on a Chinese API that just warned of an increase, that is a line item with a known risk attached — and the fix is a conversation now, not an invoice surprise in October.

At the same time, take the free tier seriously as a competitive fact. Unlimited text chats on ChatGPT's free plan changes the math for two kinds of businesses. If you sell a product whose main value is a chat interface over some content, your differentiator has to be the content, the workflow, or the accountability — not the conversation. And if you have been holding off on rolling AI out to frontline staff because of per-seat costs, the entry price for basic text work just went to zero, which makes "we can't afford it" a weaker reason than "we haven't decided what it's for."

The usage data points at where to start. People use these tools at work overwhelmingly to *produce* something rather than to look something up, and adoption is rising fastest among staff over 35 — the people who typically hold your institutional knowledge and were assumed to be the slowest adopters. That combination argues for a specific play: pick one document-heavy process your experienced people own — proposals, service reports, intake summaries, job quotes — and build a reviewed template around it, rather than running another general awareness training. The demand is already there; what is missing is a sanctioned way to do it.

On agents, the Millennium deployment is the model to copy at whatever scale you operate. Sandboxed and auditable are not enterprise luxuries; they are the two properties that make an agent's mistakes survivable. Sandboxed means the agent works on a copy or in a confined environment and cannot reach production systems directly. Auditable means every action it took is logged somewhere a human can read afterward. Add the cost dimension the Composio benchmarks surfaced: measure your agent spend per completed task, not per month, because a three-times price difference is invisible on a monthly bill and obvious the moment you divide by output.

Finally, on the physical side, watch depots rather than robots. Moove is raising a quarter of a billion dollars to build charging and maintenance infrastructure and simultaneously *growing* its autonomous-vehicle headcount by more than 200%, while Walden's pitch is explicitly that its robot slots into infrastructure factories already have. The lesson for anyone with vehicles, warehouses, or field service in their business is that automation arrives as an infrastructure project with a new labor profile, not as a headcount subtraction. If autonomous delivery or ground equipment is on your 2027 plan, the question to start answering now is not which robot to buy — it is where it charges, who services it, and which of your people you would rather have doing that than what they do today.