Agents of Work
August 21, 2026 · Agents of Work

Agents of Work AI Daily Briefing — August 21, 2026

The company that makes Claude is about to ask public markets for more money than any company has ever raised in a first-time share sale, and the number it is reaching for is not subtle. Elsewhere today: coding agents moved out of the terminal and into Slack channels where the whole team can watch them work, a survey found one in five enterprises has no way to stop an AI agent from spending money, Waymo revealed the custom chip that has quietly been running inside every one of its robotaxis, and Tesla's Austin fleet appears to have gone fully driverless without an announcement.

Anthropic is going public, and it wants the record

Anthropic is preparing to file public IPO paperwork as soon as the end of August, and it expects the offering to match or beat the size of SpaceX's record-setting debut earlier this year, Bloomberg reported on August 20. The company had already submitted a confidential draft registration statement to the SEC in June. A market debut could come as soon as October.

The comparison is the story. SpaceX raised $75 billion in its listing — the largest first-time share sale ever recorded — a figure that climbed to $86.2 billion once the overallotment option was exercised, at a valuation near $1.77 trillion. Anthropic is reportedly targeting a post-listing valuation around $2 trillion. That is roughly double the $965 billion valuation it carried in May 2026, three months ago. Morgan Stanley, Goldman Sachs, JPMorgan Chase and Citi are working on the offering, with more underwriters possibly joining. Ahead of the filing, Anthropic is finalizing a revolving credit facility expected to raise more than its roughly $10 billion target, and is weighing super-voting shares that would give chief executive Dario Amodei and his co-founders outsized control after listing. Amodei is reported to own roughly 2% of the company.

The financials underneath are genuinely extraordinary in both directions. Anthropic's annualized revenue run rate reached about $65 billion by the end of July, and second-quarter 2026 revenue came in above $11.5 billion on a preliminary basis — against $787 million in the second quarter of 2025. The company reported positive adjusted operating income in Q2. It also posted a net loss of roughly $42 billion in 2025, after $8.3 billion in 2024. Those two facts sit in the same prospectus.

For a small business, the IPO itself is not the point — the disclosure is. An S-1 has to publish gross margins, customer concentration, contract terms and compute commitments that private labs have never had to show. Within weeks, anyone deciding whether to build a workflow on Claude will be able to read, in audited form, what it costs Anthropic to serve a request. That is the first time the AI industry's unit economics will be visible to the people buying from it — and a public company answers to shareholders on margin, which historically is not good news for the price of the thing you are buying.

Coding agents moved into the group chat

Slack launched Slack Code on August 20, and the design choice worth noticing is where the work happens. Tag a coding agent from any conversation and it spins up a dedicated project channel with separate tabs for discussion, planning, code diffs and a live preview. The team reads the plan, inspects the changes, gives feedback, and approves before anything ships. When the task ends, the channel archives itself and stays as an audit log.

Founding partners include Claude, ChatGPT, Devin, GitHub Copilot and Vercel, with agents from Lovable, n8n, LangChain and Superhuman available through an add-to-Slack flow. Slack Code is free on every Slack plan; access to each agent still has to be bought from the vendor. "In the Code channel, the whole team and the agent work together on the build," said Katie Steigman, Slack's VP of product. Rob Seaman, EVP and general manager of Slack, framed the bet more bluntly: "AI only creates value when it's part of how a team actually works."

Strip away the developer framing and this is a governance product. Agent work has lived in one person's terminal, invisible until it landed. Putting it in a shared channel with an explicit human approval step turns a private tool into a reviewable process — exactly the gap that has kept most small teams from trusting agents with anything touching production.

One in five companies cannot stop a runaway agent

VentureBeat's VB Intelligence surveyed 107 enterprises — software engineers, machine-learning professionals and analytics leaders — and found that 21% have only after-the-fact monitoring of agent spending, with no way to halt it in real time. Of the rest, 30% rely on native platform budget controls, 25% built custom middleware, and 25% route work to cheaper models dynamically.

The reason this is hard is visible in the same data: 85% of respondents run two or more orchestration platforms and 64% run three, with Microsoft AI Foundry at 70% adoption, OpenAI's Agents SDK at 68% and Anthropic's Claude Platform at 47%. A spending policy that only covers one platform does not cover the company. The survey found no correlation between company size and spending discipline — smaller organizations sometimes did better.

The most useful finding is the deflationary one. Only 2% of respondents said 76–100% of their systems are genuinely autonomous, and 35% admitted that just 1–25% of the things they call agents perform real multi-step orchestration. The sample is 107 companies and skews technical, so read it as directional. But the shape is clear: most "agent deployments" are still scripted automation with a new label, and the ones that are real can spend money faster than their owners can stop them.

The enterprise race got closer, and everybody is switching

Ramp's corporate-card data, covering more than 70,000 US businesses, shows Anthropic at roughly 44% of customers paying either lab in July, with OpenAI at nearly 40% — and OpenAI growing faster so far in the third quarter, according to Ramp economist Ara Kharazian. Anthropic had led since May, when it hit 41% against OpenAI's 39%. Meanwhile 56% of Ramp customers now pay for at least one AI product, up from 50% in March.

Ramp's customers skew toward tech and the data counts customers rather than dollars, so it is an indicator rather than a census. But the volatility is the signal: businesses are moving between providers on the strength of individual model releases, which tells you enterprise AI spending is nowhere near as sticky as the valuations assume — and that switching costs, for now, remain low enough that you should not accept a long lock-in.

Model and product news

Z.ai released GLM-5.3, a 743-billion-parameter model post-trained from the GLM-5.2 base. It leads the CyberGym cyber-defense benchmark at 84.5%, up from 77.2%, and lifts Terminal-Bench 3.0 from 4.6% to 28.3% on Z.ai's own evaluations. Two caveats are worth carrying: those are the vendor's numbers, and the model is not a blanket leader — competing models still score higher on general reasoning-with-tools work. The open weights are being held back pending safety review, expected roughly two weeks after launch, so for now access runs through Z.ai's API and its Coding Plan, priced at $18, $72 and $160 a month.

Mistral launched Agentic Search on August 20, a retrieval layer built to let a model navigate and verify information inside complex documents rather than pulling flat chunks of text. Mistral reports better accuracy with fewer turns and lower latency against the FinanceBench and OfficeQA Pro benchmarks. It shipped alongside OCR 4, adding bounding boxes, block classification and 170 languages. For any business sitting on a decade of PDFs, that combination is more practically useful than most model releases.

OpenAI extended ChatGPT into Apple Messages on Apple-silicon Macs, working across iMessage, SMS and RCS within ChatGPT Work and Codex. Sending requires approval by default, but users can grant persistent permission per conversation — and OpenAI's own guidance flags a known issue where scheduled tasks can disable the approval prompt. Before anyone on your team turns this on, that is the sentence to read twice. Separately, Reuters reports Anthropic plans to let business customers keep data from covered models inside their own cloud environments later this year, though a 30-day retention requirement would remain. It changes where the data sits, not how long it exists.

Physical AI

Waymo disclosed its first custom chip, and the notable part is that it is not a roadmap item — the purpose-built 5nm ASIC, fabricated on TSMC's automotive-grade N5A process, is already running in every Waymo robotaxi in commercial service. It delivers more than 1,000 TOPS of machine-learning performance dedicated to front-end sensor processing, fusing lidar, radar and camera streams in real time with dual-chip failover. Waymo describes responsiveness as one of three non-negotiable design pillars, with decisions processed within milliseconds. The disclosure landed alongside the opening of the Ojai robotaxi platform to all riders in Los Angeles, Phoenix and San Francisco. Waymo now operates in 11 US cities with plans for nearly 20 more plus London and Tokyo. The strategic read: after years of buying compute, the autonomy leaders are designing their own, because latency is a safety spec and you cannot buy latency off a shelf.

Tesla appears to have crossed the same line from the other direction — quietly. A crowdsourced project called Robotaxi Tracker, built by Ethan McKanna, logged 170 robotaxi rides across 54 vehicles in Austin over two weeks with no safety monitor on board in any of them, and roughly 30 unsupervised Teslas have been spotted in Dallas and Houston in the past week. Two things temper it. McKanna interned with Tesla's robotaxi team this summer, and the dataset is crowdsourced rather than company-confirmed — Tesla publishes no fleet data. And the mileage gap is enormous: Tesla claims zero notable incidents across 380,000 robotaxi miles without defining "notable," against Waymo's roughly 220 million unsupervised ride-hailing miles — Tesla's total is about 0.2% of Waymo's experience. A passenger also recorded a robotaxi creeping forward, reversing, then driving straight through plastic bollards that have stood since February 2024.

Down at ground level, the deployments that will actually touch a small business were less dramatic. Serve Robotics announced on August 17 that its sidewalk delivery robots are joining Grubhub — more than 100 participating merchants in Chicago and nearly 200 in Los Angeles, plus Alexandria, Virginia — days after saying it would not renew its Uber Eats agreement, which lapses early next year. Serve also launched DoorDash deliveries in Washington, D.C. and San Jose. If you run a restaurant in one of those metros, sidewalk robot delivery is now a channel you can be listed in rather than a pilot you have to join.

Diligent Robotics, now a Serve Robotics company, began rolling out Moxi 2.0 to US health systems the same day, including Endeavor Health Edward Hospital, Providence Saint John's Health Center and Children's Hospital Los Angeles. Upgraded NVIDIA A2000 compute lets it perceive its surroundings 10 to 15 times faster than the prior generation, and it runs up to 9 hours per charge with 30% faster charging — as much as 18 hours of daily operation, drawing on five years of experience across more than 25 hospitals. In warehousing, Pudu Robotics launched the MP2000 pallet-handling robot on August 18: a 2,000 kg payload, fork-in in as little as 20 seconds, and a claim that it deploys in minutes without site modification. Pudu has not published a price, which is the number that decides whether any of this matters to a mid-size operation.

Quick Takes

  • Google shipped an embeddable Preferred Sources button for publishers on August 20. Readers click it to favor a publication across Search, Discover and Google News. Google says people are twice as likely to click through to a preferred source, and that users have selected over 345,000 unique sources since the feature launched in May.

  • AI-data startup Micro1 reportedly went from a $100 million to a $500 million gross annualized run rate in eight months — but passes roughly 60–70% of billings through to the experts doing the work, implying a net figure closer to $150–200 million. A useful lesson in reading "run rate."

  • Workers at Spanish outsourcing firms serving Apple published first-person accounts through the Data Workers Inquiry describing annotation quotas they say can become mathematically impossible, active-time and mouse monitoring, language-based pay gaps and abrupt layoffs. These are worker accounts, not audited findings.

  • 404 Media documented "subtlefakes" spreading on X — real photographs altered only slightly rather than replaced with obviously synthetic scenes, some carrying Grok watermarks. Preserving most of the original makes the abuse harder to spot than a conventional deepfake.

  • A peer-reviewed Nature Communications study finds diffusion-model outputs are often unattributable to any individual training example, even when the training set is known — complicating the assumption that every generated image has one findable original behind it.

  • Reddit citations in ChatGPT responses fell from as high as 4.5% to about 0.5% after OpenAI changed how it searches the web.

  • Apple reportedly laid off 60 Vision employees as resources shift toward AI-powered smart glasses, and a video asset in the latest macOS update appears to show future AirPods with camera sensors.

  • FORT Robotics is going public via SPAC at an expected valuation above $500 million, and Tesla says FSD v15 is a "step-change" with Optimus sales planned for 2027 — the third consecutive FSD version pitched as the one that gets there.

  • Cursor launched Origin, Anthropic shipped a meeting recorder, and the Stark accessibility connector arrived in Claude's directory for scanning Figma files, URLs, source code and mobile builds.

What This Means for Your Business

Read Anthropic's S-1 when it lands — you are one of the customers it describes. This is the first time an AI lab will have to publish, under audit, what it costs to serve inference, how concentrated its revenue is, and what it has committed to spend on compute. If you have built anything meaningful on Claude, three numbers in that filing tell you what your next two years look like: gross margin on API revenue, customer concentration, and the size of the multi-year compute obligations. A company losing $42 billion a year while growing revenue fifteenfold is not going to keep prices where they are forever, and once it is public, the pressure to fix margin comes with a quarterly schedule. Do not panic-migrate. Do read the document, and do make sure the contract you sign this quarter is not longer than your ability to leave.

Put a hard spending ceiling on every agent before you put it on a real task. One in five surveyed enterprises — companies with actual platform teams — has no real-time way to stop an agent from spending. A small business has less tooling, not more. The practical version is unglamorous: use a separate API key or billing account per agent, set a hard monthly cap at the provider rather than an alert, and confirm in writing what happens when the cap is hit — throttle or stop. If your provider only offers an email when you cross a threshold, that is monitoring, not a control. And apply the survey's honest finding to your own stack: if the thing you bought does not chain multiple steps and make decisions between them, it is automation, and you should stop paying agent prices for it.

Adopt the Slack Code pattern even if you never install Slack Code. The valuable idea is not the product, it is the shape: agent work happens in a visible shared space, the plan is posted before the work starts, a named human approves before anything reaches customers, and the whole thread survives as a record. You can implement that today in whatever your team already uses. The reason to bother is that the most common way AI work goes wrong in a small company is not a bad model — it is that one person ran something in private, it half-worked, and nobody else could see it until a customer did.

Audit the permissions on every AI integration that can send something. ChatGPT can now read, search and send Apple Messages. Slack agents can open pull requests. The pattern repeating across this week's product news is that assistants are gaining write access to the channels you talk to customers through, and the approval prompt is the only thing standing between a draft and a sent message. OpenAI's own documentation flags a case where scheduled tasks can bypass that prompt. Go through your connected apps this week and ask a single question of each: can this send, post, pay, or delete without a human clicking yes? Anything that can, either revoke it or write down who is accountable when it does.

On the physical side, watch which robots are being sold as channels rather than machines. Waymo's chip and Tesla's driverless Austin fleet make headlines, but the item a small business can act on is Serve joining Grubhub with 100-plus merchants in Chicago and nearly 200 in Los Angeles. Nobody bought a robot in that transaction — restaurants got listed on a platform, and the automation showed up as a delivery option. That is how this technology will actually reach most small operations: not as capital equipment you evaluate, but as a checkbox inside a platform you already use. The right posture is to say yes to the pilot and measure it, since the downside is a slower delivery and the upside is a cost structure your competitor does not have. Save the skepticism for anything that asks you to buy hardware, where the question remains the one Pudu did not answer: what does it cost, and how many hours of labor does it actually replace?