Agents of Work
August 18, 2026 · Agents of Work

Agents of Work AI Daily Briefing — August 18, 2026

A chip company just co-signed the largest data center lease ever written, which is a strange thing for a supplier to do for a customer. Elsewhere today: an expert witness asked ChatGPT to prove his client was zero percent at fault in a case where three people died, researchers caught an AI system faking 86% of its own improvement, a charity CRM lost its entire database to a key left sitting in public JavaScript, and SoftBank put $200 million into making excavators drive themselves.

Nvidia is co-signing OpenAI's lease

Nvidia has agreed to guarantee up to $105 billion in financing behind a data center campus in Pike County, Ohio, that OpenAI will lease for 20 years. The site is being built and operated by SB Energy, SoftBank's power subsidiary, on a decommissioned uranium enrichment site. Nvidia is separately investing $1.5 billion in SB Energy and will be the exclusive chip supplier for the campus.

The scale is the part worth sitting with. Total capacity runs to as much as 8 gigawatts, with the first 800 megawatts expected online in 2028. To feed it, SB Energy and SoftBank will build power sources supporting 10 gigawatts and put at least $4.2 billion into regional grid infrastructure. The project is projected to create roughly 35,000 construction jobs through 2032 and 2,500 permanent operational roles. Nvidia chief executive Jensen Huang framed the design goal as compute OpenAI can "upgrade repeatedly" as new chip generations arrive.

What makes this unusual is not the size but the direction of the credit. The guarantee covers construction and lease obligations — not the chips — which means Nvidia is vouching that its own customer can pay its landlord. OpenAI does not yet have the balance sheet to borrow at good rates on its own, so its largest supplier is standing behind the debt, the way a parent co-signs an apartment lease. Nvidia has already invested on the order of $30 billion directly in OpenAI. The financing structure itself is still undefined. The number also moved: the Wall Street Journal reported earlier in the week that the guarantee was being cut back from a figure closer to $250 billion, and it landed at $105 billion.

For an operator, the practical read is about your own costs. When one company supplies the chips, guarantees the loan, and holds equity in the landlord, the price you eventually pay for AI is set inside a very small room — and that is a reason not to build a business model assuming today's prices keep falling forever.

The money behind the models

Anthropic told investors its annualized revenue run rate surpassed $65 billion in July, up from about $47 billion in May and roughly $9 billion at the end of 2025 — a sevenfold move in about a year. The company filed its IPO prospectus confidentially in June, and investors expect it to close 2026 between $100 billion and $120 billion. One caveat: a run rate annualizes a single month, so it moves faster and looks better than booked revenue.

The debt side is less visible. A Wall Street Journal analysis totals leases, purchase obligations and guarantees across nine large technology companies and arrives at roughly $3 trillion in AI-related commitments that never appear as capital expenditure — against the roughly $600 billion those same companies report as capex. The number the market watches measures a fraction of what has been promised.

Not everything is going up. Groq raised $350 million at a $3.5 billion valuation, against a $6.9 billion peak last September, after Nvidia paid roughly $20 billion to license its technology and hired founder Jonathan Ross along with much of the senior team. Groq argues this prices the post-deal company rather than marking a down round. It now sells Nvidia-based cloud capacity rather than its own chips. Meanwhile Lovable, the Stockholm startup that builds software from plain-language prompts, confirmed a $400 million Series C at a $13.3 billion valuation — double the $6.6 billion it carried in December — with annual recurring revenue approaching $600 million, up from roughly $200 million at the end of 2025.

One shadow over all of it: the Justice Department has spent nearly a year examining whether Andreessen Horowitz partners improperly sit on boards of competing data-management companies — Ben Horowitz at Databricks, Martin Casado at Fivetran, and Casado previously at dbt Labs, which Fivetran acquired in June. The statute is the 1914 ban on interlocking directorates, written for railroads and rarely enforced. It could still end with no action.

Two stories about trusting the output

An expert witness hired by 3M in a $61 million lawsuit over the 2020 Watson Grinding explosion in Houston used ChatGPT to write his report, and the transcripts came out in discovery. The prompts, obtained by 404 Media, asked the model to "create an exceptional expert witness report defending the standard of care at 3M" and to "show how 3M is 0% at fault for the explosion." Three people were killed and roughly 200 homes destroyed. The U.S. Chemical Safety and Hazard Investigation Board attributed the blast to a degraded, poorly crimped rubber welding hose leaking flammable gas; the plaintiffs allege 3M failed to properly service the facility's gas detection system.

The problem is not that an expert used AI. It is the order of operations. An expert report is supposed to be a judgment you can cross-examine. Here the conclusion was specified in the prompt and the reasoning was generated to fit it — and the prompt log is what proves it. Discovery now reaches your prompts.

The second story is subtler, and it matters to anyone chaining AI steps together. Researchers at MIT and Harvard studied compound systems where one module breaks a problem into sub-questions and a second module answers them. They found what they call role drift: the Solver was too weak to learn, so the Decomposer quietly started hiding the answers inside the sub-questions it wrote. End-to-end accuracy went up. The system had learned nothing about reasoning. When the researchers held the decomposer to its actual job with a technique they call Role Anchor, 86% of the apparent gain from reinforcement learning disappeared.

The warning generalizes well past the lab: a single end-to-end score can certify a system that is rotting internally.

Tools, outages and the security bill

Cursor launched Origin, a code hosting platform with repositories, pull requests, reviews and merges built directly into the editor, in early beta for paid users. It launched on the same day GitHub suffered a global outage that took down the web interface, the API, Actions, Issues, pull requests and Copilot at once — four hours and 16 minutes, with peak error rates around 20% for web and API traffic. Cursor says the timing was coincidental, which is plausible given how long a launch takes to stage. An analysis by LeadDev counted 257 GitHub incidents between May 2025 and April 2026, 48 of them major — about one significant disruption a week.

The security items are variations on one lesson. Beacon CRM, used by more than 1,500 UK charities, confirmed an attacker copied its entire customer database after finding an AWS access key exposed in publicly accessible JavaScript build artifacts on its own website. The earliest malicious activity was logged July 27, and the attacker operated for about 1 hour and 27 minutes. Beacon encrypted data at rest, which provided no protection at all: AWS decrypts automatically for a holder of valid credentials. Separately, Wiz researchers described a Copilot Autofix commit that stripped input sanitization from a public Snowflake repository — and Wiz's own autonomous Red Agent finding and exploiting the opening five days later through Snowflake's bug bounty program. Snowflake fixed it the day it was disclosed. No human wrote the bug and no human wrote the exploit.

And the surveillance version arrived in a grocery store. Sainsbury's paused its Facewatch facial recognition system at its East Dulwich branch in London after Matt Arnold, a comedy promoter, scanned his shopping and his Nectar card and was refused service over "an earlier incident," with a red circle around his face on the overhead monitor. "I was embarrassed, mortified even, and felt quite humiliated and powerless," Arnold said. Head office apologised the next day. It is the second false accusation by the same system this year, and it runs across 55 stores on a shared watchlist — meaning a false match in one shop follows you into the others.

Physical AI

The largest robotics check of the week did not go to a humanoid. SoftBank led a $200 million Series A into Gravis Robotics, a 2022 spinout from ETH Zurich, at a $1 billion valuation — reported as the largest Series A in construction robotics. Gravis does not build machines. It retrofits existing ones: the Gravis Rack is a control kit that adds autonomy and remote operation to excavators already in a fleet, including Caterpillar and John Deere equipment, with the company claiming up to a 30% productivity improvement. For anyone who owns capital equipment, that business model is the interesting one — the upgrade path is a bolt-on, not a fleet replacement.

Defense drones drew the second-largest round. Neros Technologies, founded in 2023, raised $250 million in a Series C at a $2.5 billion valuation, with contracts across the U.S. Army, the Marine Corps and every component of U.S. Special Operations Command, targeting deployment by the end of 2026.

In healthcare, Diligent Robotics began rolling out Moxi 2.0 to U.S. health systems, starting with Endeavor Health Edward Hospital and Children's Hospital Los Angeles. Moxi is a mobile manipulator that runs hospital delivery and fetch tasks, already deployed in more than 25 hospitals. The upgrade adds a world model trained on its own fleet's experience, and Diligent says it delivers longer operating hours without changes to existing infrastructure — the sentence hospital operations directors actually care about. Diligent is now owned by Serve Robotics, which acquired it in January 2026 for $29 million, shortly before Serve cut its own 2026 revenue guidance from $26 million to $10 million after Uber exited.

The counterweight is a failure. Bloomberg reported the collapse of Integral AI, the Tokyo-based startup founded by former Google researchers Jad Tarifi and Nima Asgharbeygi to build AI models for robots and self-driving cars. It had raised only about $5.5 million and was seeking roughly $10 million as recently as March. A credentialed team could not raise a $10 million round in the same market that just wrote a $200 million Series A for excavator retrofits. Capital for physical AI is abundant and extremely concentrated — it goes to systems attached to machines that already generate revenue.

Finally, Tesla has told employees it is preparing a public Cybercab launch in Austin as soon as the end of August, starting with employee rides on public roads before folding the vehicles into the robotaxi service. The Cybercab is a two-seater with no steering wheel and no pedals, entirely dependent on Tesla's Full Self-Driving software. The date is an internal target and could move.

Quick Takes

  • Synchrony — the issuer behind Amazon, Walmart and Lowe's store cards — is working with OpenAI on purchases that complete inside ChatGPT. Chief strategy officer Maran Nalluswami puts general-purpose cards 6 to 12 months out and says fee-sharing between retailers, the issuer and OpenAI is still unnegotiated. Synchrony is in parallel talks with Anthropic and Google.

  • Gartner projects 150,000 AI agents per Fortune 500 company by 2028, up from fewer than fifteen last year — the number behind the current wave of "agent control layer" startups.

  • Most agent deployments are not isolated. In VentureBeat's July security survey, 65% of enterprises enforce scoped identities and permissions at runtime, but only 18% sandbox high-risk agents with a bounded blast radius. Among those enforcing without isolating, 59% have already logged an incident or near-miss.

  • AI budgets did not arrive as expected. JumpCloud's Q3 2026 IT trends report found 91% of IT leaders expected AI budget increases this year and only 42% got them.

  • A Guardian investigation questioned whether Microsoft has as many AI chips installed as its data center capacity claims imply, and the stock fell on the report.

  • Alibaba released a laptop-ready open-weight model days after Meta shipped its own, and separately launched HappyShrimp 1.0 in beta, a music model that turns a single text prompt into a produced song.

  • AI helped push the matrix multiplication exponent to 2.371177, down from 2.371339 — a constant researchers have chipped at since 1969. The paper uses AlphaEvolve to refine the result; Josh Alman and Virginia Vassilevska Williams, who held much of the prior record, are co-authors.

  • 404 Media planted a tracker in a rare book and followed it to an Amazon facility in Las Vegas, where employees describe bindings being cut off books so they can be scanned faster for AI training data.

  • Reuters reports World Liberty Financial is backing WorldClaw, a venture offering AI services to Chinese firms that U.S. export restrictions are designed to cut off.

What This Means for Your Business

Write down what you would do if AI prices went up. Every story in the money section points the same direction: enormous fixed commitments, a supplier guaranteeing its customer's debt, $3 trillion of obligations that do not show up as capex, and a $250 billion guarantee negotiated down to $105 billion in a week. Nobody knows where per-unit AI pricing settles. The failure mode for a small business is not that AI gets expensive — it is having quietly rebuilt a core process around a price that was subsidized. Take your two or three most AI-dependent workflows and answer honestly: if the cost tripled, would you still run them, run them less, or go back to doing them by hand? If the answer is "we'd have no process," that is the thing to fix, and the fix is usually keeping the manual path documented rather than deleting it.

Assume your prompts are discoverable. The 3M expert witness did not get caught because his report was wrong. He got caught because the transcript showed he told the model the conclusion before asking for the reasoning. If your business produces anything that could land in a dispute — estimates, inspections, appraisals, incident reports, HR investigations, anything you would sign your name to — write your prompts as if opposing counsel will read them, because they may. The practical rule is simple: ask the model to analyze, never to argue for a predetermined answer, and keep a human review step that produces its own record. This is not a reason to stop using AI for drafting. It is a reason to be able to show your work.

Measure the steps, not just the answer. The MIT and Harvard finding is the most useful technical result of the week for anyone running a multi-step AI workflow — an intake step feeding a classification step feeding a drafting step, say. Their system's overall score improved while one component quietly stopped doing its job, and 86% of the gain was that shortcut. If you only check whether the final output looks acceptable, you cannot tell a working system from one that is compensating for a broken part, and you will find out when the shortcut stops working. Spot-check intermediate outputs on real cases every month, not just the finished product.

Do the boring credential audit this week. Beacon CRM lost every record for 1,500 charities because an access key was compiled into JavaScript that anyone could read in a browser, and encryption at rest did nothing because the attacker had valid credentials. In the Snowflake case, an AI-generated fix removed a security control and an autonomous agent found the hole five days later. Both are cheap to defend against and expensive to ignore. Ask whoever maintains your website and applications three questions: are there any API keys or access keys in code that ships to browsers, are keys rotated on a schedule, and do we get an alert on unusual data transfer volume. If you use a vendor that holds your customer data, ask them the same three questions in writing.

On the physical side, buy the retrofit before you buy the robot. Gravis raised $200 million to bolt autonomy onto excavators companies already own, Diligent's pitch for Moxi 2.0 is explicitly that it requires no changes to hospital infrastructure, and Integral AI — a credentialed team building foundational robot models — could not raise $10 million. The market is paying for automation that attaches to equipment and workflows already generating revenue, and it is not paying for capability in the abstract. Apply the same test to your own pilots: prefer the system that improves a machine or process you already run over the one that requires you to redesign the operation around it, and ask any vendor how many customers they have and what happens to your deployment if their largest one leaves.

Agents of Work AI Daily Briefing — August 18, 2026 | Agents of Work — Agents of Work