Three days after shipping the model it called the start of the AGI era, OpenAI published two documents on the same Sunday: a tally of how much of its own research is now performed by AI agents, and an essay from its chief scientist arguing that no lab should be scaling at full speed. Elsewhere, a headline benchmark score turned out to depend on the harness that produced it, Claude finished a computer-checked proof of Fermat's Last Theorem, California's software tax got a start date, and a study of two million listings found AI search shows shoppers more expensive products.
OpenAI published the numbers on its own agents, then asked the industry to slow down
On Sunday, September 6, OpenAI released data from inside its research organization showing how much of the work is now being done by AI. By mid-August the company was deploying 3.1 agent-workdays for every human researcher workday, having passed the crossover point in June. The median researcher was consuming more than $600 a day of inference at API prices; the 90th percentile was above $7,000 a day. OpenAI says it has reached its "automated research intern" milestone — agents handling well-defined research tasks that would take a skilled human days — and is targeting a fully automated AI researcher by 2028. The number that should temper the rest: more than half of the successfully completed four-to-eight-hour agent tasks still required at least one human intervention. Autonomy at that horizon is real but supervised, and the supervision is not optional.
The same day, chief scientist Jakub Pachocki published an essay titled "An Alien Mind" containing the sentence that made the week: no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. His specific worry is monitorability. Chain-of-thought oversight — reading a model's verbalized reasoning to understand what it is doing — is getting less useful as models get better at reasoning about their own processes, and he says its effectiveness is progressively diminishing. He called for mandatory safety standards enforced by third parties, international coordination, and labs publishing their progress on recursive self-improvement.
Read together, the two documents describe one loop: the company is measurably better at using AI to build AI, and the person responsible for the science says the safety work underneath has not kept pace. The distinction Pachocki draws is worth borrowing. Goal alignment asks whether a system pursues the objective you gave it; value alignment asks whether the constraints you assumed survive when achieving that objective gets difficult. Every operator handing an agent a multi-hour task is making the second bet, usually without knowing it.
The 99.9% came from the harness, not just the model
OpenAI's launch numbers for GPT-6 Astra included a near-perfect score on ARC-AGI-3. The ARC Prize Foundation's own writeup splits that figure in two, and the gap is large enough to change what the number means. On the Standard harness — a neutral interface where the model carries forward only the notes it chooses to keep — Astra scored 62.7% at a cost of $26,098. On a Provider Adapter harness that uses OpenAI's native context management and preserves opaque reasoning state between requests, the same model scored 99.9% at $18,817. Higher score, lower cost, same weights.
The efficiency result underneath survives the caveat and is the genuinely striking part: in the Provider Adapter configuration, Astra used fewer actions than the human baseline on 96.0% of levels, and 51.7% fewer actions per level on average. It builds working models of unfamiliar environments faster than people do. ARC Prize is explicit that this is not a finish line, stating plainly that saturating the benchmark would not constitute proof of AGI.
The practical lesson is cheap to apply. Benchmark scores are now a property of the model plus the scaffolding around it, and vendors control both. When a number appears in a sales deck, ask which harness produced it.
Claude formalized Fermat's Last Theorem in eleven days
Anthropic published a complete, computer-checked formalization of Fermat's Last Theorem produced by a team of Claude agents in eleven days. The output is roughly 13 million lines of Lean, built from 30,300 proved intermediate theorems, about 29,500 of which appear in the final proof, consuming approximately six billion output tokens. It uses only Lean's three standard axioms and contains no omitted proofs — no placeholders standing in for unfinished steps, no convenient new axioms. Anthropic researcher Tianyi Peng led the effort; Imperial College London's Kevin Buzzard, who has spent years on the human formalization project, reviewed it and called it an extraordinary autoformalization achievement.
The caveats are in the paper. Early attempts failed outright, with agents losing track of the project's state and ceasing to collaborate. And this is translation, not discovery — Wiles and Taylor proved the theorem in the 1990s. What changed is the cost of verification. In a separate experiment, three personal Claude Max subscriptions formalized Vinogradov's Three Primes Theorem in three days, which is the number that should interest anyone whose business depends on checking work rather than inventing it.
Anthropic locked in a decade of compute, then moved its IPO
Reporting from The Information puts Anthropic's compute commitments at roughly $517 billion across 14.8 gigawatts signed in the past eleven months, against a company that previously held between one and two gigawatts — including about $200 billion with Google for TPU capacity, $45 billion with Nscale and $35 billion with Lambda. For scale, Anthropic had told investors it expected to spend roughly $180 billion on server rentals through 2029. The listing moved the other way: the prospectus, once expected in early September, has slipped to late September with investor marketing in October at the earliest, while the company finalizes a $15 billion revolving credit facility. Signing a decade of compute before you list tells public markets what you believe demand will be, and means the filing gets read as a bet on revenue that does not exist yet.
California starts taxing your software on January 1
California signed SB 122 on June 29, and on January 1, 2027 it applies sales tax to prewritten software delivered by any method — downloaded, streamed, or accessed in a browser. That covers most SaaS products and most AI application subscriptions. The rate is 7.25% at the state level plus local district taxes, landing between 7.25% and about 10.75% depending on where the buyer sits.
The exclusions matter as much as the inclusions: custom software built to order, infrastructure and platform services, and human effort delivered electronically fall outside the tax. Consumption-based AI pricing — paying per API call — remains genuinely unresolved. Remote sellers cross into a collection obligation at $500,000 in California sales. Every California customer's invoice grows by up to 10.75% on day one, the vendor collects it, and there are no input credits to recover it.
What AI search does to the price a customer sees
An e-commerce analytics study tracked more than two million product listings across more than 100,000 search result pages and AI Mode answers over 23 days in August, covering the US and UK. Where the same product appeared in both surfaces, it was 21.6% more expensive in AI Mode; across all listings, the median product in AI Mode was $149 against $100 in traditional search. Two findings are more consequential than the price gap: only 1.28% of products ranking in traditional search also appeared in AI Mode for the same query on the same day, and among matched products the main seller was different 49.6% of the time. That decouples two things most businesses still treat as one. A page's rank in classic search says almost nothing about whether it appears in the answer box, and the seller who wins the click is frequently not the seller who wins the summary.
Hiring bars start naming AI out loud
UBS has made AI proficiency a formal requirement for graduates and interns applying to start in Global Banking and Markets in 2027, according to the Financial Times. Candidates must demonstrate that they can use and experiment with AI responsibly to improve outcomes — not simply that they have opened a chatbot — and interviews now include questions about how they use the tools. Academic requirements stay in place. It is a small policy at one bank, and the first clean signal that "can you use this" has moved from a nice-to-have on a résumé into a screening criterion.
The chip loophole was a US subsidiary
A New York Times investigation published September 6 traced more than $5.6 billion in advanced technology moving from Aivres, a California-incorporated subsidiary of the Chinese server maker Inspur Group, through Southeast Asia and into Chinese hands between April 2024 and February 2026. More than $3 billion of that was computers built around Nvidia's Blackwell chips — the same silicon in current-generation American AI data centers. Commerce added Inspur Group to the Entity List in March 2023 over procurement supporting China's military; Aivres, separately incorporated in the United States, was not on the list and could buy what its parent could not. Reported destinations include firms serving Alibaba and ByteDance. Export controls written against corporate parents turn out to be enforced against corporate structure.
Physical AI
The most instructive robotics story this week is not a humanoid. Reframe Systems raised $40 million led by Energy Impact Partners, announced August 31, to expand a network of small, highly automated factories that build homes near the sites they serve. It was founded in 2022 by Vikas Enti, Felipe Polido and Aaron Small, all of whom ran teams at Amazon Robotics. The company claims roughly 35% lower cost and delivery up to three times faster than site-built construction. The number worth copying into a notebook is the plant itself: FAB1 in Billerica, Massachusetts was outfitted in under 70 days for less than $5 million in equipment, opens October 5, and is designed for up to 500 multifamily units or 250 single-family homes a year. Ten homes are complete, eight occupied, with 114 more units scheduled. That is a capital footprint a serious regional contractor could contemplate, which is not something you can say about a humanoid program.
The autonomous-vehicle story got stranger. The Financial Times reported that Atoms, the startup founded by Uber co-founder Travis Kalanick, is building robotaxi technology and has held preliminary talks about running it on Uber's network — with Uber confirming a $100 million investment made as part of the $1.7 billion round led by Andreessen Horowitz in June. The work is reportedly led by Anthony Levandowski, whose theft of Google self-driving trade secrets cost Uber roughly $350 million in settlement and contributed to Kalanick's removal as chief executive. Atoms denies the characterization, describing itself as an industrial software company with no plans to enter what it calls a saturated robotaxi market. Take the denial at face value and the personnel still tell you where the talent is pooling.
Against those headlines, the deployment record stays sober. Agility Robotics' Digit has accumulated more than 65,000 operating hours across nine customer facilities, with Schaeffler, GXO, Toyota Motor Manufacturing Canada and Mercado Libre named as commercial customers — a real number, and a small one next to the funding totals. Cost is the other filter. Frontier models are getting cheaper per task, but robotics practitioners testing them on manipulation keep landing on the same objection: a per-task cost that is trivial for a knowledge worker's afternoon is prohibitive for a machine repeating the task ten thousand times. Physical work multiplies unit cost by duty cycle, which is why warehouse robots still run narrow, purpose-built models rather than frontier ones. When a vendor quotes you a robot's capability, ask how many hours the fleet has logged in a building that looks like yours, and who was standing next to it.
Quick Takes
Jensen Huang declared that AGI has arrived, citing OpenAI's newest model training on roughly 100,000 Nvidia chips. He deleted a first version claiming 300,000 chips and reposted the smaller figure without explanation.
Microsoft rolled out Project Opal, a Copilot feature that hands multi-step office tasks such as audit prep and IT ticket triage to an agent working inside a virtual Windows PC. Limited to Frontier program testers.
Enterprise buyers are putting non-Nvidia accelerators ahead of Nvidia's next-generation GPUs on evaluation lists, as cost and supply become the binding constraints; Snowflake credited its multi-model CoCo coding assistant for part of a strong quarter.
DeepSeek reportedly plans a 160,000-chip Huawei Ascend cluster in Inner Mongolia, a domestic-silicon answer to the export-control story above.
Tokens are now loyalty rewards in China. Daily token use reached 500 trillion this year, up from 100 billion in early 2024; a Shanghai bank bundles up to 3 billion Qwen tokens with a credit card, and China Telecom sells 10 million tokens a month for about $1.40.
JetBrains urged users of its hosted Cadence service to revoke and rotate all credentials after unidentified attackers gained access.
Industrials made up 23% of the 245 companies launched so far in Y Combinator's S26 batch, against a five-year average of 6.6%; SaaS fell to 12% from a 28% average.
Perplexity shipped hybrid compute on Mac, splitting agent work between cloud models and local ones so sensitive material stays on the device.
Three hikers were rescued on Mount Shasta after using Gemini to plan supplies and packing far less food and water than the trip required.
What This Means for Your Business
Start with the intervention rate, because it is the most useful number published this week. Inside OpenAI — the most agent-saturated research organization on earth, running frontier models with unlimited budget — more than half of the successfully completed multi-hour agent tasks still needed a human to step in. If that is the state of the art, then any vendor selling you unattended agentic workflows is selling ahead of the evidence. Budget for review time as a permanent line item, not a transitional one. The right structure for a four-hour agent task is a defined checkpoint where a person looks at partial output, not a fire-and-forget queue you audit at the end of the month.
Then handle the tax, because it has a date on it. If you sell software into California, your invoices need to carry sales tax on January 1, 2027, which means your billing system, pricing page, and renewal conversations have to change before year-end — and how you structure an invoice materially affects what is taxable, since custom work and infrastructure services fall outside the tax while prewritten software does not. If you buy software, model an 8 to 10 percent increase on your California-sourced SaaS and AI subscriptions with no input credit to offset it, and use renewal season to consolidate the redundant tools you were already meaning to cut. Consumption-priced AI is unsettled; ask your vendors now what position they intend to take, because you will inherit it.
Treat the AI search finding as a distribution problem, not a pricing one. A 1.28% overlap between what ranks in traditional search and what appears in AI Mode means your search reporting is now measuring a surface many of your customers are not looking at, and a different seller wins the summary about half the time. Go run your own top ten product or service queries through AI search this week and write down what comes back — who is named, what price is shown, whether you appear at all. That is a thirty-minute exercise that tells you more than a quarter of rank tracking. If your prices are competitive but you are absent from the answer, the problem is visibility to the model, not your margin.
On evaluation, adopt one habit: ask which harness produced the number. The same model scored 62.7% and 99.9% on the same benchmark depending on the scaffolding around it, and the scaffolding is the vendor's to choose. That generalizes past benchmarks to demos. When you are shown an agent completing a workflow, ask what context it was given, what tools it had, how many attempts it took, and what it cost. Then run your own pilot on your own messy data before you sign, because the delta between a curated demo and your actual accounts receivable file is where most AI projects quietly die.
Finally, on people: UBS is asking 2027 graduate applicants to demonstrate how they use AI to improve their work, and that will not stay inside investment banking. You do not need a policy document. You need two or three named examples per role of what good use looks like in your business — the quote turned around in an afternoon instead of three days, the contract read for the clause nobody remembered, the customer email drafted from six months of history. Write them down, share them, and reward the people producing them. Announcing that you are an AI-forward company does approximately nothing; you get the behavior you reward.