Speed became a purchasable product this week, and the price of a fast answer is now a line item you can shop. Elsewhere today: investors are floating a $2 trillion IPO for Anthropic on a business that has yet to post net income, OpenAI's own 69-page study of enterprise AI could not find a link between how much a company uses AI and its revenue per employee, IBM is retraining tens of thousands of consultants on OpenAI's stack, and LG and Nvidia widened a physical-AI partnership that runs from a Tennessee washing machine plant to an 80-megawatt facility in South Korea.
OpenAI put GPT-5.6 Sol on wafer-scale silicon, and speed became a product tier
OpenAI opened a limited preview on August 13 of Ultrafast, an API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second — roughly 14 times the speed of the Standard tier — without swapping in a smaller, dumber model. The hardware is the interesting part. Ultrafast runs on Cerebras' Wafer-Scale Engine rather than a GPU cluster, holding model weights in on-chip SRAM instead of shuttling them across memory buses. That sidesteps the memory bottleneck that caps how fast a conventional GPU rack can emit tokens, which is why the speedup is a multiple rather than a percentage.
The demonstration Cerebras chief executive Andrew Feldman offered is worth holding onto: running the 2,500 questions of Humanity's Last Exam, Sol on Ultrafast finished in about 11 hours, against more than 78 hours for a comparable frontier model at similar accuracy. Same quality, a seventh of the wall-clock time. OpenAI is aiming the tier at work where latency rather than intelligence is the constraint — live voice, real-time coding, incident response, financial research, commerce — and named Jane Street, Podium, Basis, and Rogo among early testers.
For an operator, this adds a third axis to a decision that used to have two. Model selection has been a trade between capability and cost per token; how fast the answer arrives is now sold separately from which model you picked. A support agent that replies in one second is a different product from one that takes eleven, even when the words are identical — and for batch work that runs overnight, paying a premium for speed is pure waste.
Google shipped its third Flash model in three weeks and halved the price
Google began rolling out Gemini 3.7 Flash on August 13, three weeks after Gemini 3.6 Flash — an unusually short turnaround the company attributes to developer feedback and algorithmic gains. The benchmark movement is real and concentrated in agent and coding work: 43.6% on FrontierCode 1.1 Main against 34.4% for 3.6 Flash, 65.3% on DeepSWE v1.1 against 49.0%, and 30.4% on AutomationBench against 17.0%. The last of those nearly doubled in three weeks.
The pricing is the operator story. Through December 31, Gemini 3.7 Flash runs at $0.75 per million input tokens and $3.75 per million output — an explicit introductory rate. On January 1 it doubles to $1.50 and $7.50. Google is unusually transparent that the discount expires, which is more useful than it sounds: anyone building a cost model on this quarter's rate card should put the January number in now rather than discover it in a bill. The model is available through the Gemini API, AI Studio, Antigravity, Android Studio, the Gemini Enterprise Agent Platform, and to Google AI Pro and Ultra subscribers in 160-plus countries.
Investors want $2 trillion for Anthropic. The underlying business is not there yet.
Roughly half a dozen Anthropic investors have told reporters they are targeting a $2 trillion valuation for an October IPO, which would make it the largest public offering in history, ahead of SpaceX. The case rests on growth: Anthropic reached a $965 billion valuation in a May funding round, projects annualized revenue of $100 billion to $120 billion by the end of 2026, and told investors it expected Q2 revenue near $10.9 billion, more than double the prior quarter. At $100 billion of revenue, $2 trillion implies a forward price-to-sales ratio around 20 — lower than several listed AI-adjacent companies trading above 50. One investor called $2 trillion a lowball and floated $3 trillion.
The counter-argument is arithmetic. Applied to current Nasdaq 100 earnings multiples, a $2 trillion company would need something on the order of $59 billion to $79 billion in annual profit to justify the price. Anthropic is not making net income at all. Its own guidance was for an operating profit — a different and much friendlier number — and it flagged that staying there is uncertain because compute costs are scheduled to rise. Amazon, at a comparable market capitalization, posted $62.6 billion of net income in a single quarter on $200.6 billion of revenue. The valuation target also comes from investors rather than the company; senior executives are reported not to have set one even privately. That is a reason for any business standardized on a single lab to think about concentration: a vendor pricing an IPO against a revenue projection has strong incentive to protect it, and the available levers are rate cards, usage tiers, and enterprise terms.
OpenAI's own research could not find the productivity gain
Buried in a 69-page report OpenAI published on how enterprises use AI is a finding the company did not lead with: there is no statistically significant relationship between how heavily a company's employees use AI and that company's revenue per employee. The report's own language is that revenue per employee "is not meaningfully associated with output tokens per employee or messages per active user once other controls are included." What the data does show is a selection effect running the other way — firms with already-high revenue per employee tend to adopt ChatGPT earlier. That is a statement about who buys first, not about what buying does.
A second finding deserves a wider audience: executives use AI the least intensively of any group measured, by weekly messages per user. The people making adoption decisions have the least direct experience of the tool. Two of the five authors were academics working as paid contractors to OpenAI — worth knowing in both directions, since it is a disclosure and it makes the null result harder to dismiss as an outside attack.
Census Bureau survey data released this week grounds the picture. About 55% of workers say they use AI on the job. Roughly a third of those report it cuts one to two hours off a task; about a quarter report less than an hour; 15% report three hours and another 15% report four. The most common uses are searching for information (37%), writing and drafting (32%), and generating ideas (32%). The time saved is real and measurable. Whether it converts into revenue depends on what happens to the freed hour — a management question, not a model question.
The enterprise layer consolidates, and the cost fight moves to the harness
IBM announced a dedicated OpenAI practice inside IBM Consulting, with plans to train and certify tens of thousands of consultants — mostly existing employees — on OpenAI's Codex, API, and cybersecurity credentials over the coming months, plus a group of "Forward Deployed Experts" trained through OpenAI's partner network. The two will jointly market industry-specific offerings for financial services, government, telecom, and retail. It comes less than a year after a comparable IBM alliance with Anthropic, and a month after IBM cut its 2026 revenue forecast. The pattern for buyers: the integrator layer is becoming multi-lab by design — good for negotiating leverage, bad for anyone hoping a consultant's recommendation is neutral.
Writer took a different route at the same problem. Its new flagship, Palmyra X6, is a post-trained variant of Z.ai's open-source GLM-5.2, and it shipped August 13 alongside upgraded harness infrastructure — the orchestration layer that decides how many model calls a task actually needs. Writer estimates the combination cuts customer costs by about 50% on basic tasks, and says harness optimization alone produced an average 40% reduction across the models it tested. Chief executive May Habib was blunt about the demand behind it: enterprises are "absolutely sick of chasing the next benchmark" and want flattening cost. The claim to test is that harness efficiency compounds across every model you run, in a way that swapping to a cheaper model does not.
DeepSeek is pushing the same lever from the supply side. Alongside V4-Pro and the MIT-licensed DeepSeek Harness, it is abandoning flat API rates on August 16 at 16:00 UTC in favor of peak and off-peak pricing, with off-peak set at half the peak rate. Scheduling has not been a meaningful AI cost lever until now. For any batch workload without a deadline, it just became one.
Physical AI
LG and Nvidia widened their physical-AI partnership on August 14, and the shape of it says more than the humanoid headline. LG is building a humanoid on Nvidia's Isaac GR00T foundation model with Jetson Thor onboard compute and the Halos safety framework, targeting a public unveiling in the first quarter of 2027, with actuators, sensors, and batteries coming from LG Electronics, LG Innotek, and LG Energy Solution. Nvidia founder Jensen Huang framed the category as giving "every machine the ability to understand the real world, reason and act safely alongside people." The nearer-term item is the unglamorous one: LG will put its CLOiD wheel-based robots into an LG Electronics washing machine plant in Tennessee later this year for performance testing. Wheels in a working US factory this year is a more useful signal than legs on a stage next year. The deal also covers a reference AI factory on Nvidia's Vera Rubin platform in the first half of 2027 and an 80-megawatt facility in Cheonan, South Korea by the first half of 2028.
Autonomous vehicles took their largest single step into Europe. Pony.ai and Uber announced on August 14 that they will deploy more than 2,000 Pony.ai robotaxis across Europe, expanding from the commercial service the two run with Croatian operator Verne in Zagreb into four additional European cities, with Middle East deployment also planned. The structure matters more than the count: Level 4 autonomy from one company, demand from a mobility platform, day-to-day fleet operations from a local operator. That three-party split is what makes the unit economics legible, and it is the template most Western cities will see first. It also puts Chinese-developed autonomy on European streets at scale while US import policy moves the opposite direction on Chinese robotics hardware.
The silicon under all of this is no longer a single-vendor market. AMD's Ryzen AI Embedded X100 series — up to 16 Zen 5 CPU cores, an integrated GPU with up to 40 RDNA 3.5 compute units, an XDNA 2 NPU rated to 50 TOPS, and unified memory in one SoC — began customer sampling in June, with production availability expected in the fourth quarter. AMD claims up to 3x higher peak FP32 performance than Nvidia's Jetson Thor T5000 and, more relevantly for a machine that has to hit a control loop on time, up to 3.4x better real-time task reliability on OpenNAV autonomous-robotics benchmarks. Deterministic timing, not raw throughput, is what separates a robot that works from a demo that works.
Quick Takes
Apple is training its own AI model for China with Alibaba's support, giving it more control over what it can offer in the market.
X open-sourced the ranking and filtering code behind its "For You" timeline on GitHub, paired with a pilot dashboard letting eligible accounts see visibility-limiting labels applied to their posts.
OpenAI's chief revenue officer Denise Dresser is leaving after about nine months, days behind former COO Brad Lightcap; Wiz president and COO Dali Rajic takes the role.
Mistral released OCR 4.1, a document-parsing model that outputs clean JSON or Markdown from complex tables and hierarchical layouts.
Google is rolling out Sheets canvas, a Gemini feature that turns spreadsheet data into interactive mini-apps that update with the sheet beneath them, available globally in English to AI Pro and Ultra subscribers.
NHTSA granted Zoox the first-ever commercial robotaxi exemption, allowing manufacture and deployment of vehicles that do not meet eight federal safety standards written around a human driver, paired with evolving "Operational Authorizations" rather than a blanket pass.
Lockheed Martin Skunk Works and the Air Force Test Pilot School flew 27 AI-controlled air intercepts on the X-62 VISTA using live infrared sensor data, closing the sensor-to-action loop in flight rather than simulation.
Cisco beat sales expectations on record enterprise demand for AI infrastructure, and Nvidia is reportedly testing lower-memory configurations of its Rubin Ultra accelerators as high-bandwidth memory shortages persist.
Samsung Foundry pushed its 1.4nm node to 2029, focusing instead on optimizing current 2nm gate-all-around production and advanced packaging.
Anthropic published research on how individually benign agent behaviors compound into systemic failures when many agents share an environment, naming confabulation and reward hacking among the mechanisms.
Taiwan says government agencies were targeted last month by an AI-driven hacking campaign, and security researchers report criminals have moved AI from experimentation into daily operational use.
Cursor now pre-builds cloud development environments in the background at no extra cost, which it says makes agents start up to 3x faster.
Japan's Supreme Court ruled that AI cannot be listed as a patent inventor.
Twitch will train Amazon's AI on streamers' content by default unless creators opt out.
What This Means for Your Business
Sort your AI workloads into "someone is waiting" and "nobody is waiting," and price them differently. That split did not matter much when speed was whatever your model happened to give you. With Ultrafast selling latency as a separate tier and DeepSeek moving to peak and off-peak rates on August 16, the same job now costs materially different amounts depending on when it runs and how fast it returns. Take an hour and list every AI task in your business in two columns. Anything in the customer-facing column may justify paying for speed. Anything in the batch column — report generation, data cleanup, transcript processing, overnight enrichment — should be scheduled into cheap hours and should never be paying a real-time premium. Most small companies are currently paying one blended rate for both.
Do not treat OpenAI's null result as permission to stop, but do treat it as a demand for measurement. The finding is not that AI does not work; it is that usage volume by itself predicts nothing about revenue per employee, while Census data shows the time savings are real — one to two hours per task for about a third of users. The gap between those two facts is entirely about what happens to the saved hour. Pick one team, pick one recurring task, and measure three things for a month: hours before, hours after, and what those hours got redeployed to. If the answer to the third is "nothing specific," you have found your problem, and it is not the model. The same report found executives are the lightest users in their own companies, which is worth a moment of honesty if you are the one setting AI policy.
Put the January price increase in your budget now. Gemini 3.7 Flash doubles on January 1 when its introductory rate expires — from $0.75 and $3.75 per million tokens to $1.50 and $7.50. If you are building a cost model on this quarter's usage, model both numbers and see whether the workload still makes sense at the higher one. More broadly, treat every headline rate cut as dated: assume any promotional price you are enjoying reverts, and know in advance which of your workloads would need to move if it did. The teams that get hurt by a rate change are the ones who never wrote down what they were paying.
Look at your orchestration before you look at your model. Writer's claim — roughly 40% average cost reduction from harness optimization alone, before any model change — points at the cheapest available savings in most AI deployments. Every unnecessary model call, every retry from a badly specified prompt, every step where a large model does work a small one could do is a recurring charge. Audit one workflow end to end and count the actual model calls it makes. In most homegrown setups the number is two to three times what the task requires, and fixing that is free, permanent, and carries over to whatever model you switch to next.
For anyone running a plant, warehouse, or fleet, the practical move this quarter is to ask suppliers what silicon their machines run on and when the second source arrives. AMD's Ryzen AI Embedded X100 reaches production availability in Q4 with credible real-time reliability claims against Nvidia's Jetson line, which means the single-vendor era in robot compute is ending — and pricing and lead times usually loosen a quarter or two after that becomes true. You do not need to have an opinion about which chip is better. You need your integrator to know you are aware there is a choice. On the deployment side, LG putting wheeled robots into a Tennessee plant for testing this year, rather than humanoids on a stage in 2027, is the right model for your own timeline: pick the constrained, repetitive, wheel-friendly task first and prove the economics there before anyone asks you about legs.