Hot Chips 2026 turned into the most consequential hardware week of the year. OpenAI put real silicon on the table with Jalapeño, its first custom accelerator, and claimed it beats Nvidia's flagship systems on efficiency. Apple shipped its first 2nm chip in an actual product, Perplexity moved a working agent entirely onto a local machine, and two robotics companies attacked the same bottleneck from opposite directions. Meanwhile Pew put a number on something most businesses have not priced in: a third of American adults are now asking chatbots about their health.
OpenAI built its own chip, and the numbers are real
OpenAI disclosed detailed results for Jalapeño, the inference accelerator it co-designed with Broadcom, and the headline claim is that it beats Nvidia's GB200 and GB300 systems on efficiency. Across three public models — GPT-OSS, DeepSeek R1, and Kimi K2.5 — OpenAI reports 1.5x to 1.9x more work per watt and a 1.7x to 3.6x reduction in response latency. On GPT-OSS the chip reached nearly double the GB200's highest throughput point, and roughly 1,400 tokens per second per user at low concurrency.
The engineering underneath is substantial. Jalapeño is built on TSMC's N3P process for the compute die with an N3E I/O chiplet, draws 700 watts, and pairs with HBM4 memory delivering 15.4 TB/s of bandwidth. A single rack holds 128 of the chips across 16 trays, scaling to 2,048 across 16 racks using a hybrid copper and optical interconnect. Design work began in mid-2024 and taped out in November 2025 — roughly a 16-month cycle, which is fast for custom silicon. A second stepping already in the fab is expected to add about 25% more performance per watt.
The caveats matter as much as the claims, and OpenAI's own analysts flagged them. Every performance number was supplied by OpenAI. Testing used single-token prediction without speculative decoding, which is a friendlier workload than production. Most importantly, no results exist yet for multi-turn, long-context agent workloads — the pattern that actually describes how businesses use these systems. Production ramp is scheduled across 2027.
For operators, the significance is not that a new chip exists. It is that the largest buyer of AI compute in the world just demonstrated it can build its own and does not have to keep paying Nvidia's margin. That is the kind of pressure that shows up months later as lower prices on the API products small companies actually buy.
The rest of the silicon field moved at the same time
Jalapeño did not arrive in isolation. d-Matrix detailed Raptor, which stacks an N4 logic die directly on 3D DRAM to hold bandwidth near 100 TB/s; the company claims a 72-card rack could serve a three-trillion-parameter-class model at one-million-token context at roughly 1,000 tokens per second per user, though the Hot Chips report was explicit that working-system evidence is still missing. Samsung outlined "zHBM," which mounts DRAM stacks directly onto the processor die via hybrid copper bonding, removing power-hungry interface layers and cutting roughly 100 watts from a 1,200-watt accelerator.
IBM filled in the details on the dual-architecture mainframe processor it teased earlier in the week: fabricated at 2nm, 11 high-performance cores clocked above 5.7 GHz, with each core switching between IBM z/Architecture and Armv9.3-A execution within nanoseconds rather than splitting work across separate physical cores. AMD revealed its Instinct MI455X accelerator and 72-GPU Helios rack platform, and Intel previewed architectural changes in Diamond Rapids, Crescent Island, and Wildcat Lake aimed at agent workloads. The pattern across all of it is the same: the industry has stopped trying to build one chip that does everything and started building different silicon for training, for generation, and for memory-bound work.
Apple ships 2nm in a product you can order
Apple unveiled the M6, its first chip on a 2nm process, alongside the M5 Ultra — its first quad-die M-series design, supporting up to 512GB of unified memory. Apple claims the M5 Ultra delivers 4.5x the peak GPU AI compute of the M3 Ultra. Critically, both are shipping in new Mac mini and Mac Studio machines rather than sitting on a roadmap.
That combination — very large unified memory and a big jump in on-device AI compute, in a desktop that starts at Mac mini pricing — is what makes local model execution practical for a small business rather than a hobby. Separately, leaked photos surfaced of Apple's Private Cloud Compute servers for the first time: a 2U rack unit packing 32 Apple Silicon boards in four columns of eight, cooled by one fan per column. Forum analysis of the board dimensions suggests these are base M5-class packages rather than the M5 Ultra chips previously rumored.
Perplexity puts the whole agent on your machine
Perplexity shipped Portable Computer, an agent that runs entirely locally on Nvidia's DGX Spark using either a 27-billion-parameter Qwen model or Perplexity's own, keeping files and routine work on the device and consuming no credits. It escalates to the cloud only when explicitly authorized. Support for RTX Linux PCs is announced but not yet shipping.
This is the clearest expression yet of a shift worth tracking: cost, latency, and privacy are becoming a single optimization problem rather than three separate ones. For any business that has held back on AI because client data cannot leave the building — law firms, medical practices, accounting shops, anyone under a data-residency clause — local execution changes what is possible without changing what is permitted.
A third of American adults are asking chatbots about their health
Pew Research surveyed 3,488 US adults between June 22 and 28 and found 34% had used AI chatbots for at least one health or medical reason. Twenty-eight percent used them to fetch quick health information, 25% to understand what was causing symptoms, and 22% to get low-cost information. Among users, 47% called the answers extremely or very helpful — but only 29% said they were very comfortable sharing personal health data with the tools. Adoption skews sharply by group: 56% among Asian Americans and 44% among adults under 30.
The gap between "helpful" and "comfortable" is the finding to sit with. People are using these tools for consequential questions while remaining uneasy about what happens to what they type. Any business handling sensitive customer information should read that as a live expectation about disclosure, not a distant policy question. Note also what the survey measures — adoption and perception, not whether the medical answers were correct.
Anthropic makes memory the default
Anthropic unified Claude's memory across chat and Cowork, so both now write to the same store, and turned it on by default for Free, Pro, and Max users. Users can edit or delete individual topic files, and sensitive topics stay off unless explicitly enabled. The practical consequence is that a detail mentioned in casual conversation can now shape later work executed in the cloud — useful, and worth a policy conversation before your team starts pasting client details into it.
Physical AI
Two companies attacked robotics' central bottleneck this week from opposite ends, and the contrast is instructive. Skild AI unveiled S1, a foundation model built specifically for in-context learning: show it a single video of a task it has never seen, and it executes without any fine-tuning or post-training. The company demonstrated it potting a plant, brewing pour-over coffee, flipping pancakes, and assembling a kit. What makes S1 notable is the horizon — it handles tasks up to 10 minutes long, far beyond the short manipulations that similar demos have shown. Skild reports a 66% success rate on unseen tasks in internal testing against 9% for a language-prompted vision-language-action model trained on the same 100,000 hours, and estimates one video prompt is worth roughly 380 task-specific demonstrations. Independent replication is still missing, and that caveat is doing real work here.
Figure AI came at the same problem with money and crowds rather than architecture. It took Index out of stealth — a crowdsourced data platform that began quietly last September as "Project Go-Big" and has now collected more than 16 million video uploads from over 264,000 app downloads across more than 100 countries, with about 44,000 weekly active users and roughly 30 minutes of footage arriving every second. Figure has paid contributors $15 million so far to film themselves doing ordinary tasks like opening drawers and folding laundry, and has committed more than $1 billion to data and compute over the next 12 months. Submissions run through a five-stage pipeline of filtering, fraud review, deduplication, rebalancing, and annotation before training the company's Helix system.
The bet behind both is the same: language models could train on the text of the internet, but robots need physical demonstrations that simply do not exist online at scale. One company is trying to make each demonstration count for 380; the other is buying demonstrations by the million. Whichever approach wins, the cost of teaching a robot a new task is what determines whether small operators ever get access, and both are pushing it down.
There is an uncomfortable bookend to this. Amazon told workers and requesters that Mechanical Turk, the 20-year-old microtask platform that supplied labels, surveys, and evaluations across two decades of machine learning, will shut down on September 30. The hidden human labor behind AI is not disappearing — Figure is paying $15 million for it. It is just moving from typing to performing, and from a marketplace into a proprietary pipeline.
Quick Takes
Elon Musk said SpaceX aims to launch its first fleet of orbital AI data centers near the end of 2027, running on Nvidia hardware — a firmer date on the plan disclosed earlier this week.
OpenAI banned a Russian-origin ChatGPT cluster behind a fake Israeli expert community called the "International Burke Institute." In OpenAI's sample, 34 of 36 institute articles had been copied to other platforms. Reach was small; the method is the warning.
US gas-fired power capacity under construction jumped 76% in six months and now runs at roughly double China's, according to Global Energy Monitor. About half the wider US pipeline is tied to data centers.
Smart ring maker Oura is reportedly eyeing a September IPO, raising up to $3 billion at a valuation above $16 billion.
SK hynix and Intel Foundry will package next-generation HBM using Intel's EMIB-T bridge technology, easing an industry-wide advanced-packaging bottleneck.
GPT-5.6 became available in AWS GovCloud, opening the model to US public-sector workloads with stricter compliance requirements.
What This Means for Your Business
Do not buy hardware on this news, but do renegotiate on it. Jalapeño, Raptor, the AMD and Intel roadmaps, and Nvidia's own segmentation all point the same direction: inference is getting cheaper and faster, and the largest buyer of compute now has a credible in-house alternative. That pressure reaches you as falling API prices and faster response tiers, usually a quarter or two later. If you are being asked to sign a multi-year AI commitment at today's rates, the correct posture is short terms and the right to re-price.
Take local execution seriously if data residency has been your blocker. Between Apple's M5 Ultra with up to 512GB of unified memory shipping in a desktop and Perplexity running a full agent on a DGX Spark with no cloud calls, "the data never leaves the office" stopped being a theoretical architecture this week. For firms under confidentiality obligations that have kept them out of AI entirely, a single machine is now a legitimate pilot. Scope it as one workflow, measure it against what you would have paid in cloud credits, and check whether the smaller local model is actually good enough for that specific job — often it is.
Write a health-information policy if customers can talk to your systems. A third of American adults are already asking chatbots medical questions, and fewer than a third are comfortable with where that data goes. If your business touches health, wellness, benefits, insurance, or even employee assistance, assume people will type sensitive things into any chat interface you deploy. Decide now what gets logged, what gets retained, and what your assistant refuses to answer — before an incident makes the decision for you.
Set a memory policy before your team sets one by accident. Claude's memory is now on by default and shared across chat and cloud-executed work. That is genuinely useful and also means an offhand client detail can persist into later output. Spend fifteen minutes deciding what your team may put into a persistent assistant, and make sure whoever handles regulated or client-confidential material knows how to turn it off and delete topic files.
On robotics, watch the cost of teaching, not the cost of the robot. Skild is trying to make one demonstration replace 380; Figure is spending a billion dollars to buy demonstrations at scale. The hardware price gets the headlines, but the reason robots are not yet in small businesses is that every new task costs real money to teach. When that number falls far enough that a task can be taught by filming it once, the addressable market changes overnight — and that, not the next humanoid demo reel, is the signal to act on.