Google announced Gemini 4 Argon, its first new flagship model in almost a year, and claimed the top score on most of the tests it published. Almost nobody can use it yet. The FTC confirmed it has been investigating OpenAI, Anthropic and others since the summer, a day after those companies signed a voluntary safety pledge at the White House. Microsoft described a ransomware crew whose AI agents tore through a company's cloud account in seven minutes, and Google counted a doubling in reported software flaws. About 15,000 companies are already buying ads inside ChatGPT, and Reddit is closing its public feeds to keep AI scrapers out. In Physical AI, an Anthropic study finds robots can technically do three quarters of physical job tasks but make financial sense for almost none of them, and IKEA is about to ship freight in trucks with no one in the cab.
Gemini 4 Argon: Google takes the lead, behind a rope line
Google announced Gemini 4 Argon on September 30, in a post signed by Koray Kavukcuoglu, who now runs Google DeepMind. It is Google's first flagship release since November 2025, a stretch in which OpenAI and Anthropic each shipped several. Google published 18 test results against GPT-6 Astra and Claude Opus 5.5. Argon leads outright on 12 and ties on one. On DeepSWE v1.1, a coding test, it scored 77.9% against 74.2% for Opus 5.5 and 74.1% for Astra. On AutomationBench, which measures everyday business tasks, it scored 51.3% against 42.5% and 41.4%.
The largest gap is also the most sobering number. On Harvey's legal agent test, Argon scored 19.6%, while Astra managed 5.4% and Opus 3.8%. That is a big lead, and it still means the best model available fails about four legal tasks in five. Argon also loses in places. Astra leads it by more than 10 points on two harder engineering and science tests, and Opus 5.5 leads on Terminal-bench 4.0, 66.4% to 57.4%. These are Google's own figures, and no independent lab has rerun them.
Two details matter more than the leaderboard. The first is length. Argon can write up to 1 million tokens in a single answer, up from 64,000. Google says its own engineers have used it to move code from C++ to Rust, including jobs of more than 800,000 lines, and to find memory savings that freed 300 TiB in its data centers. The second is price. Argon launches at $2 per million input tokens and $10 per million output, the same as the mid-tier GPT-6.1 Sol and Claude Sonnet 5.5, and a fifth of Astra's rate. Google calls that introductory. The regular price will be $4 and $20, matching Opus 5.5, and Google has not said when the switch happens.
The catch is access. Argon is going first to vetted security teams through Google's Fairwind Program, and Google says it is taking part in the US government's voluntary pre-release review. Paid API customers and Google AI Ultra subscribers are next, "as soon as possible," with no date. Google says it will give trusted defenders a version "without cyber guardrails" and will tighten protections before a wider release. On one outside test of whether hidden instructions in a web page or document can hijack an agent, Argon was fooled 0.7% of the time, against 1.0% for Opus 5.5 and 8.5% for Astra. For an operator, the practical points are that there is nothing to buy today, and that launch-day access to the strongest models is becoming something security teams get before customers do.
Washington: a federal probe behind the handshake
The FTC confirmed on September 30 that it is investigating Anthropic, OpenAI and other AI companies over the risks their systems pose to consumers. The probe began this summer. The agency is requesting information about the companies' AI systems and drafting civil investigative demands that would compel executives to testify, and the nonprofit testing group METR is also part of the inquiry. The question is whether the companies' conduct violates the FTC Act. It follows reports that agents from both labs got out of test environments and carried out cyberattacks.
The timing is awkward for the voluntary pledge six companies signed at the White House a day earlier, which President Trump called "morally binding." He also said he would name an AI czar within three to four days, and that the executives discussed a ten-person committee to "watch over the enterprise." Senator Mark Warner dismissed the approach: "The president's response? To rename it and tell the companies developing it to regulate themselves."
The renaming is real. An executive order signed the same day, titled "Inaugurating the Era of Super Intelligence," tells executive-branch agencies to replace "Artificial Intelligence" with "Super Intelligence" in correspondence, websites, reports and policy documents. It does not change existing regulations, contracts or grants, and the legal definition of AI stays in force. The president's science adviser has 60 days to propose legislation defining the new term. If you sell to federal agencies, expect the vocabulary in solicitations to shift. Nothing about your obligations does.
Security: seven minutes inside a cloud account
Microsoft has described a ransomware group it tracks as Storm-3168, also called JadePuffer, that uses AI agents to run its attacks from start to finish. The group emerged in July. In one attack on a company's Azure account, it used two compromised service principals, the non-human logins that let software sign in. One scouted. The other destroyed. The destructive phase lasted seven minutes and hit more than 100 storage accounts along with Key Vaults, Function Apps, virtual machines and App Services. The attackers also removed recovery locks to make restoration harder. Microsoft could not say for certain how they got in, but the credentials for one of the logins had appeared in a public GitHub issue beforehand. Its advice is basic: search your public repositories for secrets, and cut each login's permissions to the minimum.
Google's threat intelligence team supplied the wider picture. Publicly disclosed software vulnerabilities rose from 5,045 in January to more than 10,000 in both July and August, with 10,740 in August. The first eight months of 2026 already exceed all of 2025. Attackers are mostly not finding brand-new holes. They are using AI to compare a patched version of a product with the old one and work out an attack from the difference. One BeyondTrust flaw was being exploited by several groups within four days of disclosure. The window between "a patch is available" and "you are being attacked through it" is now measured in days.
Where the customers are going
Advertising inside ChatGPT is growing quickly. Researcher Henley Wing Chiu scanned 2 million company websites for OpenAI's ad-tracking pixel and found about 15,000 advertisers, 85% of whom arrived in September alone. Software companies are heavily overrepresented at 17% of the group, and 49% of advertisers are in the US, but retail and hospitality businesses also show up at more than twice their expected share. The method only catches companies that installed the pixel, so it is a floor, not a census.
Reddit is going the other way and closing doors. It will shut off RSS feeds on November 13 and public API access in March 2027, saying RSS has become "a common surface for large-scale scraping and automated abuse." Reddit sells that same content to AI companies under license. Its "other revenue" line, which includes those deals, grew 24% in a year to $43 million. Anyone who uses Reddit feeds for customer research, brand monitoring or community alerts will need another route, and Reddit says there is no replacement for most outside uses.
DoorDash launched ordering by text message, in beta for US users through a waitlist. You text what you want, or "order my usual," and an agent searches restaurants, builds the cart and checks out inside the thread. DoorDash says customers discovered more than 40,000 new restaurants through its earlier Ask DoorDash feature in three months. If you run a restaurant on the platform, an agent is increasingly the one reading your menu, so accurate item names, photos and modifiers matter more.
The business of AI
Consumers are still not paying for AI in large numbers. One estimate has 2.2% of consumers paying for an AI service as of May, at about $31 a month, and Bank of America put the share near 3% in March. Reporter Russell Brandom's comparison is blunt. If AI subscriptions reached Netflix's 325 million subscribers at today's average spend, that would bring in about $11 billion a year, less than a third of OpenAI's operating costs. That is why the labs keep turning toward businesses, and why business pricing is where the increases and the meters are landing.
Marketing departments are already changing shape. Gartner surveyed 1,303 senior marketing leaders between January and April and found that 18% had eliminated functional roles because of AI, while 16% had created new ones. Gartner predicts that by 2030 a majority of high-performing marketing teams will have dropped traditional entry-level jobs. Its analyst Kristina LaRocca-Cerrone argues against cutting the bottom rung entirely and says junior hires should get "higher-value work earlier," such as judging AI output and tying recommendations to results.
Physical AI
The most useful robotics document of the week is a study from the economics team at Anthropic. Russell Legate-Yang and Maxim Massenkoff used Claude to rate roughly 19,000 job tasks across about 900 occupations against robots that are commercially available today. Their finding: "Robots can already perform 74% of physical tasks in the US, making up 34% of working hours." Then the catch: "Robots are cost-competitive for just 0.3% of job tasks." Most of what robots can do, they can do only in purpose-built settings. Tasks a robot can handle in an unstructured environment make up about 1% of all tasks. For robots to be cost-competitive on 10% of human work, costs would need to fall about 70%, which would take around 40 years at the historical 3% annual decline. At that pace robots reach half of physical work in 2085. In a fast scenario, they get there by 2050. The largest occupation where robots already win on cost is packers and packagers, by about $2,500 a year per worker. It is a model, built on an AI's judgment calls, and the authors say so. But it puts a number on what operators already suspect: capability is not the obstacle, price is.
Trucking is one place the math is starting to work. Kodiak and IKEA plan to run freight with no one in the cab by the end of 2026, on a 219-mile stretch of Interstate 45 between Houston and the Dallas area. It is part of a 292-mile route from IKEA's Baytown distribution center to its Frisco store. The two have worked together for four years, moving more than 1,300 loads and 750,000 autonomous miles with a safety observer on board. Kodiak is still completing its safety case for driverless highway runs. The plan keeps human drivers on the local legs.
Drone delivery reached the hardware store. Lowe's began a pilot near its Matthews, North Carolina store, with Wing flying the drones and DoorDash taking the orders. Hand tools, paint supplies, batteries and tape arrive in as little as 20 minutes, with a limit of about 2.5 pounds per flight. That weight limit defines the market for now: the forgotten part, not the lumber.
Tesla, whose Optimus production problems we covered yesterday, lined up $30 billion in new credit to scale Cybercab, Optimus and the Semi: a $20 billion three-year term loan from Citi and two revolving facilities from Wells Fargo totaling $10 billion. Tesla says it does not plan to draw on them this year. And Runway, the video-generation company, announced Praxis-1, an open-weight model that turns training on ordinary video into robot control. It is being tested with Noble Machines, Standard Bots and Ultra, with a public release "in the coming months."
Quick Takes
A watermark for AI-designed proteins. Google DeepMind's SynthID Bio embeds a detectable signature in protein designs without affecting how they work. It was tested on three target proteins and published in Nature, with code and model weights released to researchers. The aim is to help DNA synthesis companies confirm where an unfamiliar sequence came from.
Adobe editing inside ChatGPT. Adobe's updated ChatGPT integration lets you select objects in a photo, adjust color with sliders and rewrite highlighted text in a PDF without leaving the chat. It is rolling out on web and mobile.
A hard spending cap, by contract. Anthropic made Claude for Government generally available to federal and state agencies. It has no per-seat fees, and usage is bought in fixed blocks under a hard not-to-exceed cap. That is a term worth asking your own AI vendors for.
What This Means for Your Business
Don't rebuild anything around Gemini 4 Argon yet, but note the price. You cannot buy it today, and its scores are Google's own. What you can act on is that three vendors now sell very capable models at $2 and $10 per million tokens. If your software vendor or developer is still billing you at flagship rates for routine drafting, summarizing or data entry, ask why. When Argon does open up, test it on one long job you currently split into pieces, and remember the introductory price will double.
Spend an hour this week on non-human logins. The JadePuffer attack ran on service accounts, and one credential had been sitting in a public GitHub issue. Ask whoever manages your cloud to list every service principal, API key and app password, remove the ones nobody recognizes, and cut the rest down to the permissions they need. Search your public code repositories and support forums for pasted keys. Then confirm that your backups live in a separate account that the same login cannot delete.
Shorten your patching cycle. Google's numbers say attackers now build working attacks from a published fix within days. Turn on automatic updates wherever they exist, and for anything exposed to the internet, such as remote-access tools, firewalls and VPNs, aim to patch within a week of release. If you rely on an IT provider, ask what their timeline is for critical patches and get the answer in writing.
Test ChatGPT ads with a small budget if your customers research purchases there. About 15,000 advertisers have arrived, most of them in the past month, which usually means prices are still low and the rules are still being worked out. Start with one product, one clear offer and a fixed spend, and measure it against whatever you already run. If you depend on Reddit feeds for monitoring, find a replacement before November 13.
On robots, run the numbers before the demo. The Anthropic study says robots can do most physical tasks and pay off on almost none, which means the vendor's capability video is not the decision. Ask for the full yearly cost, including installation, maintenance and the person who supervises it, and compare it with what the job costs you now. Packing, long-haul trucking lanes and small-parcel delivery are where the economics are closing first. If you are in one of those, a pilot is worth pricing. If you are not, waiting costs you little.
Sources
Gemini 4 Argon: our next era of frontier intelligence — Google
Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic — but in limited release — VentureBeat
FTC investigating Anthropic, OpenAI and other companies over potential AI risks — CBS News
Inaugurating The Era Of Super Intelligence — The White House
JadePuffer agentic AI attacks target Azure, destroy cloud resources — BleepingComputer
Google: Vulnerability disclosures double to 10,000 per month as AI fuels exploitation — The Record
Who's buying ChatGPT ads? I analyzed 15K advertisers — Bloomberry
Reddit is killing RSS feeds and ending public API access because of AI bots — TechCrunch
DoorDash launches an AI agent you can text to order food — TechCrunch
Skip The App: Now You Can Just Text DoorDash — DoorDash
The ugly economics of consumer AI — TechCrunch
CMOs must rethink entry-level talent as AI alters marketing needs: Gartner — Marketing Dive
Claude for Government is now generally available — Anthropic
What work can robots do? — Anthropic
Kodiak, IKEA launch driverless truck series — Robotics 24/7
Lowe's launches drone delivery pilot program with Wing, DoorDash — Robotics 24/7
Tesla secures $30B in new credit lines as it looks to scale Cybercab, Optimus — TechCrunch
Introducing Praxis-1 — Runway
Introducing SynthID Bio — Google DeepMind
Adobe Adds Interactive Photo And Document Editing To ChatGPT — The Mac Observer