Four AI labs shipped models inside the same window, and the interesting number is not the benchmark score — it is what a finished task now costs. Elsewhere today: Anthropic is in talks to spend $6 billion on a company that makes chips run cheaper, a poisoned software package that was live for forty minutes turns out to have swept up credentials from more than 2,500 organizations, paid AI coding subscriptions are quietly becoming a hiring filter, and North American robot orders hit $622 million in a quarter where most of the buyers were not car companies.
The model labs synced calendars, and the fight is now over cost per finished task
SpaceXAI released Grok 4.6 on August 12, aimed squarely at long-running agent work rather than single-answer cleverness. Artificial Analysis scored it 61 on its Intelligence Index, up from 56 for Grok 4.5 and level with GPT-5.6 Sol. It leads on CursorBench at 69.9%, DeepSWE at 65.9%, and APEX-Agents at 57.5%. It is available now through Cursor, Grok Build, the API, OpenRouter, Vercel, and Cloudflare, with double the included usage for the first week. API pricing starts at $2 per million input tokens and $6 per million output.
The leaderboard placement is the least useful part. The number that matters for anyone running an agent all day is that Artificial Analysis measured Grok 4.6 at roughly $0.84 per completed task and placed it on its cost-performance frontier across every agentic evaluation in the index. On the AA-Briefcase evaluation, Grok reached top-tier quality while averaging about 53 turns and 0.5 billion input tokens, against roughly 103 turns and 2.0 billion for Claude Opus 5 Max. A model that reaches the same answer in half the turns is not marginally cheaper over an eight-hour job — it is a different line item. Elon Musk says Grok 4.7 is three to four weeks out and already meaningfully better.
The same week, DeepSeek began rolling out V4-Pro-0813 at $0.435 per million input tokens and $0.87 per million output — roughly a seventh of Grok's output price — claiming wins over Opus 4.8 on Terminal Bench 2.1, Cybergym, DeepSWE, and AutomationBench. Alibaba shipped Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model free to download, though its license carries a revenue tripwire aimed at companies that resell hosting or build coding assistants on it. Microsoft added MAI-Thinking-1 for cost-efficient enterprise workloads. Four releases, one theme: the frontier is no longer a single price point, and the spread between the cheapest credible model and the most expensive is now more than sevenfold for work that looks similar on a benchmark table.
Anthropic's $6 billion bid is for cheaper compute, not a better chatbot
Anthropic is in early talks to acquire Decart, an Israeli startup, for about $6 billion, according to Bloomberg and Reuters. It would be by far the largest acquisition the company has made. Decart builds world models that simulate physical environments and real-time generative video, but the piece Anthropic reportedly wants is the unglamorous one: software that makes chips work more efficiently and cuts the cost of training and serving models. Decart's team would join Anthropic's inference and performance group. The startup last raised at a valuation of nearly $4 billion, up from $3.1 billion in August 2025, with Nvidia among its backers. Talks are early and could still collapse.
Read it alongside Anthropic's own numbers. The company told investors it expected its first profitable quarter in Q2, with revenue roughly doubling to about $10.9 billion from $4.8 billion in Q1 and an operating profit near $559 million — while cautioning that staying profitable is uncertain because compute costs are scheduled to rise. Spending $6 billion to lower the cost of every future token answers exactly that problem. For customers, the price declines showing up on rate cards are not a promotional phase; they are vendors buying their way down the cost curve because their own margins depend on it.
A forty-minute compromise, and a bill that arrived five months later
Two malicious releases of LiteLLM, versions 1.82.7 and 1.82.8, sat on PyPI for about forty minutes on March 24 starting at 10:39 UTC. The project now advises treating any install up to 16:00 UTC that day as suspect. This week CloudSEK published the exposure map, and it is worse than the window suggests: roughly 434,000 captured files, mapping potential exposure to more than 2,500 organizations, with entries tied to Nvidia, Cisco, Deloitte, Volkswagen, FedEx, Siemens, and X Corp. The researchers are careful that this is captured loot, not confirmed compromise — appearing in the dataset does not prove anyone's credentials were used.
What the malware took is worth reading twice: environment variables, SSH keys, cloud credentials, Kubernetes tokens, database passwords, and model API keys including OpenAI and Anthropic keys. Both CloudSEK and LiteLLM say the correct response is to rotate credentials now rather than wait for evidence of misuse. Most teams cannot say from memory what their build pipeline installed in March — which is the point. Rotation is cheaper than finding out.
Two other security items landed in the same 48 hours. Researchers at A Security used a publicly available AI tool to find a Zoom screen-sharing flaw and build a working exploit in fewer than twenty prompts; the bug let anyone on a call silently take over other participants' devices with no interaction from the victim, and Zoom has fixes rolling out. And Israeli firm Dream Security documented a multi-agent system that autonomously mapped and compromised government entities in Asia over four days, running self-directed "Learning Cycles," adapting mid-operation, and expanding on its own from the primary targets to supply-chain vendors and energy companies — what Dream calls the first publicly known largely autonomous AI attack on a government target. Offensive capability is now cheap and fast; the defensive work is still manual.
The tools are becoming a hiring filter
LeadDev reported something small-company operators should sit with: fluency with paid agentic coding tools is quietly becoming a screening criterion. James Lowman, an engineering team lead at Starboard, asks candidates to describe the last three Claude Code skills they wrote. A college student interviewing for an internship could not answer — not from lack of ability, but because they could not afford anything past the free tier. That intern later took Lowman's part-time offer largely for the continued access to the paid tool. James Rowe, an engineering manager, spends $100 a month out of pocket on agentic coding tools to stay competitive while job hunting, and candidates have started negotiating token budgets alongside salary.
Matthew Sharp of the Oxford Martin AI Governance Initiative names the trap directly: "you're not just testing aptitude; you're also partly testing who can afford to practice." For a small business this cuts two ways. Screen on tool fluency and you are filtering on who had a budget, not who is good — and competing for that narrow pool against companies who can outbid you. Supply the tools instead and a $100-per-month seat becomes one of the cheapest hiring advantages available. OpenAI's own enterprise research points the same direction: the highest-usage firms now generate 8.3 times more output tokens per active user than typical enterprises, up from 2.6 times — a gap about access and workflow, not talent.
Elsewhere in models and money
Anthropic upgraded Claude in Chrome so the browser side panel runs a full Cowork session, with conversations saving to the account and resuming on desktop, web, or mobile, and existing skills and connectors working without setup. Vibe-coding startup Lovable raised $400 million at a $13.3 billion valuation, more than doubling its price since December. AI coding company Cognition entered talks at a $40 billion-plus valuation after reaching roughly $1 billion in annualized revenue. Lenovo posted record quarterly revenue of $26.9 billion, up 43% year over year, with AI-related revenue up 60% and now roughly 35% of the total — the hardware layer is compounding as fast as the software one.
On the applied side, China's Fengwu model called Typhoon Dolphin's mainland landfall five days ahead, missing by 30 kilometers and 30 minutes, for the most powerful typhoon to hit China this year. Its developers claim it beat Google's GraphCast across 80% of tested metrics — meaning it won more categories, not that it is 80% more accurate — and it still trails traditional systems on storm intensity. Five days of warning shows up as fewer ruined shipments and better-timed closures rather than as a product launch.
Physical AI
The most useful robotics number this week is not a humanoid demo. The Association for Advancing Automation reported that North American companies ordered 8,940 robots worth $622 million in the second quarter, up 4.3% in units and 21.3% in dollars over Q2 2025. First-half totals reached 17,995 units at $1.166 billion. The composition is the story: non-automotive customers accounted for 56% of all units ordered, the point at which general industry formally outgrew the car plants that built this market. Semiconductors, electronics, and photonics rose 38% year over year in the quarter; life sciences, pharmaceuticals, and biomedical were up 32% in units for the half; food and consumer goods rose 18%. Automotive OEM orders fell 25% in the first half, and the rest of the economy more than covered it.
Collaborative robots are where smaller operations are actually buying. Cobots accounted for 2,774 units and $114 million — 15.4% of units but only 9.8% of order value, the arithmetic of a machine costing a fraction of a caged industrial arm. In healthcare they made up 43.7% of robot orders and in electronics 36.5%. Tate Manufacturing deployed 58 Hirebotics cobot welders across multiple facilities, addressing exactly the skilled-trade vacancy most manufacturers cannot fill. Agility Robotics, meanwhile, says its Digit v5 humanoid will run more than 20 hours a day on fast charging when early units ship in December.
Drone operators should be paying closer attention than robot buyers. The FCC published Public Notice DA 26-758 in the Federal Register on August 3, proposing to expand its drone restrictions to cover LiDAR, thermal imaging, aerosol dispensing, swarms, and heavy airframes — and, critically, to apply it to models it already approved. That would pull the DJI Air 3S, Avata 360, and Mini 5 Pro from the US market despite existing FCC authorization. DJI calls it a total reversal, and it does break the pattern every prior action followed: cleared hardware stayed cleared. Drones already owned are unaffected, and the public comment deadline is September 2, 2026. If your business flies for surveying, roof inspection, agriculture, or real estate photography, the replacement question just moved from someday to this quarter.
Defense money continues to set the pace for the category's engineering. Cambridge Aerospace, founded in 2024 with about 250 people mostly in the UK, closed a $300 million Series C on August 10 at a $3.4 billion valuation, led by DFJ Growth with Lux Capital, Accel, Lakestar, and Elad Gil participating, bringing its total past $630 million. Its Skyhammer low-cost interceptor entered rapid deployment with the UK Armed Forces and Gulf partners in April, built to knock down cheap one-way attack drones. Hadrian raised $1.37 billion to scale US defense and aerospace manufacturing, and Moove raised $250 million for autonomous-vehicle fleet infrastructure. The counter-drone and automated-machining work funded here becomes affordable industrial equipment later, on the same lag that carried military GPS into every delivery van.
Quick Takes
Spotify is labeling AI-generated artist profiles with AI Persona badges and excluding them from algorithmic and editorial recommendations by default.
Amazon will train generative AI on Twitch streamers' content by default unless creators opt out.
Apple is discussing a nine-figure budget for multiyear news deals that would pay publishers when Siri's AI uses their content.
Google unveiled the Pixel 11 lineup at $899 and up, on a Tensor G6 chip, alongside DeepMind's SL2T model letting Deaf users sign ASL directly into the phone instead of typing.
A Windows zero-day called ShieldBreak turns Microsoft Defender into a privilege-escalation path to SYSTEM on fully patched Windows 11 25H2 and Server 2025; it requires local access and a user running a malicious app, and is unpatched.
Vantage Data Centers is exploring an IPO at roughly a $100 billion valuation, potentially raising around $10 billion, and legal AI company Legora is seeking funding above a $10 billion valuation, up from $5.6 billion four months ago on roughly $150 million in recurring revenue.
L&T won a contract worth up to $1.57 billion to deploy roughly 10,000 Nvidia B300 GPUs for Together AI in India.
Databricks acquired Electric, continuing the consolidation of data infrastructure underneath the agent stack.
Geoffrey Hinton, Fei-Fei Li, and Andrew Ng jointly argued for keeping AI open, warning that closure hands a small number of firms permanent control.
A new benchmark, AA-AnalystAgent, scores agents on real spreadsheet and document work with pass-all-five-attempts reliability as the headline metric; Claude Opus 5 leads at 54%, which is a fair picture of how often unattended analyst work actually completes.
What This Means for Your Business
Rotate your API keys this week, and stop treating that as a security-team chore. The LiteLLM disclosure is the clearest signal yet that AI tooling has become supply-chain surface area: a forty-minute window in March produced 434,000 captured files touching 2,500-plus organizations, and the stolen material specifically included OpenAI and Anthropic keys. You almost certainly cannot reconstruct what your build pipeline installed five months ago. Rotating model API keys, cloud credentials, and database passwords takes an afternoon and closes a question you otherwise cannot answer. Set a recurring rotation schedule while you are in there — quarterly is enough, and it converts this class of incident from an emergency into a non-event.
Start measuring your AI spend in cost per finished task, not per million tokens. The rate-card number is now actively misleading: Grok 4.6 costs three times DeepSeek's headline output price and can still be cheaper on a long job if it converges in half the turns, which is exactly what the AA-Briefcase results showed. Pick one repetitive workflow you already run — a weekly report, a batch of support triage, a data cleanup — and run it end to end on two models, timing it and recording total spend per completed run. That measurement takes an hour and will tell you more than every benchmark published this week. Re-run it quarterly, because with a sevenfold spread between credible options, the right answer keeps moving.
Buy the tools for your team rather than screening for who already has them. The hiring-filter reporting describes a real and fast-moving distortion: candidates are being evaluated on fluency with $100-a-month subscriptions, and an intern took a job partly to keep the tool. Small businesses lose that bidding war and should not enter it. Supplying seats instead flips it into an advantage — you get access to capable people who were filtered out elsewhere for reasons that have nothing to do with skill, at a cost that rounds to nothing against a salary. If you are already paying for seats, make sure people have time to actually practice; OpenAI's 8.3x usage gap between leading and typical firms is a workflow gap, not a talent one.
If you fly drones commercially, file a comment before September 2 and get a hardware plan on paper. The FCC's proposal would retroactively pull already-approved DJI models from the market, breaking the assumption that cleared equipment stays sellable. Your existing fleet keeps flying, so this is not a panic — but it does mean the replacement units you assumed you could buy next year may not exist, and the used market will price accordingly. Two concrete actions: inventory which of your airframes appear on the proposed list, and get a quote on a non-affected alternative now so you know the real number rather than a guess. The comment window is the only part with a deadline.
For anyone running a plant, shop, or warehouse, the robot order data says the entry point has moved. Non-automotive buyers are now the majority of the market and cobots are 15.4% of units at under 10% of dollars — meaning the machines being bought are cheaper, smaller, and aimed at specific unfilled jobs rather than at line-wide automation. Tate's 58 welding cobots are a template, not an outlier: pick the single role you have failed to hire for over the past year, and price a collaborative machine against the fully burdened cost of the position you cannot fill. That is a far more tractable question than whether humanoids are ready, and unlike the humanoid market, the equipment is available now and made by suppliers no import rule is currently threatening.