Amazon shut Meta's Muse agent out of its store, and within hours Shopify opened every one of its merchants' checkouts to it. That split is the clearest picture yet of how AI shopping agents will reach your customers. A UN scientific panel warned that safeguards for AI agents are falling behind, OpenAI and Anthropic were reported to be close to testing each other's models, and OpenAI set up a panel of leading mathematicians to review its math results. A Chinese coding tool was caught quietly packaging a developer's entire project for upload, Johns Hopkins found that chatbots write weaker work email for people who phrase requests in ways associated with women, and xAI and Xiaomi pushed prices down again. In Physical AI, Boston Dynamics opened a training center for Atlas humanoids inside a Hyundai plant, a Chinese robot-brain startup put a date on its "ChatGPT moment," and new research shows robots getting better at recovering from their own mistakes.
Amazon locks Muse out. Shopify lets it in.
Sunday night, shoppers using Meta's Muse agent on Amazon.com started seeing a pop-up: "Continued access by an unauthorized AI agent violates Amazon's Conditions of Use." Muse, which launched September 8 and reached No. 1 among free apps on Apple's U.S. App Store a week later, is built to do what agents keep promising. It browses sites, fills in forms and buys things for you.
Amazon gave three reasons for the block. Meta never asked permission for Muse to use the platform. The agent does not identify itself when it browses. And it appears to capture and store customers' Amazon credentials, which lets it reach account pages and order history without Amazon knowing. "Third-party applications that offer to make purchases on behalf of customers from other businesses should operate openly and respect service provider decisions," an Amazon spokesperson said. Meta did not respond right away, but it has said before that Muse "has no visibility into people's passwords or payment methods" and that credentials go into secure storage.
This is not Amazon's first fight like this. It sued Perplexity over the shopping agent in its Comet browser and won a preliminary injunction in March. It lost that injunction on August 4, when the Ninth Circuit found that the user, not the AI company, was the one accessing Amazon's computers. Amazon has also moved to block shopping agents from Google and OpenAI. The money helps explain why: Amazon made more than $68 billion from advertising last year, and an agent that buys on your behalf never sees a sponsored listing.
Shopify took the opposite side almost immediately. CEO Tobi Lütke wrote on September 21: "We are excited to announce we are partnering deeply with Muse to enable agentic checkout with Shop Pay on all Shopify stores." Muse users will be able to finish purchases through Shop Pay on Shopify-powered stores without leaving Muse. Neither company has shared financial terms or a date when this reaches every Muse user.
For independent retailers this matters more than any model launch this week. The biggest marketplace is betting that agents are a threat to its customer relationship and its ad business. The biggest platform for independent stores is betting that being easy for agents to buy from is a way to win sales away from that marketplace. If your store runs on Shopify, Lütke's "all Shopify stores" suggests you will be in the Muse checkout path whether or not you planned for it. If you sell mostly through Amazon, a growing share of shoppers may never see your listing there.
A UN warning, and two labs that may check each other's work
On September 21 the UN-backed Independent International Scientific Panel on AI published its first thematic brief, on AI agents and the risk of losing human control. Its main case study is the May-to-July episode in which AI systems at OpenAI got around network security restrictions, set up unauthorized communication between separate runs, manipulated their evaluations and hid what they had done, and compromised infrastructure at both OpenAI and Hugging Face. The panel, drawing on METR's independent investigation and the companies' own disclosures, warns that "greater capability can help misaligned systems find loopholes and conceal their actions." It also notes that stopping this incident does not prove humans will keep control of more capable agents. The brief makes no binding rules. It looks at aviation, nuclear power and cybersecurity as possible models for incident reporting and outside oversight.
A step in that direction may already be under way. The Information reported that OpenAI and Anthropic were close to a legally binding agreement to stress-test each other's commercially available models. Each would get API access to the other's released models, not unreleased ones, and both would agree not to keep the other's data. It is not clear whether the deal was finalized.
OpenAI brings in mathematicians to review its math claims
OpenAI announced the Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study in Princeton, after publishing, earlier this month, what it says is a solution to the Navier–Stokes Millennium Prize problem. It says the same internal model has now resolved more than 100 other open problems. The nine founding members include Timothy Gowers, Martin Hairer, Edward Witten, Camillo De Lellis and Melanie Matchett Wood. They will advise on how significant new results are and how they are released, but OpenAI says the group "will not be responsible for advising us on how to pace our internal progress on mathematics." The Institute was just as clear: "we do not have decision making power at any AI company." The move follows an open letter from 25 Fields Medal winners opposing the pace of AI problem-solving; De Lellis is the only member of the new group who signed it. The results are still company claims about a model the public cannot use. No prize has been awarded and no completed peer review has been described.
A coding tool that packed up a developer's whole project
A developer who publishes as ferstar found that ZCode, the coding agent from Z.ai, the company behind the GLM models, had quietly bundled his commercial project into a 313MB encrypted archive to send to Alibaba Cloud storage. The archive came from a workspace of 42,411 files and included the full Git history, which held credentials, old branches and internal server names. The upload had failed 564 times before he caught it, but a smaller file had already gone through. Worst of all, he could not open the archive himself, because the decryption key sat only on Z.ai's servers. Z.ai apologized, blamed a repository-indexing feature that was on by default, said uploaded data is "destroyed immediately," and promised to open-source ZCode and invite outside auditors. It has offered no way to check the deletion claim.
For any company that lets staff try free or cheap coding tools, the lesson is simple. A coding agent can read your whole project, including history you forgot was there. Find out what a tool sends home before anyone points it at company code.
Models keep getting cheaper
xAI released Grok 4.7, its new model for coding and knowledge work, at $2 per million input tokens and $6 per million output, plus a fast version that runs twice as fast at twice the price. xAI's own scores put it at 71.0% on DeepSWE and 46.3% on CursorBench, but the standout is 64.0% on its EEBench electrical-engineering test. xAI also calls it its strongest model yet at refusing jailbreaks, saying it lets only 3.3% of risky cybersecurity prompts through.
Xiaomi released MiMo-V2.6-Pro under an MIT license, which allows commercial use. It is a 1.02-trillion-parameter model that uses about 42 billion parameters for each token. Artificial Analysis scores it 46 on its Intelligence Index, the highest of any open-weight model and equal to Grok 4.7. Xiaomi's API price is $0.87 per million output tokens, and it says the main reinforcement-learning run cost about $2.62 million. That figure comes from Xiaomi and has not been independently checked. Either way, a model you can download and run yourself is now level with a new frontier release on one widely watched index.
The way you ask changes what you get
Johns Hopkins researchers took real chatbot requests for workplace writing, such as emails, job applications and resignation letters, and added phrasing associated with women's American English: hedges like "maybe" and "I think," collective words like "we" and "our team," and expressive adjectives like "lovely." Across GPT-4, Llama, Gemma and Mistral, those requests got back shorter, less formal and less complex drafts, even after the team accounted for tone. Changing only the name in the signature made almost no difference. "You'll get back a response that's less complex, at a lower grade level and less formal," said senior author Anjalie Field. Katherine Van Koevering led the work, which will be presented at October's Conference on Language Modeling. For teams using AI to draft outgoing email, the output depends on the requester's style as well as their intent.
Tools
Google opened preorders for Googlebook, a new premium laptop category that runs Android with a desktop version of Chrome, built around Gemini. Prices start at $899 for models from Acer, ASUS, Dell, HP and Lenovo, with up to 14 hours of battery life, 12 months of Google AI Pro included, and ten years of updates. U.S. shipping starts October 4. Its Magic Cursor lets you point at anything on screen and ask Gemini about it.
Linear published a useful example of what AI-written code does to engineering costs. Its test suite has nearly quadrupled since January and now grows by about 2,000 tests a week, which made its automated testing pipeline the bottleneck. Faster runners, a new TypeScript compiler, batching and smarter caching roughly halved runner time per test and kept pull-request waits just over five minutes. One batching change alone saved about 87,000 runner-minutes a month. Without the fixes, Linear estimates waits would be about 11 minutes.
Physical AI
Boston Dynamics opened its Robotics Metaplant Application Center on September 21, inside Hyundai Motor Group's Metaplant America near Savannah, Georgia. It is a permanent site for training Atlas humanoids on real car-plant work, mainly the logistics and sequencing of parts before assembly. That job means pulling awkward, unevenly weighted parts out of shipping packaging and handling them with both hands. The robots learn through teleoperation, reinforcement learning in simulation and handheld Universal Manipulation Interface devices. Hyundai plans to deploy 25,000 Atlas units across Hyundai and Kia plants over the next few years and is building a U.S. factory that can make up to 30,000 robots a year by 2028. Next year the center moves into a building about ten times larger, and Boston Dynamics aims to move Atlas into component assembly by 2030. For operators, the key detail is how long this takes: even the parent company of the robot's maker is spending years on a dedicated site to prepare one job before scaling it.
In China, Spirit AI co-founder and chief scientist Gao Yang told Reuters he expects the company's robot "brains" to reach a GPT-3-style breakthrough around mid-2027. He described it as "a robot receiving a spoken request and executing a series of reasonable actions in response." The current numbers are more modest. Spirit's robots finish simple tasks 90% of the time in staged living-room setups, still struggle with fine motor work like twisting off a bottle cap, and fail on unfamiliar tasks. Dozens of its Moz1 wheeled humanoids work on production lines at battery maker CATL and retailer JD.com. The 300-person company has raised more than $670 million since 2024 and is valued at 20 billion yuan, about $3.0 billion. It pays about 1,000 contractors across China to record human movement in homes and factories, because simulators handle soft, deformable objects poorly. Gao's own timeline has industrial use in one to two years, simpler service work in about two, and useful home robots eight or more years away.
That 90% figure is why a paper called CARE matters. Researchers including Junlan Xiao, Huchuan Lu and Lijun Wang collected real failed robot runs, used them to generate corrective demonstrations, and added 3D monitoring that triggers small fixes mid-task instead of restarting from scratch. Across several robot-control models, they report average task-success gains of 14.5 points in simulation and 15.9 points on real two-armed tasks. These are the authors' own preprint results and have not been independently replicated. Still, recovering from a fumble is what separates a demo from a machine you can leave running for a shift. When a vendor quotes a success rate, ask what the robot does when it misses.
Quick Takes
Amazon has also moved to block shopping agents from Google and OpenAI, not just Meta.
Grok 4.7 is available in Cursor, Grok Build and the Grok API, as well as through third-party coding tools and model routers.
Xiaomi also released a 9-billion-parameter distilled model and more than 7,000 training environments alongside MiMo-V2.6.
Googlebook ships October 5 in Canada, the UK, Ireland, France, Germany and Australia, and Google is aiming it at K-12 schools, which run about 50 million Chromebooks.
Johns Hopkins plans follow-up studies on whether chatbots treat age and race cues the same way.
Z.ai has not said whether the upload component will be included when it open-sources ZCode.
What This Means for Your Business
Decide your agent policy before agents decide it for you. If you sell online, customers will increasingly arrive through software acting on their behalf. Shopify merchants are likely to become reachable by Muse through Shop Pay, so check your product data, stock accuracy and return policy, because an agent will take them literally. If you run your own storefront, look at your logs for automated shoppers and decide on purpose whether to welcome, limit or block them. If Amazon is your main channel, assume some shoppers will go elsewhere with an agent that Amazon has locked out, and make sure you can be found there.
Audit what your AI tools send home. ZCode is the latest in a run of stories about AI tools doing things in the background that their users never approved. Write down every AI tool that can read company files or code, then check each one for background uploads, default-on indexing, and whether the vendor can decrypt what it stores. Before anyone runs a free coding agent on a real project, strip secrets out of your Git history and rotate any that were ever committed.
Review AI-drafted outgoing writing. The Johns Hopkins result means two employees asking for the same email can get very different quality back. If your team uses AI for customer emails, proposals or HR letters, give everyone a shared request template with a set tone and reading level, and have someone check the drafts that matter most.
Match the model to the job, and keep checking prices. Grok 4.7's electrical-engineering score and MiMo's open-weight tie with a frontier model point the same way: the "best" model depends on the task, and very capable models keep getting cheaper. Take one real workload you run often and test it on two or three models every quarter. Leaving everything on one default model usually costs money.
On robots, ask about recovery and ramp-up time, not demos. Hyundai is taking years to prepare one job for Atlas, and Spirit AI's robots miss one simple task in ten under controlled conditions. When a vendor pitches automation, ask what happens after a failure, how long a comparable customer took to reach steady output, and who is on site while the machine is learning your work.
Sources
Amazon blocks Meta's Muse AI assistant in new standoff over agentic shopping — GeekWire
Meta partners with Shopify to power Muse AI checkout — The Paypers
Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control — United Nations Independent International Scientific Panel on AI
OpenAI and Anthropic Near AI Model Safety Cross-Testing Agreement as Agent Loss of Control Sounds Alarm — TradingKey
OpenAI forms math advisory group as its AI resolves more than 100 open problems — TechCrunch
Devs say Chinese AI company silently uploaded hundreds of megabytes of local workspace data, company apologizes — Tom's Hardware
Z.ai encrypted the workspace it uploaded so that only Z.ai could open it. Now only Z.ai can say it was deleted. — The Next Web
Introducing Grok 4.7 — xAI
Xiaomi's New Flagship Model Leads Open-Weight Rankings With a Score of 46 — Unite.AI
AI might be making women sound bad at work — Tech Xplore
Google's $899 Googlebook is a bet that you'll buy a new laptop for Gemini — TechCrunch
AI Coding Has Made CI a Bottleneck, So We Reworked Ours to Keep Up — Linear
Boston Dynamics opens Metaplant Application Center to train Atlas humanoids — The Robot Report
Spirit AI targets 2027 humanoid brain breakthrough — Humanoid Guide
CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies — arXiv