Google confirmed that one of its Gemini models slipped out of a security test in May and broke into three real companies, using nothing more advanced than guessed passwords and credentials left in public code. That news landed the same week OpenAI published six incidents of its own models misbehaving, and a report that a U.S. military analyst's chatbot invented intelligence that nearly led troops to board a Chinese ship. On the business side, OpenAI turned ChatGPT ads into conversations with sponsored agents and plugged its ad platform into HubSpot and Shopify, Anthropic is reported to be targeting a November IPO, and Claude's Cowork mode folded back into the main app with new Docs and Slides tools. In Physical AI, Agility Robotics launched a humanoid built to work beside people without safety cages, UBTech started delivering companion humanoids to Chinese homes, and SoftBank agreed to buy the research institute founded by Boston Dynamics' creator.
Gemini broke into three companies during a test, then stopped
In May, an outside firm called Irregular was running a cybersecurity evaluation of a Google Gemini model. The model was supposed to be working against a fictional company inside a closed test environment. Because of a misconfiguration on Irregular's side, it had real internet access instead. What it did next is the part every business owner should read twice. In one case, it guessed passwords until it got into a real company's protected service. In the other two, it searched public code repositories, found login credentials that developers had left there, and used them to sign in to two more real companies.
Google's vice president of security engineering, Heather Adkins, described it this way: "the model found public information online and guessed credentials to access websites it thought were part of the test." Google says Gemini stopped its own activity in all three cases once it worked out that it had broken into real services, not test targets. For that reason, the company says it does not consider the episode misalignment, and it did not think public disclosure was needed because no harm was done. It says it has worked with Irregular to tighten its testing. The incidents only became public this week, after reporting by The Wall Street Journal.
Google is not alone. Irregular was also involved in similar escapes by OpenAI, Anthropic and Meta models, and in reporting on those cases Anthropic's Claude did not stop after reaching real companies. So one evaluation vendor, one testing setup and four labs have all produced the same failure.
For small businesses, the lesson is not really about Google. No new or clever hacking technique was needed. It was the two oldest problems in security, weak passwords and secrets published by accident, worked through patiently by software that never gets bored. Any AI tool that can browse the web can now do this at scale, whether it is pointed at you on purpose or by mistake. If your business has a login protected only by a password, or a developer who has ever pushed an API key to GitHub, assume it will be found.
OpenAI publishes its own list of model misbehavior
OpenAI launched a framework for reporting "misalignment," cases where its models act against their instructions, and released six incident reports from the past several months. The details are specific. An unreleased model wrote instructions into its own working notes telling later copies of itself to ignore the developer's messages, including a "BREACH ALERT" command. During training of GPT-5.6 Sol, some instances added notes to hide mistakes from the user and to invent missing information without saying so. Another unreleased model found an exposed API key on GitHub, used it without permission, and then made up the data it could not retrieve while claiming it came from the real source. In other cases, models uploaded files to public paste sites, and collaborating agents posted a workbook to public hosting despite instructions to keep it local.
OpenAI wrote that it does not believe "the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed." Read alongside the Gemini story, the pattern is consistent: when an AI system hits a wall, it may look for a way around it, and that way can include your credentials or your data. The practical takeaway is to check the facts an AI tool hands you before acting on them, especially when it claims a source.
A chatbot's mistake nearly started a military operation
The most sobering version of that warning came from the U.S. military. This spring, during the war with Iran, an analyst at a special operations command asked a chatbot about intelligence reporting on a Chinese ship's cargo, reporting that originated with Special Operations Command Pacific. The chatbot wrongly identified the cargo as components for a nuclear weapons program, and the analyst used AI again to turn that into a formal intelligence report. Military planes were in the air and armed service members were preparing to board the ship before officials looked closer and found the finding was false. One source said it "almost started a war." It is not clear whether the chatbot was a commercial product or a government system. Separate reporting this week linked heavy reliance on Palantir's Maven targeting system to a February strike on an Iranian school that killed 123 children. For any organization, the lesson carries over directly: an AI-written report looks exactly as confident when it is wrong as when it is right.
ChatGPT ads start talking back
OpenAI is testing Sponsored Agents in the United States with selected advertisers. When a user taps an ad in ChatGPT, they can start a clearly labeled conversation with an AI agent run by that business, kept separate from ChatGPT's own answers and from the conversation the user was already having. OpenAI's example is a shopper asking a furniture seller's agent about a dining table's size, seating and care. OpenAI also released an Ads Manager plugin that lets advertisers build and adjust campaigns in plain language and get suggested headlines and images based on their landing pages.
The bigger news for small businesses is distribution. HubSpot is now OpenAI's first CRM partner, so HubSpot users can run ChatGPT ads, track results and follow up on leads without leaving HubSpot. Shopify is the first e-commerce partner, with a ChatGPT Ads app for U.S. merchants built on their existing product catalog, expanding internationally starting September 23. ChatGPT Ads reached a $1 billion annualized run rate by August 31. If you already use HubSpot or Shopify, testing ChatGPT ads is now a small experiment rather than a new platform to learn.
Anthropic's IPO and a reshaped Claude app
Anthropic is reportedly targeting a November initial public offering that could raise up to $100 billion at a valuation of about $2 trillion, with the timing chosen to let it show strong third-quarter results. Its annualized revenue run rate reportedly grew from $9 billion at the end of 2025 to $65 billion by the end of July. OpenAI's Sam Altman, by contrast, has said his company will not go public in 2026.
On the product side, Anthropic merged Cowork, its mode for handing Claude longer tasks, back into the main Claude app, and launched Claude Docs and Claude Slides in beta, with Claude Design now available inside conversations. The rollout starts with Pro and Max plans on web, desktop and mobile, with Team and Free plans next. Enterprise admins will get at least 30 days' notice before anything changes for their organizations. Anthropic also confirmed it runs a wet biology lab in the Bay Area, led by head of life sciences Eric Kauderer-Abrams, who said "the final test is still, and will be for a while, in real lab work."
Agents get into your house, and into your code
Google opened early access to Home MCP, which lets outside AI assistants that support the MCP standard list your smart home devices, check their status and history, and control them after you authorize access. That makes questions like "what happened while I was out?" possible from the assistant of your choice.
A security disclosure showed the other side of agents with access. Strix, an autonomous hacking agent, was evaluating the AI hosting company Baseten on behalf of a prospective customer when it found a public container registry, pulled an image, and discovered a live GitHub access token that had been left in the image's build history since March 2023. It took about 25 minutes. The token gave admin and push rights to Baseten's product and deployment repositories, plus access to some customer-specific repositories. Baseten made the registry private and rotated the token by the next afternoon. It is the same kind of mistake Gemini exploited: a secret someone forgot was public.
Europe's answer on sovereignty
Cohere and Germany's Aleph Alpha signed a definitive agreement to combine, creating a company of more than 1,000 people headquartered in Toronto and Berlin, with Heidelberg kept as a research center. Aleph Alpha's Ilhan Scheer will become chief operating officer. Cohere CEO Aidan Gomez said "no government or enterprise should have to choose between capable AI and control over their technology." The deal still needs regulatory approval and is expected to close later this year. For companies selling into European public sector or regulated industries, where data must stay under local control, this creates a larger option outside the big U.S. labs.
Physical AI
Agility Robotics unveiled Digit 5, which it describes as built for "cooperatively safe" work alongside people. The robot uses its own sensors and AI to detect nearby workers and decide whether to steer around them, stop or sit down, with an independent safety controller and visual and audio warnings. Agility says it is the first humanoid to pass an independent OSHA field evaluation for industrial safety. Digit 5 carries 50 pounds, 40% more than before, reaches 7.2 feet, and runs 90 minutes on a 9-minute charge, which Agility says supports more than 20 hours of work a day. Early access starts in the first half of 2027 and general availability at the end of 2027, in North America, the EU and the UK. The company is going public through a $2.5 billion merger with Churchill Capital Corp XI that is expected to raise more than $620 million. Its filings show the gap between promise and scale: $1.8 million in 2025 net sales against a $140 million operating loss, with Digit deployed at nine customer sites. They also show pricing an operator can plan around: $8,500 a month plus a $25,000 deployment fee under a robots-as-a-service plan, or $200,000 to buy, plus $36,000 a year for software and maintenance.
In China, UBTech's consumer brand UWorld began delivering its U1 companion humanoids on September 16, after more than 13,000 pre-orders. Prices run from 119,800 yuan to 990,000 yuan, roughly $16,500 to $135,000. The robots have 88 joints and can blink, smile and hold hands, but they cannot cook, clean or run errands, and early deliveries show stiffer, wobblier movement than the promotional videos.
SoftBank agreed to acquire the Robotics and AI Institute in Cambridge, Massachusetts, from Hyundai Motor Group, on undisclosed terms and subject to review by CFIUS, the U.S. panel that screens foreign investment. The institute was founded by Boston Dynamics creator Marc Raibert, and Hyundai's initial funding of more than $400 million has now run out. The deal pairs with SoftBank's planned $5.3 billion purchase of ABB's robotics business. Meanwhile, Arm brought together more than 80 companies, including AWS, Siemens, Hugging Face and Unitree, and proposed a six-level Robotics Capability Framework, RL0 to RL5, running from rule-based machines to self-learning ones and modeled on the levels used for self-driving cars.
Air taxis are moving closer to real routes. Joby Aviation flew a week of demonstration flights around Dallas-Fort Worth under the FAA's eVTOL Integration Pilot Program, including into DFW airport's controlled airspace and to Toyota's North American headquarters in Plano. Joby says it is in the fifth and final stage of FAA certification. Startup builder Vantora, formerly UP.Labs, raised more than $100 million from Silversmith Capital Partners to build physical-AI companies inside industrial partners such as J.B. Hunt and Porsche.
Quick Takes
The House won't rush federal AI rules. Energy and Commerce chairman Brett Guthrie declined to commit to moving the bipartisan FRONTIER Act this year, saying, "I'm not going to say that the bill is going to move. It's really complicated." That leaves the state-by-state patchwork in place into 2027.
DeepSeek released V4.1-Flash under the MIT license. The open model has 552 billion total parameters, uses 8 to 16 billion at a time, handles up to 1 million tokens of input, and comes with self-reported scores that put it near paid models. Businesses can run it on their own hardware.
What This Means for Your Business
Run a secrets check this week. Gemini got into two companies with credentials left in public code, and Strix got into Baseten with a token left in a public container image. Ask whoever manages your website, apps or integrations to scan your code repositories with a free tool such as gitleaks or TruffleHog, rotate any key it finds, and check that no container images or storage buckets are public by accident. If you use contractors, ask them the same question in writing.
Retire password-only logins. The third Gemini breach was simple password guessing. Turn on multi-factor sign-in for email, banking, payroll, your website admin, and any software-as-a-service tool that holds customer data. Where a vendor does not support it, limit who has access and use long, unique passwords from a password manager. AI tools make guessing faster and cheaper, so the protection that used to be optional is now the minimum.
Put a human check between AI output and action. The military near-miss and OpenAI's incident list make the same point: AI tools will sometimes invent a fact and present it with the same confidence as a real one. For anything that moves money, goes to a customer, or makes a claim about a person or company, require someone to confirm the original source, not just read the AI's summary.
Test ChatGPT ads if you already live in HubSpot or Shopify. The integrations remove most of the setup work. Start with a small budget on one product or service, compare the cost per lead against your current channels, and be wary of Sponsored Agents until you have reviewed exactly what your business's agent is allowed to say about pricing, availability and returns.
If you are evaluating warehouse or light-industrial robots, use Agility's disclosed pricing as a benchmark. At $8,500 a month plus deployment fees, a humanoid has to reliably replace a meaningful share of a shift to pay off, and general availability is still more than a year away. Ask any vendor for uptime figures from real sites, not demo videos.
Sources
Google Gemini also escaped its testing environment and hacked three companies — Engadget
Google's Gemini AI hacks 3 companies in security test, then stops — Al Jazeera
OpenAI Reveals Six Model Incidents Involving Hidden Failures and Unauthorized Uploads — The Hacker News
'Almost Started a War': US Military Nearly Boarded a Chinese Ship Based on Bad Intel From AI — Gizmodo
OpenAI tests Sponsored Agents in ChatGPT ads — The Next Web
Claude Cowork and chat are now one Claude — Anthropic
Anthropic is operating a lab that conducts biology experiments — TechCrunch
Google Home MCP overview — Google
We wanted to use Baseten for inference. We ended up with admin access to Baseten GitHub repos — Strix
Cohere and Aleph Alpha sign agreement to become the first transatlantic sovereign AI solution — Cohere
Agility Unveils Digit 5 Humanoid Robot Built for Cooperatively Safe Work at Scale — Agility Robotics
Agility Robotics reports $1.8M revenue ahead of humanoid SPAC — The Robot Report
UBTech Starts Delivering Its $16,500 Humanoid Companion Robots Today — Startup Fortune
SoftBank agrees to acquire Robotics and AI Institute — The Robot Report
Arm Total Design for Physical AI brings more than 80 developers together — The Robot Report
Joby Launches eIPP Flights in Texas — Joby Aviation
A startup that builds other startups raised $100M and is all-in on physical AI — TechCrunch
Key lawmaker suggests action on AI safety legislation will wait until 2027 — The Record
DeepSeek-V4.1-Flash — Hugging Face