Agents of Work
Let's Talk
September 16, 2026 · Agents of Work

Agents of Work AI Daily Briefing — September 16, 2026

Salesforce put its sales tools inside Anthropic's Claude and, the same day, launched its own reasoning model with Nvidia. OpenAI confirmed weeks of safety talks with Anthropic and Google DeepMind, and investors are reportedly lining up to value it at $1.2 trillion even as its IPO slips to 2027. A startup founded by a ChatGPT researcher launched a model that makes decisions rather than writing text, Google shipped voice models that talk while they work, Meta turned its apps into a subscription business, and a dispute over an AI price tracker pulled Washington into the market for rented computing power. In Physical AI, Waymo opened driverless rides in Las Vegas, and Odyssey and Nvidia released new tools for training robots.

Salesforce hedges its bets: Claude on one side, its own model on the other

Salesforce launched Salesforce in Claude on September 15, a plugin that brings a seller's accounts, opportunities and pipeline into Anthropic's assistant. It ships with 37 skills that cover much of an account executive's day: a morning briefing on meetings, deals closing and at-risk opportunities; call preparation with stakeholder and deal history; deal scoring and close plans; post-call follow-up; pipeline review and forecasting; and drafting updates to Salesforce records. It is in beta on all paid Claude plans. Sellers sign in with their own Salesforce credentials, so Claude sees only what their existing permissions allow, and by default Claude asks for approval before writing any change back to the CRM. A companion Slack connector pulls in deal channels and account threads. Anthropic says data is not used for model training on Team and Enterprise plans. Salesforce says 7,000 of its own sellers already use the plugin, and it is deployed at GitLab, Siemens and Legora. "Sellers can start the day with pipeline review already done and account history there," said Alexa Vignone, Salesforce's president and chief revenue officer.

On the same day, Salesforce and Nvidia introduced Koa, a reasoning model built on Nvidia's open-weight Nemotron and further trained for sales, marketing and customer-support work inside Agentforce. Salesforce says Koa was trained only on synthetic data, not customer records. The two companies simulated a customer-service environment, complete with irate customers, to generate it. The model is designed to use less computing per task, which lowers the cost of multi-step jobs compared with sending every request to a frontier model. "Reasoning has always been something that we've relied on the frontier model providers for. Until now," said Jayesh Govindarajan, Salesforce's executive vice president for AI.

Put together, the two launches show where business software is heading. The big CRM vendor is not choosing between the major labs and its own model. It is using both: a cheaper model it controls for high-volume routine work, and a frontier assistant where people want a smarter partner. For a small business on Salesforce, the practical change is that the AI features you pay for may run on different models depending on the task, and the approval-before-writing default is worth keeping on.

The labs start coordinating on safety

OpenAI confirmed on September 15 that it has been working with Anthropic and Google DeepMind on AI safety for several weeks. "It's better to try to work together to prioritize safety," Chris Lehane, OpenAI's global policy chief, said at a briefing in Washington, adding that the company does not believe the three need an antitrust waiver to coordinate. The Information has reported that the three are discussing an industry standards body, and Lehane said OpenAI supports provisions in the bipartisan FRONTIER Act that would require independent verification organizations to monitor model development at the top labs. No agreement has been announced.

The move follows Anthropic CEO Dario Amodei's call to "pace the frontier," which we covered on Monday and Tuesday. The political split is unchanged: President Trump has called the safety concerns a "hoax," and his AI adviser David Sacks has said fears of existential risk are exaggerated. What has changed is that the three companies whose models most businesses use are now openly aligning on how those models will be tested. For buyers, outside audits are likely to become a standard part of how AI vendors prove their products are safe.

OpenAI's valuation climbs as its IPO slips

OpenAI has held early talks with investors about a private funding round that could value it at more than $1.2 trillion, according to a Financial Times report. That would be about 41% above the $852 billion valuation from its March round, when it raised $122 billion. Investors started the talks, not OpenAI. The company's annualized revenue topped $40 billion last month, and it spent $34 billion on training models last year. Sam Altman has said a public listing is unlikely before 2027, calling this an ill-advised moment to go public given concerns about AI risk. That means OpenAI will keep publishing less financial detail than a listed company would, so buyers will know less about its costs and pricing.

A model that decides instead of writes

TypeSafe AI, founded by Diogo Almeida, who helped develop the instruction-following research behind ChatGPT at OpenAI, opened early access to Jev, which it calls a "System One Model." Jev does not generate text. It takes a structured question and returns a typed answer along with probabilities, so software can act on the decision directly. TypeSafe says responses take 70 to 500 milliseconds end to end, 40 to 200 times faster than frontier models on comparable tasks, and it prices input at $0.042 per million units of text, with output free.

The company is candid about the limits: its speed tests were run largely from laptops on the West Coast, the test workflows were designed by its own team, the results come from internal evaluations rather than public ones, and it cannot yet prove its pricing is not subsidized. For a business, the idea is what matters. Much of what companies use AI for is small, repeated judgment calls: is this invoice a duplicate, should this ticket escalate, which of these leads is ready. A narrow, very cheap decision engine for those calls could cut AI bills sharply if the claims hold up in real use.

Google's voice models can talk while they work

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The first handles fast spoken conversation in 97 languages, switching languages mid-conversation, and can run tasks in the background without stopping the dialogue. The second reasons and speaks at the same time, using natural cues such as "Let me check that…" while it works through multi-step jobs. Google cites a top score of 82.6 on Artificial Analysis's speech-to-speech quality index, but also 68.6% on the τ-Voice task-completion test and just 35.1% on Sierra's banking version of that test, a reminder that voice agents still fail at many real service calls. Both are available through the Gemini API, in private preview in Gemini Enterprise, and to AI Pro and Ultra subscribers in Gemini Live, Docs, Gmail and Keep.

Meta sells subscriptions, and builds its own chips

Meta launched Meta One, a subscription service across Instagram, Facebook, WhatsApp and Meta AI with more than 50 features and higher AI usage. Consumer plans run from $2.99 a month for WhatsApp Plus to a $19.99 bundle. Business and creator plans start at $14.99 a month, including a verified badge and access to Meta's Business Agent, and climb through $49.99 and $149 to $499 for teams. Meta reports 15 million subscriptions and trials so far. For small businesses that sell through Instagram or WhatsApp, the free tier stays, but the better AI tools now cost money.

Meta also said it will begin putting its third-generation in-house AI chip, the MTIA 450, code-named Arke, into its data centers in the first half of 2027. Twelve samples arrived from TSMC on September 1 and performed within 2% to 3% of Meta's simulations, running Meta's own models plus models from DeepSeek and Alibaba on the first day. "Each chip generation delivers better performance per watt of energy and per dollar spent," said Yee Jiun Song, who leads Meta's custom silicon work. Its successor, MTIA 500 (Astrid), is due in data centers by the end of 2027. Meta shares closed up 0.88%.

Washington and the price of AI computing

Semafor reported that the Commerce Department ordered Kalshi last month to take down a product that combined several markets betting on the cost of renting Nvidia chips into one forward price curve for AI computing, citing national security. Kalshi quietly complied, though the individual markets remain open. Commerce also reportedly pushed the Commodity Futures Trading Commission to effectively freeze approvals of new computing contracts for 60 days. A Commerce spokesperson called the story false and said the department "has never once asked Kalshi to take down this market." Traders suggested one possible worry: thinly traded contracts could be manipulated to show older chips losing value quickly, which could unsettle AI stocks and the debt that funds data centers. Either way, a public signal on where AI computing costs are heading has become politically sensitive.

Physical AI

Waymo began offering fully driverless public rides in Las Vegas on September 14, across a service area of nearly 24 miles that runs from Boulder Junction through the Strip to Sahara Avenue. More than 100,000 people had signed up, and Waymo will admit riders gradually. Las Vegas joins Denver and San Francisco as cities where most of the fleet is Waymo's Ojai vehicle running its sixth-generation driver, with Allegiant Stadium and The Venetian as partners. Waymo says its driver is involved in 94% fewer crashes causing serious injury or worse than human drivers over the same distance. For hotels, venues and event businesses on the Strip, driverless pickup points are now part of planning for how guests arrive.

Odyssey, founded in 2023 by Oliver Cameron and Jeff Hawke, introduced Odyssey-3, a "world model" that learns how the physical world behaves from video and can then be used to control robot arms, humanoids, vehicles and drones. The company says it can train a robot arm with tens of hours of demonstrations and a driving system with 20 hours of simulated data. Its own caveat matters more than its demos: driving policies trained only in simulation reached about 77% of the performance of those trained on real roads. Odyssey says it will release the model publicly in the coming weeks and has not announced pricing.

Nvidia open-sourced OSMO, a system that runs the whole robot-development pipeline — generating synthetic data, training, simulation and hardware testing — from a single YAML configuration file. It schedules work across on-premises clusters, AWS, Azure, Google Cloud and Jetson edge devices, and AI coding assistants can read and monitor its pipelines. Free, vendor-neutral tooling like this lowers the cost for small robotics teams and integrators to build custom automation.

Safety is getting its own shared infrastructure. The Open Source Safety Consortium, with Polymath Robotics among its visible members, is pooling hazard analyses, safety arguments and qualification evidence for open-source software used in safety-critical machines, so each company does not have to repeat that work. Its first deliverable, announced August 25, is an assurance case for an open-source emergency stop. Sensing is attracting serious money too: earlier this month Lyte, led by CEO Alexander Shpunt, raised a $165 million Series C at a $1.6 billion post-money valuation, bringing its total to $272 million. Its LyteVision platform combines 4D vision, imaging and motion sensing and is shipping to robotics customers in inspection, logistics and manufacturing.

Quick Takes

  • AIUC, a startup founded by early Anthropic employee Rune Kvist and former METR COO Rajiv Dattani, audits AI agents against its AIUC-1 standard, modeled on the SOC 2 security standard. It runs agents through about 5,000 tests for jailbreaks, hallucinations and data leaks, then produces a roughly 100-page report that humans check. Customers include Cursor, Lovable, Harvey and ElevenLabs; it has raised $55 million.

  • Nous Research used 1,393 AI sub-agents running Anthropic's Fable 5.1 to refactor a Python codebase of more than 1 million lines in about 19 active hours for roughly $25,000, cutting non-test code by 34.4%. Human reviewers still caught removed names that outside plugins relied on, and regressions in error handling at about 65 places.

  • Cloudflare added a "Disallow AI Training" setting on all plans at no extra cost. It lets sites stay indexed by Googlebot, Applebot and Bingbot while refusing to let those same crawlers use the content for AI training.

  • Profound, which helps brands show up in AI search answers, raised $180 million at a $1.8 billion valuation led by Sequoia and Kleiner Perkins, reporting 3x revenue growth in six months and more than 1,000 enterprise customers including Walmart and Comcast.

What This Means for Your Business

If your sales team lives in Salesforce, test Salesforce in Claude on a small group before rolling it out. It is in beta and runs on the Salesforce permissions you already have, so first make sure those permissions are right: a rep who can see every account in Salesforce will see every account in Claude too. Keep the approve-before-writing default on, and if you are on an individual Claude plan rather than Team or Enterprise, check your training-data settings before connecting customer records. Ask your other software vendors the question Koa raises: which model runs which feature, and does the cheaper model handle the routine work well enough?

List the small, repeated decisions your team or your tools make every day: duplicate invoices, ticket routing, lead scoring, fraud checks. That list is where AI costs pile up, and where decision-only models like Jev claim to be 40 to 200 times faster. Don't switch on vendor claims alone. When early access opens, run 100 of your own past decisions through it and compare the results with what actually happened. In the meantime, you can copy the core idea with the tools you already have: ask for a fixed answer plus a confidence score, and send low-confidence cases to a person.

Start asking AI vendors for independent audit results. With OpenAI backing third-party verification, the three big labs coordinating, and firms like AIUC publishing SOC 2–style standards for agents, "has this been independently tested?" is now a reasonable procurement question. Before you give an agent the ability to send email, spend money or change records, ask for the audit report, not just a sales deck.

Check two settings this week. If you run a website on Cloudflare, decide whether you want your content used for AI training, now that you can stay in search results while opting out. If you sell through Instagram or WhatsApp, compare Meta One's $14.99 Essential business plan with what you currently get free, and see whether the Business Agent actually answers your customers' common questions before paying. And if you run events or hospitality in Las Vegas, add driverless pickup and drop-off to your guest instructions.

Sources