ACM

admin9675

Nvidia’s Switchyard router reshuffles AI models mid-task, cutting task costs to a third in its own tests

Enterprises running always-on AI agents keep hitting the same tradeoff. Send every task to a frontier model and the bill climbs fast. Build custom routing logic to send easy tasks to cheaper models and that becomes its own engineering project, one that has to be maintained every time a workflow changes. Nvidia is proposing a …

Nvidia’s Switchyard router reshuffles AI models mid-task, cutting task costs to a third in its own tests Read More »

OpenAI launches GPT-5.6-Cyber with reduced refusals, 95% completion on advanced cybersecurity tasks

Earlier today, OpenAI launched GPT-5.6-Cyber, a specialized model designed to perform advanced vulnerability research and exploit development for approved defenders — including categories of work that its general-purpose models will often refuse. GPT-5.6-Cyber is a fine-tuned version of OpenAI’s most advanced general model, GPT-5.6 Sol, unveiled back in June, but trained specifically to improve performance …

OpenAI launches GPT-5.6-Cyber with reduced refusals, 95% completion on advanced cybersecurity tasks Read More »

AWS Continuum integrates with OpenAI Codex and Anthropic Claude Code in major AI security push

Amazon Web Services is threading its AI-powered security infrastructure directly into the coding environments built by two of its fiercest rivals — and in doing so, it is making a bold bet that controlling the security layer matters more than controlling the model. AWS announced at Black Hat USA 2026 this month that its Continuum …

AWS Continuum integrates with OpenAI Codex and Anthropic Claude Code in major AI security push Read More »

Brex assumes its AI agents could do anything — so it watches the network, not the code

Brex CEO Pedro Franceschi offered a blueprint for one of the pressing challenges facing the enterprise today at VB Transform 2026: securely deploying AI agents, like the open-source OpenClaw, into production environments. Unlocking this enterprise value requires a mindset shift. The industry needs to move past vague terminology and focus on concrete enterprise roles.  “People …

Brex assumes its AI agents could do anything — so it watches the network, not the code Read More »

Meta returns to open source with Muse Glimmer, an Apache 2.0 licensed 30B parameter AI model optimized for agents — available now

Meta today released Muse Glimmer, a 30-billion-parameter open-weight model designed to run autonomous AI agents directly on consumer hardware — pushing agentic workloads that normally depend on cloud infrastructure onto high-end Macs and PCs. Just as notable as what the model does is how it’s licensed. Glimmer arrives under the permissive, industry-standard Apache 2.0 open …

Meta returns to open source with Muse Glimmer, an Apache 2.0 licensed 30B parameter AI model optimized for agents — available now Read More »

Your agent didn’t hallucinate; it exceeded its authority

Content filters can block unsafe output. They cannot tell you whether an agent was authorized to issue that refund, touch that production system, or commit the company to an external action. Those are different problems, and most enterprises are only solving the first one. An AI agent can follow its instructions perfectly and still take …

Your agent didn’t hallucinate; it exceeded its authority Read More »

Four AI agents coordinating in real time outperformed Claude Opus 4.8 on enterprise coding tasks

As enterprise codebases grow, AI agents tasked with analyzing them are buckling under the weight of long-horizon tasks that require multiple interactions and tool calls. Dividing the work among a team of agents seems like the obvious fix, but it introduces a fatal flaw: most multi-agent systems are not designed for agents to coordinate among …

Four AI agents coordinating in real time outperformed Claude Opus 4.8 on enterprise coding tasks Read More »

Tencent’s Team Memory shares AI agent memory across a team — with no governance yet for when it’s wrong

A VB Pulse survey this June found that 57% of enterprises had traced a confidently wrong agent answer back to missing or inconsistent context — the latest sign of how central context has become to whether AI agents can be trusted to act on their own. Most of the fixes so far have solved a …

Tencent’s Team Memory shares AI agent memory across a team — with no governance yet for when it’s wrong Read More »

Stanford is running 37,000 AI agents as a virtual biotech — and one of its drug designs got independently confirmed by Merck

For developers, the operating assumption has been one engineer, one agent — the model Claude Code and similar tools. At VB Transform 2026, James Zou, associate professor of biomedical data science at Stanford University, argued that assumption is about to break: the next frontier isn’t a single, more capable agent, it’s tens of thousands of …

Stanford is running 37,000 AI agents as a virtual biotech — and one of its drug designs got independently confirmed by Merck Read More »