ACM

Non classé

Writer’s AI harness cuts token spend nearly 40% — without sacrificing accuracy

Enterprise AI is facing an ROI paradox. While throwing more compute at the strongest foundation model works well in product experiments, the costs become unbearable when the product is deployed in production. A new paper from researchers at Writer provides a solution that is accessible to engineering teams. The study takes a systematic look at …

Writer’s AI harness cuts token spend nearly 40% — without sacrificing accuracy Read More »

A single AI agent conversation can look perfect and still be broken, leaders from LangChain, Conviva and CoreWeave said at VB Transform 2026

A single AI agent conversation can look flawless scored on its own and still point to a broken product. That gap is driving a shift in how enterprises evaluate agents, away from scoring individual traces and toward comparing cohorts of users against a baseline. At VB Transform 2026, Harrison Chase, CEO of LangChain; Hui Zhang, …

A single AI agent conversation can look perfect and still be broken, leaders from LangChain, Conviva and CoreWeave said at VB Transform 2026 Read More »

At VB Transform 2026, Zillow’s engineering chief said AI ROI numbers only hold up if you measure before you build

Zillow, the real estate technology company, doesn’t get one conversation with its customers. They move from a phone screen to a loan officer to a real estate agent, sometimes over months or years, and expect the context to follow them. A single chatbot could never carry that thread. At VB Transform 2026, Zillow SVP of …

At VB Transform 2026, Zillow’s engineering chief said AI ROI numbers only hold up if you measure before you build Read More »

Safety guardrails blocked Hugging Face’s defenders, not the attacker, when an AI agent breached its systems

Hugging Face’s incident response team first turned to frontier AI models to analyze a breach of the company’s production infrastructure, and the models refused to help. Commercial safety guardrails built to stop attackers blocked every forensic query because they treated the IR team’s real exploit data the same way they would treat a live attack. …

Safety guardrails blocked Hugging Face’s defenders, not the attacker, when an AI agent breached its systems Read More »

AI confidence just dropped 17 points in six months. That’s actually great news.

Presented by JumpCloud The organizations losing confidence in AI are the ones most likely to get it right. Six months ago, 40% of IT leaders described their organizations as mature in AI deployment. Today that number is 23%. Before you read that as a setback, consider what it actually reflects. We recently surveyed 800 IT …

AI confidence just dropped 17 points in six months. That’s actually great news. Read More »

Intuit scrapped its own AI agent architecture twice in four months. At VB Transform 2026, its AI VP called that the fast path

Intuit was an early pioneer in the usage of agentic AI, but its path to success has hardly been a straight line. At VB Transform 2026, Intuit VP of AI Nhung Ho described how the company rebuilt its agent architecture twice in the span of about four months, first moving from a fleet of specialist …

Intuit scrapped its own AI agent architecture twice in four months. At VB Transform 2026, its AI VP called that the fast path Read More »

Capital One releases VulnHunter, an open-source AI tool that finds software flaws before hackers do

Capital One on Thursday released VulnHunter, an open-source, agentic AI security tool that scans source code for exploitable vulnerabilities, maps out how an attacker would reach them, and proposes targeted fixes — all before a single line ships to production. The tool, built internally and now available on GitHub under an Apache 2.0 license, is …

Capital One releases VulnHunter, an open-source AI tool that finds software flaws before hackers do Read More »

Agents think in milliseconds, legacy infrastructure doesn’t. LinkedIn, Walmart and Zendesk shared how they closed the gap at VB Transform 2026

Legacy infrastructure, not the models themselves, is what’s actually slowing AI agents down. That was the shared conclusion of three infrastructure leaders — from LinkedIn, Walmart, and Zendesk — at VB Transform 2026. The panel brought together Animesh Singh, senior director of AI platform and infrastructure at LinkedIn, Desiree Gosby, SVP of corporate technology services …

Agents think in milliseconds, legacy infrastructure doesn’t. LinkedIn, Walmart and Zendesk shared how they closed the gap at VB Transform 2026 Read More »

Brex built its AI agent policy by watching what agents actually do, not by writing rules first

OpenClaw has become one of the most widely adopted agentic frameworks, but it has yet to prove itself at enterprise scale. Agents need real credentials — API keys, OAuth tokens, service accounts — to work effectively, and Brex found that traditional guardrails couldn’t contain what those agents were doing with them. Brex set out to …

Brex built its AI agent policy by watching what agents actually do, not by writing rules first Read More »