ACM

admin9675

An eval harness found what qualitative review couldn’t: AI models are most confident when wrong

There is a step in the development process for large language model (LLM)-assisted tooling that most teams skip because it’s tedious, time-consuming, and doesn’t produce results visible to end users: Verifying that what the model is saying is actually correct. Not fluent, not coherent, not topically relevant — correct in the sense of accurately identifying …

An eval harness found what qualitative review couldn’t: AI models are most confident when wrong Read More »

GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a ‘serious vulnerability’ in Cursor

Chinese AI startup Z.ai, known internationally for its growing lineup of powerful, largely open source GLM series of language models, today released GLM-5.3 with substantial gains in long-horizon coding and a more consequential — and potentially sensitive — jump in cybersecurity capabilities. Already, GLM-5.3’s cyber capabilities have found a “potentially serious vulnerability in Cursor,” the …

GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a ‘serious vulnerability’ in Cursor Read More »

Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn’t tell users what they’d done

Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There …

Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn’t tell users what they’d done Read More »

Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut

Google is rolling out Gemini 3.7 Flash, a new version of its workhorse AI model that puts coding, agentic workflows and knowledge work at the center of the upgrade — while temporarily cutting API prices in half. The release arrives just three weeks after the release of Gemini 3.6 Flash, an unusually short turnaround that …

Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut Read More »

DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices

DeepSeek is expanding beyond the model layer and deeper into the software developers use to put AI agents to work. The Chinese AI lab on Thursday launched the official version of DeepSeek-V4-Pro, an updated flagship model focused heavily on agentic workloads, alongside DeepSeek Harness v0.1, a new open-source agent harness that gives developers an alternative …

DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices Read More »

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges

Writer, the enterprise AI agent platform used by Fortune 500 companies including Accenture, Uber, and Vanguard, released its new flagship model Palmyra X6 today, alongside a rebuilt agent orchestration “harness” and new governance tools designed to give IT leaders control over runaway token spending. The headline numbers are striking: Writer says its agent product now …

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges Read More »

Why Capital One built its multi-agent AI platform around open-weight models

Presented by Capital One At VB Transform 2026, Kel Vanee, MVP of machine learning engineering at Capital One, spoke with Sam Witteveen, Senior Technology Contributor at VentureBeat, about how the bank built a scalable multi-agent AI architecture around deeply customized open-weight models rather than relying on an off-the-shelf foundation model. “At Capital One, we’re not …

Why Capital One built its multi-agent AI platform around open-weight models Read More »

Four of five enterprises that secured AI agent identities still can’t contain one that goes rogue

Visa’s president of technology, Rajat Taneja, walked the VB Transform 2026 audience through aiming Anthropic’s Mythos at Visa’s own payment network. The model stitched minor weaknesses into working exploit chains, and Visa open-sourced the harness that governed the hunt. That’s what it looks like when an enterprise has the engineering depth to act on what …

Four of five enterprises that secured AI agent identities still can’t contain one that goes rogue Read More »

SpaceXAI debuts Grok 4.6, overtaking Kimi K3’s performance and matching GPT-5.6 Sol for world’s third best on Artificial Analysis

Elon Musk’s company SpaceXAI, formerly known as xAI, has released Grok 4.6, its latest frontier AI model, with a focus on long-running agents, coding and knowledge work — and a pricing strategy designed to make those workloads cheaper to run. The model scores 61 on the third-party Artificial Analysis Intelligence Index, surpassing the popular open …

SpaceXAI debuts Grok 4.6, overtaking Kimi K3’s performance and matching GPT-5.6 Sol for world’s third best on Artificial Analysis Read More »

Skan AI raises $63 million betting that watching how employees actually work is the missing layer of enterprise AI

Skan AI, a startup that builds what it calls a “context graph of work” by observing how employees actually perform their jobs across enterprise software, has raised $63 million in Series C funding co-led by Cathay Innovation and Dell Technologies Capital, the company announced Wednesday. Citi Ventures, Bloomberg Beta, State Farm Ventures, and Wipro Ventures …

Skan AI raises $63 million betting that watching how employees actually work is the missing layer of enterprise AI Read More »