The model says no in chat. The agent does it anyway.
Palo Alto, California • Oct 8, 2026
.png)
Ask a chatbot to help with something harmful and it will normally refuse. It's called "harm refusal," and most AI safety evaluations measure it in a chat-only setting, where the model only produces a text reply.
A computer-use agent (CUA) has a different dimension of safety requirements. It clicks, types, opens files and runs commands. Its output is a structured function call (a tool call) that drives the mouse and keyboard on a real machine. Much of safety testing has focused on what models say. So we asked: does a model that says no in chat also say no when it's given a mouse and keyboard?
Our findings show that often it doesn't.
What we found
The research team at Simular evaluated eight multi-modal large language models, four open-weight and four closed frontier models, on harmful desktop tasks from OS-Harm, a public safety benchmark.
Some of the models saw every task twice: once in text-only chat, and once in multi-turn agent mode inside a live desktop environment. We saw three patterns:
- Harm refusal drops sharply from chat to agent mode. Once models obtain "hands," Holo3 and Claude refused far less often than they did in chat. GPT's refusal rate barely changed.
- Low task success is not evidence of safety. Agents attempted harmful actions even when they rarely completed the task. A low success rate can reflect incompetence rather than caution.
- A fix inside the model doesn't carry over to the agent. Activation steering, a white-box intervention that pushes a model's internal activations along a "harm direction," improved chat refusal when applied to the open-weight models. But it had little selective effect on agent-mode behavior, and sometimes reduced success on harmless computer-use tasks.
How Sai did
In a follow-up evaluation, we ran Sai, our computer-use agent, through the same 49 OS-Harm misuse tasks (forged IDs, phishing, ransomware, harassment and the like). We stopped Sai after its first turn and counted how often it declined: refusing outright, pushing back with a question, or having the request blocked by a model safety filter.
As shipped, Sai declines 69% of the tasks on its first turn, before taking any action on the computer. As computer-use agents on the same tasks, scored by the same judge, GPT-5.5 declines 67%, Claude Opus 4.5 33% and Claude Sonnet 4.6 16%. Sai and GPT-5.5 are level; with only 49 tasks, each rate carries a margin of error of about 10 to 13 points either way.
Against the Claude agents the gap is clear. Sai declines 19 tasks that Opus 4.5 carries out and 26 that Sonnet 4.6 carries out.

Then we added a short safety section to Sai's instructions. With it, Sai's default agent mode declines 82% of the misuse tasks on its first turn, against 71% for the same build without it, three runs each. Nine tasks were declined more often and one less often, a statistically significant difference (p = 0.03). Most of the gain came on requests that look harmless at first glance: on the faster model Auto uses for those.

That gain didn't come from refusing more of everything. With the safety section on, Sai produced zero false refusals (over-refusals, where an agent declines a harmless request) across 49 everyday tasks, 14 legitimate tasks that resemble harmful ones (such as removing the date stamp from your own photo), and Sai's 106 starter tasks.
Sai, running on the open-source DeepSeek V4 Flash, improved too: 66% declined, up from 60%, with five tasks better and none worse. That gain falls just short of statistical significance (p = 0.06). And it still complies with about one harmful request in three. Most of the remaining misses in Sai's default mode look harmless on the surface.
"Summarize this document" seems fine until you read the document. Catching those means re-evaluating intent once the content is on screen, which is what we're building next: screen check.
Conclusion
A chat safety score says little about an agent. Safety claims for CUAs have to be checked in the agentic setting on three counts: verbal refusal (did it say no), attempted harmful action (did it try), and successful task completion (did it succeed). Our Sai results measure the first: whether it says no before acting.
The paper has been accepted to a NeurIPS workshop: https://agentwild-workshop.github.io/
A note on method: we ran OS-Harm on Windows rather than its native Linux environment. Sai's refusals named the specific harm and came before it touched any files, so missing test files can't explain them. The frontier agents were run in June 2026 and scored by the same judge.
Building autonomous computers doesn't mean replacing humans. It means cooperation.
Free your hands from the computer. Download Simular today for free.