什么是“机器人秘书”?一种能够操控多个自主运行计算机的计算机代理程序。
Palo Alto, California • Oct 8, 2026

来自 Simular 的 Sai 是全球首款机器人秘书。它不仅是一个计算机操作智能体,还能自主指挥整个计算机集群。它像人类一样使用计算机——点击按钮、填写表格、阅读屏幕并在应用程序之间切换。Sai 接收指令,在你现有的软件中工作,并直接交付最终成果,而不是仅仅提供操作步骤。
“秘书”的演变
“秘书”一词曾意为“机密保管人”。几百年来,它经历了心腹、内阁成员、打字员和助理等多种身份。在 AI 时代,我们正在赋予它全新的定义。
- Harm refusal drops sharply from chat to agent mode. Once models obtain "hands," Holo3 and Claude refused far less often than they did in chat. GPT's refusal rate barely changed.
- Low task success is not evidence of safety. Agents attempted harmful actions even when they rarely completed the task. A low success rate can reflect incompetence rather than caution.
- A fix inside the model doesn't carry over to the agent. Activation steering, a white-box intervention that pushes a model's internal activations along a "harm direction," improved chat refusal when applied to the open-weight models. But it had little selective effect on agent-mode behavior, and sometimes reduced success on harmless computer-use tasks.
How Sai did
In a follow-up evaluation, we ran Sai, our computer-use agent, through the same 49 OS-Harm misuse tasks (forged IDs, phishing, ransomware, harassment and the like). We stopped Sai after its first turn and counted how often it declined: refusing outright, pushing back with a question, or having the request blocked by a model safety filter.
As shipped, Sai declines 69% of the tasks on its first turn, before taking any action on the computer. As computer-use agents on the same tasks, scored by the same judge, GPT-5.5 declines 67%, Claude Opus 4.5 33% and Claude Sonnet 4.6 16%. Sai and GPT-5.5 are level; with only 49 tasks, each rate carries a margin of error of about 10 to 13 points either way.
Against the Claude agents the gap is clear. Sai declines 19 tasks that Opus 4.5 carries out and 26 that Sonnet 4.6 carries out.

Then we added a short safety section to Sai's instructions. With it, Sai's default agent mode declines 82% of the misuse tasks on its first turn, against 71% for the same build without it, three runs each. Nine tasks were declined more often and one less often, a statistically significant difference (p = 0.03). Most of the gain came on requests that look harmless at first glance: on the faster model Auto uses for those.

That gain didn't come from refusing more of everything. With the safety section on, Sai produced zero false refusals (over-refusals, where an agent declines a harmless request) across 49 everyday tasks, 14 legitimate tasks that resemble harmful ones (such as removing the date stamp from your own photo), and Sai's 106 starter tasks.
Sai, running on the open-source DeepSeek V4 Flash, improved too: 66% declined, up from 60%, with five tasks better and none worse. That gain falls just short of statistical significance (p = 0.06). And it still complies with about one harmful request in three. Most of the remaining misses in Sai's default mode look harmless on the surface.
"Summarize this document" seems fine until you read the document. Catching those means re-evaluating intent once the content is on screen, which is what we're building next: screen check.
机器人秘书、通用 AI 智能体与 RPA 的对比
科技行业长期以来一直致力于解决桌面软件自动化问题。RPA(如 UiPath、Automation Anywhere)一旦构建完成,运行精准且成本低廉,但一旦界面发生变动就会失效。通用 AI 智能体具备良好的推理能力并能操作屏幕,但每次点击都需要支付模型调用费用,且受限于供应商提供的环境。机器人秘书旨在同时兼顾这两者的优势:既拥有 GUI 原生操作能力,又能实现确定性、可重复的执行。
最重要的是,机器人秘书不仅仅是一个智能体;它是一个中央指挥官,能够指挥云端或本地的计算机集群。你无需忙于协调数十个智能体,只需与你的机器人秘书 Sai 对话,它便会指派其下属智能体团队在各自的计算机上完成任务。