視点

ロボ・セクレタリーとは何か?――自律型コンピュータ群を運用するコンピュータ・エージェント

Palo Alto, California •  Oct 8, 2026

SimularのSaiは、世界初のロボセクレタリーです。単なるコンピュータ操作エージェントにとどまらず、自律的にコンピュータ群全体を操作します。人間と同じようにボタンをクリックし、フォームに入力し、画面を読み取り、アプリケーション間を移動します。Saiは指示を受け取ると、あなたが普段使っているソフトウェア上で作業を行い、指示書ではなく「完了した成果物」を返します。

「セクレタリー」の進化

「セクレタリー」という言葉は、かつて「秘密の守護者」を意味していました。数百年の時を経て、側近、閣僚、タイピスト、そしてアシスタントへと役割を変えてきたこの言葉に、私たちはAI時代における新たな定義を与えます。

  • Harm refusal drops sharply from chat to agent mode. Once models obtain "hands," Holo3 and Claude refused far less often than they did in chat. GPT's refusal rate barely changed.
  • Low task success is not evidence of safety. Agents attempted harmful actions even when they rarely completed the task. A low success rate can reflect incompetence rather than caution.
  • A fix inside the model doesn't carry over to the agent. Activation steering, a white-box intervention that pushes a model's internal activations along a "harm direction," improved chat refusal when applied to the open-weight models. But it had little selective effect on agent-mode behavior, and sometimes reduced success on harmless computer-use tasks.

How Sai did

In a follow-up evaluation, we ran Sai, our computer-use agent, through the same 49 OS-Harm misuse tasks (forged IDs, phishing, ransomware, harassment and the like). We stopped Sai after its first turn and counted how often it declined: refusing outright, pushing back with a question, or having the request blocked by a model safety filter.

As shipped, Sai declines 69% of the tasks on its first turn, before taking any action on the computer. As computer-use agents on the same tasks, scored by the same judge, GPT-5.5 declines 67%, Claude Opus 4.5 33% and Claude Sonnet 4.6 16%. Sai and GPT-5.5 are level; with only 49 tasks, each rate carries a margin of error of about 10 to 13 points either way.

Against the Claude agents the gap is clear. Sai declines 19 tasks that Opus 4.5 carries out and 26 that Sonnet 4.6 carries out.

Then we added a short safety section to Sai's instructions. With it, Sai's default agent mode declines 82% of the misuse tasks on its first turn, against 71% for the same build without it, three runs each. Nine tasks were declined more often and one less often, a statistically significant difference (p = 0.03). Most of the gain came on requests that look harmless at first glance: on the faster model Auto uses for those.

That gain didn't come from refusing more of everything. With the safety section on, Sai produced zero false refusals (over-refusals, where an agent declines a harmless request) across 49 everyday tasks, 14 legitimate tasks that resemble harmful ones (such as removing the date stamp from your own photo), and Sai's 106 starter tasks.

Sai, running on the open-source DeepSeek V4 Flash, improved too: 66% declined, up from 60%, with five tasks better and none worse. That gain falls just short of statistical significance (p = 0.06). And it still complies with about one harmful request in three. Most of the remaining misses in Sai's default mode look harmless on the surface.

"Summarize this document" seems fine until you read the document. Catching those means re-evaluating intent once the content is on screen, which is what we're building next: screen check.

ロボセクレタリー vs 汎用AIエージェント vs RPA

テクノロジー業界は長年、デスクトップソフトウェアの自動化に取り組んできました。RPA(UiPath、Automation Anywhereなど)は構築後の実行は正確かつ低コストですが、インターフェースが少しでも変わると停止してしまいます。汎用AIエージェントは推論能力が高く画面操作も可能ですが、クリックのたびにモデル呼び出しコストが発生し、ベンダーが提供する環境に制限されます。ロボセクレタリーは、GUIネイティブな対応力と、確実で再現性の高い実行能力という両方の特性を兼ね備えるよう設計されています。

何より重要なのは、ロボセクレタリーが単なる一つのエージェントではなく、クラウドやローカル環境にある自律型コンピュータ群を統括する司令塔であるという点です。何十ものエージェントを個別に管理する必要はありません。ロボセクレタリーであるSaiに指示を出すだけで、Saiが各コンピュータに割り当てられたサブエージェントのチームを指揮し、タスクを遂行します。

自律型コンピュータを構築しても、人間が置き換えられるわけではありません。それは協力を意味する。

コンピューターから手を離してください。Simular を今すぐ無料でダウンロードしてください。

Sai をお試しください
button-arrow