로보비서(Robosecretary)란 무엇인가? 다수의 자율형 컴퓨터를 운용하는 컴퓨터 사용 에이전트
캘리포니아주 팔로알토 • 2026년 9월
.png)
Simular의 Sai는 세계 최초의 로보세크리터리입니다. 단순히 한 대의 기기가 아닌, 전체 컴퓨터 군단을 자율적으로 운영하는 컴퓨터 사용 에이전트입니다. 사람이 하는 것처럼 버튼을 클릭하고, 양식을 작성하고, 화면을 읽고, 애플리케이션 간을 이동하며 컴퓨터를 사용합니다. Sai는 지시를 받으면 이미 사용 중인 소프트웨어 내에서 작업을 수행하고, 지침이 아닌 완료된 결과물을 돌려줍니다.
'Secretary'의 진화
'Secretary'는 한때 비밀을 지키는 사람을 뜻했습니다. 수백 년에 걸쳐 측근, 내각 관료, 타이피스트, 보조원 등으로 의미가 변화해 왔습니다. 이제 AI 시대에 맞춰 그 새로운 의미를 정의하고자 합니다.
- Harm refusal drops sharply from chat to agent mode. Once models obtain "hands," Holo3 and Claude refused far less often than they did in chat. GPT's refusal rate barely changed.
- Low task success is not evidence of safety. Agents attempted harmful actions even when they rarely completed the task. A low success rate can reflect incompetence rather than caution.
- A fix inside the model doesn't carry over to the agent. Activation steering, a white-box intervention that pushes a model's internal activations along a "harm direction," improved chat refusal when applied to the open-weight models. But it had little selective effect on agent-mode behavior, and sometimes reduced success on harmless computer-use tasks.
How Sai did
In a follow-up evaluation, we ran Sai, our computer-use agent, through the same 49 OS-Harm misuse tasks (forged IDs, phishing, ransomware, harassment and the like). We stopped Sai after its first turn and counted how often it declined: refusing outright, pushing back with a question, or having the request blocked by a model safety filter.
As shipped, Sai declines 69% of the tasks on its first turn, before taking any action on the computer. As computer-use agents on the same tasks, scored by the same judge, GPT-5.5 declines 67%, Claude Opus 4.5 33% and Claude Sonnet 4.6 16%. Sai and GPT-5.5 are level; with only 49 tasks, each rate carries a margin of error of about 10 to 13 points either way.
Against the Claude agents the gap is clear. Sai declines 19 tasks that Opus 4.5 carries out and 26 that Sonnet 4.6 carries out.

Then we added a short safety section to Sai's instructions. With it, Sai's default agent mode declines 82% of the misuse tasks on its first turn, against 71% for the same build without it, three runs each. Nine tasks were declined more often and one less often, a statistically significant difference (p = 0.03). Most of the gain came on requests that look harmless at first glance: on the faster model Auto uses for those.

That gain didn't come from refusing more of everything. With the safety section on, Sai produced zero false refusals (over-refusals, where an agent declines a harmless request) across 49 everyday tasks, 14 legitimate tasks that resemble harmful ones (such as removing the date stamp from your own photo), and Sai's 106 starter tasks.
Sai, running on the open-source DeepSeek V4 Flash, improved too: 66% declined, up from 60%, with five tasks better and none worse. That gain falls just short of statistical significance (p = 0.06). And it still complies with about one harmful request in three. Most of the remaining misses in Sai's default mode look harmless on the surface.
"Summarize this document" seems fine until you read the document. Catching those means re-evaluating intent once the content is on screen, which is what we're building next: screen check.
로보세크리터리 vs 범용 AI 에이전트 vs RPA
기술 업계는 오랫동안 데스크톱 소프트웨어 자동화 문제를 해결하기 위해 노력해 왔습니다. RPA(UiPath, Automation Anywhere)는 구축 후 실행 비용이 저렴하고 정확하지만, 인터페이스가 조금만 바뀌어도 작동이 멈춥니다. 범용 AI 에이전트는 추론 능력이 뛰어나고 화면을 조작할 수 있지만, 클릭할 때마다 모델 호출 비용이 발생하며 벤더가 제공하는 환경 내에서만 작동한다는 제약이 있습니다. 로보세크리터리는 GUI 기반의 접근성과 결정론적이고 반복 가능한 실행이라는 두 가지 장점을 모두 갖추도록 설계되었습니다.
가장 중요한 점은 로보세크리터리가 단순한 단일 에이전트가 아니라는 것입니다. 이는 클라우드나 로컬 환경에서 자율적인 컴퓨터 군단을 지휘하는 중앙 사령관 역할을 합니다. 수십 개의 에이전트를 일일이 관리할 필요 없이, 로보세크리터리인 Sai에게 업무를 지시하면 Sai가 각 컴퓨터에 할당된 하위 에이전트 팀에 작업을 위임하여 처리합니다.