Sai navigates online museum 3D viewers, controls 360-degree orbit rotations and zoom levels, and inspects anatomical structures without manual drag-and-drop.
The recording is a real session. The sheet on the right is what it produced.
Sai opens each profile, pulls the signal, and writes the row, live, in a real browser.

Eight columns, sorted by score, with a source link behind every claim.
Target 3D museum portal or model repository link (e.g. Smithsonian 3D, Sketchfab) Target specimen name, catalog ID, or historical period Specific anatomical features or angles required for inspection
Multi-perspective high-resolution visual capture (anterior, lateral, dorsal, close-up) Structured extraction of specimen dimensions, metadata, and curator annotations Formatted scientific inspection dossier with direct links to original 3D coordinates
Run as an event-driven research task whenever a new fossil batch or archaeological artifact is digitized in public repositories.
Autonomous WebGL 3D specimen exploration is the programmatic manipulation of web-based 3D model viewports (powered by WebGL, Three.js, or Babylon.js) by an AI agent that controls virtual camera orbits, executes multi-axis rotational dragging, adjusts focal zoom depths, and extracts morphological data from spatial annotations without manual human interaction.
Global cultural and scientific institutions—such as the Smithsonian Institution, the British Museum, and the Natural History Museum—have digitized hundreds of thousands of rare paleontology fossils, archaeological relics, and biological specimens into interactive 3D formats. However, conducting academic review, visual asset harvesting, or educational cataloging requires researchers to manually drag, tilt, zoom, and screenshot models from multiple orientations. Autonomous spatial agents bridge this capability gap by operating within the WebGL canvas, capturing standardized orthographic and perspective views, and reading interactive 3D spatial callout pins.
Standard browser automation frameworks break when encountering 3D canvases because the individual anatomical elements (bones, teeth, surface fractures) do not exist as distinct HTML DOM nodes. The entire model renders as a flat GPU context, making standard element selectors completely ineffective.
Sai overcomes this limitation through embodied computer use and multimodal vision reasoning:
Free your hands from the computer.
Run this Task