Article
Google DeepMind SIMA: Research Design and Results
A research profile of SIMA and SIMA 2, including their virtual environments, language-driven interfaces, reported progress, and physical-transfer limits.
Google DeepMind SIMA Research Profile
SIMA is Google DeepMind research on agents that follow language instructions and act across varied virtual 3D environments. The first SIMA used image observations and language instructions to produce keyboard-and-mouse actions. SIMA 2 later added Gemini-based reasoning, conversation, more complex tasks, and virtual-world self-improvement experiments.
SIMA is a research program, not a public robot controller or proof of physical-world deployment. Its results should remain attached to the named environments, task definitions, data, agent version, evaluation, and report.
flowchart LR
A["Image observations"] --> C["SIMA agent"]
B["Language instruction"] --> C
C --> D["Keyboard and mouse actions"]
D --> E["Virtual 3D environment"]
E --> A
E --> F["Task result and evaluation"]
Research record
| Field | SIMA 1 |
|---|---|
| Name | Scalable Instructable Multiworld Agent |
| Announced | March 13, 2024 |
| Organization | Google DeepMind |
| Setting | Commercial games and curated research environments |
| Inputs | Image observations and language instructions |
| Outputs | Keyboard-and-mouse actions |
| Primary question | Can one agent follow free-form instructions across varied 3D worlds? |
| Physical deployment | Not established |
The research state was checked on July 28, 2026. Later papers, datasets, access, benchmarks, environments, model versions, and physical-world experiments require a fresh primary source.
The research question
Google DeepMind introduced SIMA in March 2024. The team shifted from optimizing an agent for one game score toward a general instructable agent across environments.
That distinction corrects a common misunderstanding. SIMA was not described as an agent that wandered without purpose. A person supplied language instructions, and the research evaluated whether the agent could ground those instructions in perception and action.
The target was broad interaction through a shared interface. The system saw images rather than privileged game state and acted through keyboard and mouse rather than a custom control API for every environment.
Why games were useful
Games can provide responsive 3D environments with navigation, objects, changing goals, visual variation, and real-time action. They also make experiments repeatable and keep early failures inside a virtual system.
The researchers partnered with game developers and used commercial titles alongside research environments. That diversity was part of the generalization question.
A game still represents a designed world. Its physics, objects, rules, controls, data, and visual cues are not the physical world. The environments are useful because they are rich and bounded, not because they perfectly reproduce reality.
What the first report established
The SIMA technical report describes the motivation, environments, data collection, architecture, evaluation, and preliminary results. Google DeepMind's public summary says the agent learned more than 600 language-following skills.
Exact performance claims require the report's metric and split. A success rate in one task group should not be repeated as a universal agent score. Human baselines, single-environment agents, held-out environments, and instruction categories answer different questions.
The research supports the statement that one agent showed instruction-following behavior across multiple studied virtual environments. It does not support unrestricted competence across any game or any physical setting.
What SIMA did not establish
SIMA did not demonstrate consciousness, sentience, a self-generated purpose, legal personhood, or artificial general intelligence as a completed state.
It did not demonstrate physical manipulation, robot safety, a household deployment, or transfer to an uncontrolled environment. Keyboard-and-mouse action is embodied in a virtual sense, but it does not include motor torque, collision energy, sensor wear, maintenance, or human physical safety.
The E008 transcript speculated about those broader possibilities. The public research profile keeps them outside the result.
SIMA 2 expanded the program
Google DeepMind introduced SIMA 2 in November 2025. The follow-on integrated Gemini capabilities and added dialogue, abstract and higher-level instructions, explanations, and learning experiments inside new virtual environments.
The SIMA 2 technical report reports improved performance, generalization to held-out environments, and self-improvement experiments in which Gemini generated tasks and rewards.
Those are meaningful changes to the virtual-agent research question. They remain first-party research results. The authors describe eventual physical-world relevance as a path, not as a completed transfer result.
How to read the self-improvement claim
Self-improvement can sound like unrestricted autonomous evolution. The report describes a bounded experiment inside a virtual environment with a specified model, task-generation process, reward process, and evaluation.
The correct questions concern who defined the environment, which actions were possible, how tasks and rewards were generated, what data was retained, what metric improved, and which safety or access boundaries remained fixed.
A self-improvement result inside one experimental loop does not establish open-ended improvement across systems or the ability to rewrite every part of the agent.
What physical transfer would require
A physical experiment would need a target body, sensors, control interface, task, environment, success measure, duration, failure threshold, intervention boundary, and safety case.
Researchers would have to measure the gap between images and game controls on one side and real sensors, latency, contact, force, wear, people, and maintenance on the other. A promising virtual policy could become one input to that process. It would not skip it.
NIST's AI Resource Center provides testing, evaluation, verification, and validation resources that can help structure the AI record. The physical system would add its own domain standards and operating controls.
Venture Step context
E008 connects SIMA to Blackwell and Figure 01 without claiming that the systems were integrated. E010 and E068 develop the simulation and robotics-model thread. E032 and E035 explore tools and source-grounded AI systems, while E025 supplies a reusable operating-domain boundary.
Read [[Simulation Accelerates Embodied-System Development]] for the learning loop and [[What Blackwell Figure 01 and SIMA Demonstrated in 2024]] for the source-era comparison.
This research profile was developed with AI assistance from the preserved E008 transcript and the linked Google DeepMind technical reports, announcements, and NIST resource. Google DeepMind sources establish the team's methods and reported results, not independent physical-world performance. Research, benchmark, transfer, safety, accessibility, editorial, and founder review remain required. Publication is unauthorized.
Sources
Follow the evidence.
- NVIDIA Rubin announcementnvidianews.nvidia.com
- SIMA 2 technical reportstorage.googleapis.com
- NIOSH Center for Occupational Robotics Researchcdc.gov
- OSHA robotics overviewosha.gov
- NIST AI Resource Centerairc.nist.gov
- Google DeepMind SIMA 2 announcementdeepmind.google
- NVIDIA Blackwell Ultra announcementnvidianews.nvidia.com
- BMW Figure 02 trialpress.bmwgroup.com
- Spotify episode recordpodcasters.spotify.com
- Google DeepMind SIMA announcementdeepmind.google
- Figure news indexfigure.ai
- Figure 03 introductionfigure.ai
- NASA Systems Engineering Handbooknasa.gov
- Figure Helix 02figure.ai
- NVIDIA Blackwell launchinvestor.nvidia.com
- BMW Figure 03 projectpress.bmwgroup.com
- SIMA technical reportstorage.googleapis.com
- Figure company pagefigure.ai