Ilustración editorial para Gemini Robotics ER 2: qué planifica el modelo y qué debe validar el robot
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

What Gemini Robotics ER 2 is and what embodied reasoning means

Gemini Robotics ER 2 is a vision-language model intended for robotics applications. In this context, ER means embodied reasoning: interpreting information about a physical environment and reasoning about a task in that environment. Google’s official documentation describes Gemini Robotics ER models as models that enable robots to perceive and interact with the physical world. For ER 2, the model page attributes multi-step task planning to the system and distinguishes that function from subsequent motor execution, which is assigned to a lower-level system.

That distinction matters because “reasoning about a robotics task” is not the same as “moving a robot autonomously and safely.” The available description supports the idea that the model can participate in a chain of interpretation and planning. It does not, by itself, show that the model directly controls motors, that a proposed sequence can be executed by every platform, or that a task will be completed at any particular success rate.

It is therefore useful to view ER 2 as a potential component in an architecture, not as a complete robot specification. Perception may depend on the camera and input format; execution may depend on controllers, sensors, and physical limits; and supervision may depend on integration decisions made outside the model. The information reviewed is not sufficient to establish which hardware, software, and environmental combinations have been validated end to end.

02

Separate perception, planning, and movement

To analyze the model’s role, it helps to divide the system into observable responsibilities. First, a visual input must represent the scene well enough for the task: relevant objects and positions, obstacles, or changes. Next, a component interprets the instruction and decides which steps might achieve the goal. Finally, a lower-level controller converts instructions or references into physical movement. The documentation reviewed describes ER 2’s planning function and its separation from execution, but the available excerpts do not specify every detail of that interface.

This decomposition avoids attributing to the model achievements that depend on the system as a whole. If an object is not detected because it is occluded, a later decision may be unsuitable even if the textual reasoning appears coherent. If a plan is reasonable, a controller may still be unable to execute it because of reach, precision, or configuration limits. And even when the controller performs the movement, it remains necessary to check whether a hazardous condition arose during the action.

The following workflow is an analytical guide, not an exhaustive description of Google’s implementation. Teams should identify which component receives each piece of data, which component produces each decision, and which controls prevent an unvalidated proposal from reaching an actuator. Without that allocation of responsibilities, errors can be obscured by broad labels such as “model failure” or “robot failure.”

Responsibility chain for evaluation

  1. 01Input and observation: document which images, instructions, and additional data the system receives, and under what capture conditions.
  2. 02Interpretation and plan: record the instruction as understood, the proposed steps, and any uncertainty the system communicates.
  3. 03Validation: check that the plan complies with operating limits and safety rules defined by the team before authorizing movement.
  4. 04Execution and supervision: identify the responsible controller, record the robot’s state, and stop or review the action if deviations occur.
03

Documented capabilities and questions about the API

The Gemini Robotics ER Interactions API documentation presents the family as vision-language models that let robots perceive and interact with the physical world, and identifies ER 2 in that context. Another official page describes a way to use the model through generateContent. The existence of documentation for more than one interface does not, without reviewing the complete specifications, establish that both offer the same operations, formats, limits, or behavior for ER 2.

The official model page attributes multi-step task planning to ER 2. That is a capability description, not an evaluation protocol or an exhaustive list of approved tasks. Nor does it imply that the model can resolve every ambiguous instruction, work with every camera, or produce a motor trajectory that is ready for execution. Details about inputs, outputs, tool calls, and error handling should be checked in the current documentation for the chosen API.

Before integrating the model, a team should answer concrete questions: What structure does the response have? Can it express uncertainty or ask for clarification? Which tools can it invoke, and with what permissions? How is an invalid plan represented? What happens when a response is incomplete or the network connection is interrupted? The information provided for this analysis does not answer all of these questions. They should not become implementation assumptions simply because a page describes interaction with the physical world.

What can be concluded and what needs verification

TopicSupported conclusionOpen question
Model typeGoogle describes it as a vision-language model for robotics.Exact inputs, supported formats, and requirements for the operating environment.
PlanningThe ER 2 page says it plans multi-step tasks.Specific tasks, success conditions, and reproducible results.
Motor executionThe description separates planning from execution by a lower-level system.Interface, controller, physical constraints, and validation before movement.
API and accessOfficial documentation exists for the Interactions API and generateContent.Current status, identifier, permissions, quotas, and differences between interfaces.
04

Limits, mitigations, and physical safety

The Gemini Robotics ER 2 model card is the relevant source for reviewing known limitations and mitigations. However, the verifiable material available for this article only confirms that the card is specific to ER 2 and that model cards are intended to provide information about limitations and mitigations. It does not allow us to rigorously list specific risks, evaluation conditions, or measures for this version. It would therefore be inappropriate to attribute particular safeguards to ER 2 without examining the complete card.

It is also important to distinguish mitigation from guarantee. A warning, filter, or evaluation may reduce certain risks under particular conditions, but it does not certify the safe behavior of an entire robotics installation. Physical safety depends, among other factors, on the robot, workspace, sensors, speeds, attached tools, and stopping mechanisms. This is an engineering consideration for deployment, not a claim that ER 2’s documentation has validated those dimensions.

A team should keep critical controls outside a free-form instruction to the model. For example, authorization to start a movement, limits on force or speed, and stopping when someone enters a work area require mechanisms whose response can be verified in the deployed system. This recommendation does not mean the model is not useful; it means a generated output should be treated as a proposal that passes through explicit validations before becoming movement.

05

What the claim about Safety Instruction Following and Human Proximity means

Google announced improvements for ER 2 in Safety Instruction Following and Human Proximity compared with ER 1.6 and other models. That comparison should be attributed to the manufacturer. The announcement establishes that Google makes this claim, but the search information available does not provide numerical scores, sample size, exact tasks, operational definitions of the metrics, or execution conditions. Without those elements, it is not possible to determine how much the result improved, whether the differences matter for a particular use case, or whether conditions were comparable across systems.

The names of the categories are not enough to reconstruct the protocol either. “Safety Instruction Following” may refer to a defined task for following safety instructions, while “Human Proximity” suggests an evaluation related to proximity to people; but we should not infer from the names what scenarios, distances, movements, or thresholds were used. The full results page and methodological details would be needed to interpret the metrics.

A responsible evaluation should preserve the distinction between an announcement and verifiable quantitative evidence. The claim is relevant when deciding what questions to ask or which tests to replicate, but it does not support promises of reduced incidents or extrapolation to a different robot. Nor does it justify concluding that ER 2 is safer in every scenario: even if the published comparison is confirmed in detail, it would apply only to the tasks and conditions specified.

A cautious reading of the announced comparison

ElementWhat is knownWhat is missing to assess the result
ComparatorsGoogle mentions ER 1.6 and other models.The identity and versions of all compared models, and whether configurations were equivalent.
CategoriesThe announcement mentions Safety Instruction Following and Human Proximity.Definitions of each task, scoring criteria, and included scenarios.
ResultsGoogle says ER 2 performs better.Scores, sample size, variability, and data that would allow the comparison to be reproduced.
Application to a robotThe claim can inform a local evaluation.Testing on the hardware, sensors, environment, and safety procedures of each deployment.
06

Access, availability, and price: what cannot be confirmed

Google Cloud publishes a Gemini Robotics ER 2 documentation page within Gemini Enterprise Agent Platform, and Google provides documentation for the Interactions API and generateContent interfaces. These references show that official documentation is available for platform and API channels. On their own, they are not enough to confirm current availability for all users, access requirements, the exact identifier to invoke, quotas, or applicable restrictions.

Nor is there a verifiable, specific official price for Gemini Robotics ER 2 available here. The price should not be inferred from rates for other models, a different interface, or a third-party calculation. Cost may depend on the channel, billed units, and applicable terms, but the evidence available for this article does not establish those details. Before preparing a budget, consult the current page for the chosen channel and confirm that the price applies to the specific model and usage mode.

The same caution applies to usage limits and access terms. The existence of documentation does not mean access is open or universal. A team deciding whether it can begin a test should verify the relevant official console or documentation, record the date and region of the check, and confirm the applicable terms. In the absence of that verification, the correct conclusion is that access and price have not been confirmed—not that the model is free, publicly available, or priced identically across platforms.

07

How to design a useful acceptance test

A local test should measure the complete system, not merely judge whether a response sounds reasonable. Start with a bounded task, a controlled environment, and an observable outcome. Define which initial states are valid, which steps are allowed, and which conditions require the test to stop. Keep separate records of the interpretation, proposed plan, validation decision, and executed movement; this makes it possible to locate where a failure originated.

Include variations that matter in the intended environment, such as changes in object position, occlusions, or incomplete instructions, but introduce them gradually and under control. Do not let the first test of an unfamiliar condition involve physical movement with consequences. If the integration supports it, first check the response in an observation or simulation mode, and require human review or deterministic rules for actions that could cause harm. These are evaluation recommendations, not capabilities attributed to ER 2.

Before expanding use, agree on auditable acceptance criteria: completion rate under specified conditions, interpretation errors, plans rejected by the validator, stops, human interventions, and movement deviations. The decision does not need to be reduced to a single average. An infrequent but serious failure can matter more than many correctly completed tasks. Acceptance should include criteria for restricting or withdrawing the system if errors outside the agreed limits are observed.

Practical sequence for a controlled evaluation

  1. 01Define a bounded task and environment; specify initial states, the expected outcome, and stop conditions.
  2. 02Record inputs, interpretation, and plan before allowing movement; retain the data needed to review each decision.
  3. 03Test first under supervision and without physical consequences where the setup allows it; then enable limited movements with independent controls.
  4. 04Introduce variations gradually and record failures, rejections, interventions, and stops—not just completed tasks.
  5. 05Approve only scenarios that meet written criteria; maintain supervision and reassess after changes to the model, API, robot, or environment.
08

Conclusion: testing the model does not certify production performance

The available evidence supports describing Gemini Robotics ER 2 as a vision-language model for robotics to which Google attributes multi-step task planning, separated from motor execution by a lower-level system. It also supports reporting that Google announces better results than ER 1.6 and other models in Safety Instruction Following and Human Proximity. With the information retrieved, however, it does not allow us to reconstruct the protocols or verify scores quantifying those improvements.

For a technical team, this is enough to formulate an evaluation hypothesis, not to make a production decision. Testing must determine whether the model interprets relevant inputs, proposes plans that the system can validate, and integrates with suitable controls for the specific robot. Safety and reliability must be measured in the deployed architecture, with physical limits and supervision procedures defined by the team.

Official platform and API documentation also does not fully clarify availability, terms of use, or applicable pricing. Those details should be verified directly in the current channel before estimating costs or committing to an integration. In short: ER 2 merits a bounded evaluation if its described capabilities fit the problem, but the documentation and benchmark announcement verifiable here do not justify assuming that the model will execute physical tasks reliably in production.

Open questions

  • The current availability status, model identifier, access terms, and applicable restrictions could not be confirmed in detail.
  • A specific official price for Gemini Robotics ER 2 could not be verified.
  • The available information does not detail all inputs, outputs, operations, or differences between the Interactions API and generateContent.
  • No scores, sample size, complete protocol, or comparable conditions are available for the announced improvements in Safety Instruction Following and Human Proximity.
  • The material checked does not allow specific limitations and mitigations from the ER 2 model card to be listed.
  • Physical safety or production performance cannot be inferred without testing the specific robot, its sensors, controller, and environment.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction