How We Trained Robot 2.1

Robot 2.1 is the internal architecture we use to make our AI workers much faster without making the main model smaller or less capable.

The basic idea is simple: intelligence should not run at one speed.

You don’t think about your feet

When you walk across a room, you make one conscious decision: where you want to go.

You do not think about your ankles. You do not plan the exact pressure under each foot. You do not run a careful calculation every time your body makes a tiny balance correction. That would be impossibly slow.

Instead, your brain runs in layers.

One slow, powerful loop decides the goal: walk over there. A faster loop turns that goal into movement. Even faster loops handle balance, timing, and the small corrections that keep you upright. Most of the work happens below conscious thought, but it is still intelligent work. It is just the right kind of intelligence at the right speed.

We built Robot 2.1 around that same idea.

Big models are smart, but they are slow at small things

Large models are very good at understanding what is happening on a screen and deciding what to do next. That is the part we want them to do. They should look at the task, understand the page, reason about the goal, and choose the right action.

But after that decision has been made, a lot of the output is mechanical.

For an AI worker, the model might need to produce structured text for an action, a click, a coordinate, or a short reasoning trace. In a normal system, every single token requires another full pass through the large model. That means we are asking the expensive, powerful brain to wake up again and again for work that is often already decided.

That is like using your entire conscious mind to decide how to move each toe while walking.

It works. It is just wasteful.

Robot 2.1 adds faster loops underneath

Robot 2.1 keeps the large model as the main decision-maker. We do not replace it. We do not distill it into a weaker model. We keep the strong model in charge.

Then we attach smaller, faster sidemodels around it.

The large model runs the slow loop. It sees the screen, understands the task, and decides the direction. It updates every few steps.

The first sidemodel runs a faster loop. It reads the internal state of the large model and predicts the next block of output tokens. In other words, it learns to continue the large model’s motion for a few steps when the direction is already clear.

Smaller loops can sit below that. A click decoder can specialize in actions. A vision encoder can specialize in the visual details that matter for a UI. A model head can handle formatting or coordinates. Each loop is simpler than the one above it, but much faster.

This is why we call it a side-model architecture. The sidemodels are not competing with the main model. They sit beside it, learn from it, and handle the work that does not need the full slow loop every time.

The important part is that these loops are gated by confidence. The sidemodel is only allowed to move ahead when it is confident. If it becomes unsure, control goes back to the large model.

So the system behaves like a human body: automatic when the path is clear, deliberate when the situation changes.

How we trained it

Training Robot 2.1 was not just a matter of teaching a small model to copy a large model.

We started by running the large model on real AI-worker traces: actual browser tasks, screenshots, multi-step sessions, and action outputs. During those runs, we recorded not only the final text, but also internal hidden states from the large model. Those hidden states are useful because they contain the model’s unfinished thought before it becomes text.

The sidemodel learned from those traces. Given the large model’s internal state, it learned to predict the next block of output.

That gave us a good first version, but there is a common problem with this kind of training: the sidemodel has only seen perfect trajectories from the large model. Once it starts running inside the real system, it visits slightly different states. Small differences compound.

The fix was to train it on its own trajectory.

We let Robot 2.1 run live, with the sidemodel actually making accepted predictions. Then we collected the states it reached and trained it again against the large model’s answer in those exact states. We repeated that loop: run, collect, retrain, measure, repeat.

That was the main unlock. Each round made the sidemodel more confident on the states it really visits, so it could safely predict longer blocks.

What changed

At the start, the sidemodel could only predict a few tokens per large-model step reliably.

After iterative training, Robot 2.1 reached the range we were aiming for: 10-14+ tokens per step in sustained generation, with the large model still acting as the source of truth.

On short, decision-heavy agent turns, where the main model genuinely needs to stay involved, we see meaningful wall-clock speedups while preserving action quality. On longer stretches, where the path is clear and the fast loop can keep moving, the speedup is much larger.

In the full stack, where the same idea is applied not only to text generation but also to action decoders, vision encoders, and task-specific control loops, the efficiency gain can be much bigger. For the kinds of robotic and computer-control workloads we care about, this points toward systems that can be up to 14x more efficient than a naive one-token-at-a-time setup.

The most important result is not just that it is faster. It is that the system gets faster in a controlled way. The sidemodel does not replace the large model’s judgment. It accelerates the parts of the process where the judgment has already been made.

Why this matters for real-world agents

A lot of AI infrastructure tries to make systems faster by making the model smaller.

That can be useful, but it comes with a cost: smaller models are less capable. For real work, especially visual browser work and robotic control, we do not want to throw away the intelligence of the strong model. We want to use it more efficiently.

Robot 2.1 is a different approach.

Keep the strong model. Add fast specialist loops around it.

That means we can add a small decoder for clicks. Or a fast vision encoder for a specific screen environment. Or a structured-output head for a particular action space. The main model can remain general and powerful, while the lower loops become very good at the repeated mechanics of the job.

This is especially natural for robotics and computer control. A robot does not need a giant reasoning model to decide every motor adjustment. An AI worker does not need a giant reasoning model to open a new browser tab. The high-level loop should decide. The lower loops should execute.

That is the part we find most interesting: Robot 2.1 is not just an optimization trick. It is a way to make AI systems more like useful machines. General where they need to be general. Specialized where specialization pays off. Fast where speed matters.

The bigger picture

We think this is one of the right ways to build practical AI systems.

Not one model doing everything at one speed.

A hierarchy:

  • slow loops for reasoning,
  • faster loops for action,
  • very fast loops for reflexes,
  • and clean interfaces between them.

That is how humans work. It is also how efficient machines should work.

Robot 2.1 is our current version of that idea. It lets our AI workers think with a powerful model, but act with fast learned reflexes around it. The result is a system that is more efficient, more modular, and better suited to real-world control.

The large model decides where to walk.

The sidemodels move the feet.