EN / 中文

Self-Evolving Embodied AI for Scientific Scenarios: Data Centers Transition to Teleoperation Centers

by burixin·September 29, 2026

In his keynote speech titled "Self-Evolving Embodied AI for Scientific Scenarios," Dr. Pan Yuxi, CTO of Synkrotron, did not dwell on grand visions of the future of embodied AI or Artificial General Intelligence (AGI). Instead, he emphasized how to achieve early deployment in a vertical domain, establish a data closed loop as soon as possible, and thereby enable continuous iteration.

When discussing data collection centers, his exact words were:

"The current business model of data collection centers is being questioned and is gradually cooling down. We will see that operational data from real-world scenarios, which holds greater value and provides better auxiliary capabilities for training, will become more prominent. We predict that these data collection centers may eventually evolve into teleoperation centers—remote operation hubs similar to the remote driving facilities used for Robotaxis. By using humans as a fallback to generate closed-loop feedback data, we can support the iteration of the embodied brain."

This analogy is not purely imaginative; the autonomous driving industry has already paved this way:

Waymo's vehicles are equipped with a Fleet Response backend mechanism. When encountering construction zones, police hand signals, or unconventional obstacles, they send assistance requests to remote support personnel. Similarly, Baidu Apollo's Apollo Go and various Robotaxi operators have all established remote safety operator stations.

The value of remote safety operators has never been about their "high driving skills," but rather that every takeover they intervene in represents a precious Corner Case sample.

The same logic applies to the robotics industry. A remote operation center is not about having humans do the work for machines; it is about letting machines do the work while humans are responsible for catching the parts that fall through the cracks.

Human-in-the-Loop Fallback + Data

Dr. Pan provided two reasons for the trend of data collection centers evolving into remote operation centers:

Reason 1: For robots to truly be deployed, human fallback is essential.

Synkrotron integrates data, the brain, and the operating system into a single composite robot, supporting dexterous manipulation in laboratory scenarios—such as grasping transparent objects, and chemical experiment actions like pouring, stirring, and filtering liquids.

They have summarized approximately 50 atomic actions. Each action requires only dozens of post-training teleoperation data samples to achieve relatively good deployment, reaching a success rate of around 95%.

In many scenarios, 95% is already excellent. However, in scientific experiment scenarios, especially those involving pharmaceutical pre-processing, customers expect a success rate close to 100%.

So what about the remaining 5%?

"At this point, we need to use a human-in-the-loop feedback mechanism, where remote teleoperators serve as the human fallback to achieve rapid capability growth," said Pan Yuxi.

Reason 2: Data fed back from the field is the lifeline for robot iteration.

Compared to data collected from staged setups in laboratories, Pan Yuxi believes that operational data from real-world scenarios provides better auxiliary capabilities for training.

A sentence he repeatedly emphasized during his sharing is worth extracting separately:

"In operational scenarios, the trajectory data we generate includes correction data and failure data. This data is more representative than teleoperation data or ego data used in cold starts. It can reveal exactly in which scenarios the robot performs poorly, and in which scenarios it can improve its capabilities through online data generation and feedback learning."

The 100% successful trajectories demonstrated repeatedly in staged laboratories record the correct approaches; whereas the data from every takeover and repair by remote operators in real-world positions records the experience of failures.

The former outlines the center of the distribution, while the latter sketches the boundaries of the distribution. What ultimately determines whether a model can be delivered has always been the boundaries, not the center.

Human-in-the-Loop Circuit: Agent Harness is the Protagonist

This theoretical logic is sound, but what is missing in practice is the engineering framework that can truly put the "human fallback" into action.

Pan Yuxi repeatedly emphasized that beyond the core brain, what is more important is the Agent Harness established around it.

The Agent itself is just a model; the Harness is the scaffolding that turns the model into a usable system—comprising toolboxes, memory, routing, permissions, validation, and rollback.

When applied to robotics, it refers to the entire layer of scheduling and governance framework built around the embodied brain.

This harness structure consists of three layers:

The first layer is the atomic skill library. Based on segmentation tools like SAM (Segment Anything Model) and several small-scale VLAs (Vision-Language-Action models), a batch of "atomic actions" is generated and organized into a skill library—these are the smallest executable units that can be directly invoked, combined, and parameterized.

The second layer is high-level planning and orchestration. More advanced Vision-Language Models (VLMs) like GPT are responsible for overall task planning and skill invocation. A key design here is the decoupling of cognition and execution: the cognitive layer answers "what to do next," while the execution layer answers "how exactly the arm should move and with how much force." The former requires world knowledge and reasoning capabilities, while the latter is handled by expert models like the Pi series action expert. Meanwhile, historical experience—how an action was decomposed last time and whether it was done well—is abstracted into a RAG (Retrieval-Augmented Generation) memory library, preventing the system from having to understand the task from scratch every time.

The third layer is the human fallback entry point. In situations where the robot cannot complete the task autonomously, the system will invoke remote human-in-the-loop teleoperation. This layer is not just a "task safety net"; it is simultaneously a critical action for data collection.

But what truly enables this third layer to "act only when necessary" is an even more crucial design in Pan Yuxi's words—the semantic validator.

"How do you determine whether the current task is completed, not completed, failed, or if an abnormal state has occurred requiring human intervention?"

This is not a technical issue, but a governance issue. Without a validator, there are no trigger conditions for human-in-the-loop—either humans have to stare at the screen constantly, making the cost so high that it loses its meaning; or the robot runs through to the end with errors, losing the opportunity for correction.

Thus, the complete closed loop consists of these seven steps:

Autonomous execution → VLM validation → Problem diagnosis → Self-repair → Human intervention → In-Context learning → Experience consolidation

The brilliance lies in the sixth step: the data from human correction is not just accumulated for the next offline training session, but immediately enters the context window of the current task, allowing the model to learn how to correct it within minutes.

This is the so-called "self-evolution."

The technical foundation supporting this has already emerged. NVIDIA's RoboTTT (Test-Time-Training Robot Policies) integrates the TTT layer into the VLA, turning the model's recurrent state into a set of "fast weights." It scales the visuomotor context to 8,000 time steps, which is three orders of magnitude higher than the best policies at the time, without inference latency increasing with the context length; on real-robot long-horizon manipulation tasks, it achieves an overall performance improvement of 87% compared to the single-step context baseline.

In other words, a human takes over once, the machine provides immediate feedback, and benefits from it in the current run—this is no longer just a vision technically.

Another fundamental base supporting this real-time link is the open-source project Dora, initiated by Synkrotron at the OpenAtom Foundation in 2022. In terms of data transmission latency, Dora is roughly 1/17 to 1/20 of that of ROS (Robot Operating System).

An order-of-magnitude difference in latency turns "training while working" from impossible to possible. This is also the prerequisite for full-duplex human-in-the-loop to be viable: a robot's slipping gripper can afford to wait for a manual command like "grip tighter."

Finally, the entire mechanism operates on three time scales: second-level real-time correction, day-level memory consolidation and skill library updates, and week/month-level offline retraining for new versions.

"Only through this kind of validation do the validated correction trajectories become valuable data that is crucial for model training," said Pan Yuxi.