REAL ROBOTS · STREAMING UNDERSTANDING

RoboChrono.

A Real Robot Benchmark for
Streaming Task Understanding

The world doesn’t wait for the video to end.
Can a model understand an action as it unfolds?

See the robots in action
OBSERVATION STREAMTIANJI / CAM 01
STACK CUBES 2× PREVIEW
00:00 / —
No future frames. Only the story so far.
THE PEOPLE BEHIND ROBOCHRONO10 affiliations
Yuzhou Wu1,2,*Longteng Fan3,*Zimeng Li4,*Yu Wanchan4,*Ting Zhang2,*Yiyang Ma5,*Shihao Li3,*Wei Ying6Jianbin Qin2Jiajian Jing4Fangwen Chen4Yifan Wu3Zichen Zhang3Ruiqi Yang3Weibin Kong3Yihang Xu3Haoran Liu3Zonghang He7Xuyang Liu8YiFan Xiong9Siteng Huang10Tao Xu3Zhuo Xu3,†Long Chen3,†Ruoxiang Li2,†
* Equal contribution† Corresponding authors
View all affiliations
  1. Tianji Tec.
  2. Shenzhen University
  3. General Intelligence Machine
  4. Huazhong Agricultural University
  5. Yanbian University
  6. South China Agricultural University
  7. Shanghai Jiao Tong University
  8. Hong Kong Polytechnic University
  9. Beijing Jiaotong University
  10. Alibaba Group

REAL-WORLD PLATFORMS.
SHARED RESEARCH.

Bringing temporal understanding
into the physical world.

A WORLD IN MOTION

Small actions.
Endless stories.

Ten glimpses of real manipulation: making tea, packing, cleaning, and arranging. Watch the same world through different hands.

Preview excerpts from yyyyywv/egocentric · CC BY-NC 4.0 · Trimmed and compressed for the web.

01 / THE IDEA

Understanding is
a matter of time.

A single frame tells you what is visible. An unfolding interaction tells you what is happening. RoboChrono asks models to reason from the observations available up to a queried moment, without looking into the future.

39↗Manipulation scenarios
34,713Evaluation instances
4Embodiment settings
7Understanding tasks
ROBOCHRONO AT A GLANCE
Overview of RoboChrono embodiments and streaming task understanding capabilities
OVERVIEW Real manipulation. Diverse embodiments. A shared test of streaming task understanding.

02 / THE BENCHMARK

Seven tasks.
Three ways to reason.

From identifying the current action to locating it in time, each task isolates a different part of temporal understanding. Six tasks use multiple-choice evaluation; one measures action time localization.

01

RECOGNITION

Read the action.
Anticipate the next.

Recognize actions, anticipate the next step, and identify a frame within the observed history.

  • 01 Current Action Recognition
  • 02 Next Action Prediction
  • 03 Goal-Conditioned Next Action Prediction
  • 04 Frame Matching
02

ALIGNMENT

Connect the views.
Recover the order.

Match observations across cameras and reconstruct the chronological sequence of video frames.

  • 05 View Matching
  • 06 Frame Ordering
03

TEMPORAL GROUNDING

Find the action.
Pinpoint it in time.

Identify when an action starts and ends within the observation history, using explicit temporal boundaries.

  • 07 Action Time Localization
Evaluated with R@1 at tIoU ≥ 0.5.

A causal observation boundary. Models see only the visual evidence available up to the query time. Future frames stay out of reach.

FOUR EMBODIMENTS / ONE SHARED TASK

Same world.
Different hands.

Robot grippers, dexterous hands, and egocentric human demonstrations. Explore the shared “Stack Cubes” task across all four embodiment settings.

TIANJI GRIPPER / THREE CAMERA VIEWS

Tianji gripper

Teleoperated gripper demonstrations capture the same execution from a top camera and two wrist-mounted cameras.

14scenarios
01 / Top view
02 / Left wrist
03 / Right wrist
THE SHAPE OF THE DATASET
Distribution of RoboChrono scenarios, embodiments, and manipulation actions
DATASET 1,816 successful videos. Approximately 25 hours of real manipulation.

04 / WHAT WE LEARNED

Seeing isn’t the same
as understanding time.

Zero-shot evaluation of 18 vision-language models reveals a capability gap: strong visual matching does not consistently translate into strong temporal ordering.

VISUAL MATCHING → TEMPORAL ORDERING

The temporal gap.

Accuracy on Frame Matching vs. Frame Ordering.

GPT-6-Astra
Frame Matching
98.3%
Frame Ordering
68.3%
RynnBrain1.1-122B-A10B
Frame Matching
95.4%
Frame Ordering
32.9%

Paper-reported results (Table 3). Evaluation coverage and aggregation vary by model; see the paper for details.

WHAT IF THE MODEL CANNOT SEE?

Different tasks.
Different evidence.

Accuracy change when visual observations are removed.

−22.1pp

Current Action
Recognition

−0.7pp

Next Action
Prediction

Next-action prediction can draw on task and action priors, even when visual evidence is unavailable.

Mean drops across five open-weight models on a balanced 312-stem subset; percentage points (pp).

CAPABILITY PROFILES & GOAL CONDITIONING
Paper plots comparing recognition and alignment, frame matching and ordering, and goal-conditioned prediction
ANALYSIS Replotted from Table 3: capability profiles and goal conditioning across 18 models. Reported aggregation is retained; evaluation coverage varies by model.
PLATFORMS & EVIDENCE ABLATIONS
Paper plots comparing GIM and Tianji performance and the effects of removing or shuffling visual inputs
ABLATIONS How platform, visual evidence, and temporal order affect model performance.
Full paper

05 / BUILD WITH ROBOCHRONO

Give your model
a sense of time.

Explore the data, examine the evaluation, and build toward better understanding of real-world robot interaction.

CITE ROBOCHRONO

@misc{wu2026robochrono,
  title = {RoboChrono: A Real Robot Benchmark
           for Streaming Task Understanding},
  author = {Yuzhou Wu and Longteng Fan and
    Zimeng Li and Yu Wanchan and Ting Zhang and
    Yiyang Ma and Shihao Li and Wei Ying and
    Jianbin Qin and Jiajian Jing and Fangwen Chen and
    Yifan Wu and Zichen Zhang and Ruiqi Yang and
    Weibin Kong and Yihang Xu and Haoran Liu and
    Zonghang He and Xuyang Liu and YiFan Xiong and
    Siteng Huang and Tao Xu and Zhuo Xu and
    Long Chen and Ruoxiang Li},
  year = {2026},
  note = {Public preprint},
  url = {https://github.com/mfan-res/ROBOCHRONO/blob/main/docs/paper/RoboChrono.pdf}
}