Manipulation behaviors vary widely across objects and scenes, yet they share a small set of reusable skills. Vision-Language Models describe individual manipulation events well, but can they organize a stream of events into reusable skills? We study this question through:
Video2Skill in two minutes.
Want more detail? Watch the full 6-minute walkthrough.
Music: “Digital Lemonade” by Kevin MacLeod (incompetech.com), licensed under CC BY 4.0.
grasp, transfer, and place; Video 2 reuses all three despite changes in objects and scene; Video 3 reuses grasp and transfer while adding throw for a new transformation.
The library thus preserves shared structure across observations while expanding to accommodate novelty.
Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.
Click to jump to each section.
A model processes an ordered stream of videos in chunks. At each step it receives a new chunk together with its current state: a persistent skill library and an optional unfinished event carried over from the previous chunk. It outputs any new schemas, the completed events in the chunk, and an updated unfinished event:
$$y_t \sim \pi_\theta(\cdot \mid x_t, z_{t-1}), \qquad y_t = (\mathcal{D}_t, \mathcal{E}_t, u_t), \qquad \mathcal{L}_t = \mathcal{L}_{t-1} \cup \mathcal{D}_t.$$
Each event specifies its temporal boundaries, a schema, and argument bindings, e.g. grasp(object, source). A schema has a free-form name, a definition, and typed argument slots; it describes a physical transformation rather than an executable controller. The library starts empty and persists across videos, so every earlier creation and reuse decision shapes how later observations are interpreted.
Creating duplicate schemas fragments recurring experience, while reusing an overly broad schema merges distinct transformations.
Source data. Video2Skill spans human and robot manipulation: HD-EPIC provides long egocentric recordings of unscripted kitchen activity, and RoboInter provides robot tabletop videos drawn from DROID and RH20T. Median reference segments in the verified test splits last 1.9 s and 4.5 s, respectively.
Curation. Annotation prioritizes visual evidence over source labels. Each candidate segment is checked against the video; labels whose stated manipulation is not visually supported are discarded. Retained events are typed skill calls with arguments drawn from seven roles (object, source, destination, instrument, direction, substance, state). Review corrected 241 of 12,429 verified calls, which are organized into 48 canonical transformation classes that define the reference partition. Models never see this class list and name their schemas freely.
Splits. HD-EPIC: 42 training videos (5,726 calls) and 40 test clips from 29 held-out recordings (1,137 segments, 46 classes). RoboInter: 2,992 training videos (5,032 calls) and 132 test videos (534 segments, 15 classes). Training and test splits are disjoint at the source-video level.
Evaluation covers three complementary dimensions:
We evaluate 19 open-source VLMs without task-specific training under two paradigms. Unified models jointly ground events and update the library from visual input; factorized models first propose events without library access, then use the same model for text-only library updates.
We fine-tune Qwen3.5-4B and LLaVA-OneVision-2-8B under both paradigms with three supervised strategies that differ in which library states the model sees during training:
We introduced Video2Skill to study whether models can turn streaming visual experience into a persistent library of reusable skills. Across robot manipulation and egocentric kitchen activity, evaluations of 19 VLMs reveal complementary weaknesses: unified models tend to merge distinct transformations, while factorized models frequently assign recurring transformations to duplicate schemas. Oracle-history SFT, counterfactual library-state rebalancing, and on-policy correction partially alleviate these failures, but do not reliably resolve the boundary between novelty and reuse. Our findings underscore the need to evaluate event coverage and abstraction quality jointly, and establish a testbed for learning coherent skill libraries that remain useful as experience accumulates.
We thank the Cambrian authors for providing this webpage template.
@article{zhang2026video2skill,
title={Video2Skill: From Streaming Experience to Reusable Embodied Skills},
author={Zhang, Jianshu and Zhang, Ce and Yang, Xiyuan and Xu, Chenwei and Lu, Haoran and Li, Yijiang and Xie, Yaqi and Sycara, Katia P. and Liu, Han},
journal={arXiv preprint},
year={2026}
}