Instructions to use leninangelov/act_picking_up_a_white_cube_v1.6 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use leninangelov/act_picking_up_a_white_cube_v1.6 with LeRobot:
- Notebooks
- Google Colab
- Kaggle
ACT Picking Up a White Cube — v1.6
An Action Chunking with Transformers (ACT) policy trained with Hugging Face LeRobot to perform a tabletop manipulation task: visually locate, approach, grasp, and lift a white cube.
The policy was trained from demonstrations collected with a 6-DoF follower robot and a front RGB camera. The dataset collection strategy deliberately varied the cube's position and orientation across the robot's reachable tabletop workspace.
Authors: Lenin A. Villanueva Q. and Thomy J. Villanueva Q.
Model Summary
| Property | Value |
|---|---|
| Policy | Action Chunking with Transformers (ACT) |
| Framework | Hugging Face LeRobot |
| Task | Pick up a white cube from a tabletop |
| Published checkpoint | 160,000 training steps |
| Parameters | 51,597,190 (~51.7M) |
| Vision backbone | ResNet-18 |
| Backbone initialization | ImageNet-1K V1 pretrained weights |
| Visual observations | 1 front RGB camera |
| Image resolution | 640 × 480 |
| Robot state dimension | 6 |
| Action dimension | 6 |
| Training dataset | leninangelov/act-training-data-picking-up-the-white-cube-v5.0 |
| Dataset episodes | 54 |
| Dataset frames | 46,898 |
| Dataset control/video rate | 30 FPS |
| Training hardware | NVIDIA GeForce RTX 3070 Laptop GPU |
| Main training device | CUDA |
| License | Apache-2.0 |
The model is task-specific and was trained through offline imitation learning from human-generated robot demonstrations.
Task
The objective of the policy is to execute a complete cube-picking behavior from visual and proprioceptive observations.
At inference time, the model receives:
- one RGB image from the front camera;
- the current 6-dimensional robot state.
It outputs a sequence of 6-dimensional robot actions representing:
shoulder_pan.posshoulder_lift.poselbow_flex.poswrist_flex.poswrist_roll.posgripper.pos
The policy therefore learns the visuomotor mapping from the observed configuration of the cube and robot to the joint/gripper commands required to approach and grasp the cube.
Dataset
The model was trained on:
leninangelov/act-training-data-picking-up-the-white-cube-v5.0
Dataset statistics
| Property | Value |
|---|---|
| Episodes | 54 |
| Frames | 46,898 |
| Tasks | 1 |
| Frame rate | 30 FPS |
| Camera streams | 1 |
| Camera | Front RGB |
| Resolution | 640 × 480 × 3 |
| Robot state | 6 dimensions |
| Actions | 6 dimensions |
| Video codec in dataset | AV1 |
| Audio | None |
Spatial design of the demonstrations
The demonstrations were not collected by repeatedly placing the cube at approximately the same location.
Instead, the reachable tabletop workspace was conceptually divided into six imaginary spatial regions. These regions were used as a collection protocol to encourage broader spatial coverage.
The collection strategy was designed around three principles:
- Spatial coverage: demonstrations were distributed across the six regions of the table rather than concentrated around a single pickup location.
- Approximate balance: the number of cases associated with the different regions was intentionally kept balanced so that one part of the workspace would not dominate the training distribution.
- Within-region variation: cube placement was varied around each conceptual region instead of using a single fixed coordinate.
- Orientation variation: the cube was presented at different orientations, adding variation independently of its position on the table.
The regions were collection aids rather than labels supplied to the ACT policy. The policy receives images and robot state; it is not explicitly told which region contains the cube.
This design was intended to reduce positional memorization and encourage the visual policy to use the observed cube location when selecting its trajectory.
Model Architecture
This model uses the ACT architecture implemented in LeRobot.
Inputs
Visual observation
observation.images.front
shape = [3, 480, 640]
A single RGB front-camera observation is processed by a ResNet-18 visual backbone initialized from ImageNet-1K V1 weights.
Robot state
observation.state
shape = [6]
The state contains the positions of the five arm joints plus the gripper.
Outputs
action
shape = [6]
Each action controls the same six robot dimensions represented in the state.
ACT configuration
| Parameter | Value |
|---|---|
| Observation steps | 1 |
| Action chunk size | 100 |
| Action steps | 100 |
| Transformer model dimension | 512 |
| Attention heads | 8 |
| Feed-forward dimension | 3,200 |
| Transformer encoder layers | 4 |
| Transformer decoder layers | 1 |
| Feed-forward activation | ReLU |
| Dropout | 0.1 |
| Pre-norm | False |
| Vision backbone | ResNet-18 |
| VAE | Enabled |
| VAE latent dimension | 32 |
| VAE encoder layers | 4 |
| KL weight | 10.0 |
| Temporal ensembling | Disabled |
| State normalization | Mean / standard deviation |
| Action normalization | Mean / standard deviation |
| Visual normalization | Mean / standard deviation |
The policy predicts an action chunk of 100 future robot actions from the current observation. With the dataset operating at 30 Hz, 100 control steps correspond nominally to approximately 3.33 seconds of action horizon.
Training
Published checkpoint
This repository represents the checkpoint saved after:
160,000 optimization steps
The surviving training log at the checkpoint reported approximately:
step: 160,000
epoch: 27.29
loss: 0.045
grad norm: 2.417
learning rate: 1.0e-5
Important note about the 200K configuration
The saved train_config.json contains:
steps: 200000
This represents the configured target of the resumed training run, not the checkpoint represented by this model version.
The v1.6 model itself corresponds to the 160,000-step checkpoint. The same training lineage was subsequently continued toward later checkpoints/model versions.
Optimization configuration
| Parameter | Value |
|---|---|
| Optimizer | AdamW |
| Learning rate | 1 × 10⁻⁵ |
| Backbone learning rate | 1 × 10⁻⁵ |
| Weight decay | 1 × 10⁻⁴ |
| Adam β₁ | 0.9 |
| Adam β₂ | 0.999 |
| Adam ε | 1 × 10⁻⁸ |
| Gradient clipping | 10.0 |
| Batch size | 8 |
| Data-loader workers | 4 |
| LR scheduler | None |
| Random seed | 1000 |
| Automatic mixed precision | Disabled |
| cuDNN deterministic mode | Disabled |
| Log frequency | Every 200 steps |
| Checkpoint frequency | Every 5,000 steps |
Image-augmentation transformations were present in the LeRobot configuration but were disabled for this run.
Training environment
Training was performed on a personal laptop equipped with an:
NVIDIA GeForce RTX 3070 Laptop GPU
The project was developed in a Windows/WSL-based LeRobot workflow. The surviving training/resume logs for this checkpoint lineage show the later training stages running from a Windows Conda environment with CUDA enabled.
This is therefore an example of training a real-robot visuomotor Transformer policy on consumer laptop hardware, rather than a datacenter GPU.
Training Time
The model was trained through several resumed sessions rather than one uninterrupted process.
The surviving logs allow the following approximate reconstruction:
| Training segment | Logged process time |
|---|---|
| 0 → 10K steps | ~11 h 17 min |
| 10K → 100K steps | ~9 h 09 min |
| 100K → 115K steps | ~45 min |
| Resume from 110K → 160K | ~2 h 38 min |
The checkpoint at 115K encountered a disk-space error while saving optimizer state. Training was subsequently resumed from the previous valid checkpoint, causing part of the 110K–115K interval to be recomputed.
The recovered logs therefore account for approximately 23 h 49 min of active process time leading to the 160K checkpoint, including the repeated work caused by that interruption.
The first surviving training session started on April 4, 2026, and the 160K checkpoint was saved on April 5, 2026 at approximately 15:23, giving a calendar span of roughly 35 h 10 min including interruptions and inactive periods.
These numbers should be treated as a reconstruction from the original logs, rather than as a controlled GPU-training benchmark.
Inference Configuration
The policy configuration stored with the model provides the following inference characteristics:
| Parameter | Value |
|---|---|
| Device used by policy | CUDA |
| Precision | FP32 |
| AMP | Disabled |
| Front-camera tensor | [3, 480, 640] |
| Robot state | 6-D |
| Policy output | 6-D actions |
| Observation history | 1 step |
| Action chunk | 100 steps |
| Executed action steps configuration | 100 |
| Temporal ensembling | Disabled |
| State normalization | Mean / standard deviation |
| Action normalization | Mean / standard deviation |
The saved LeRobot processor configuration performs the equivalent of:
RGB image + robot state
│
▼
batch construction
│
▼
move tensors to CUDA
│
▼
mean/std normalization
│
▼
ACT policy
│
▼
100-step action chunk
│
▼
action unnormalization
│
▼
move action output to CPU
│
▼
robot execution
The exact historical lerobot-record command used for the physical inference experiment was not preserved in the material currently available, so it is intentionally not reconstructed here from assumptions.
Inference Demo
The following video shows the trained policy operating the physical robot during inference.
The video is intended as a qualitative demonstration of the learned behavior. It should not be interpreted as a statistically controlled success-rate evaluation.
Evaluation
Physical inference was used to inspect the policy behavior on the real robotic setup.
This card does not report a numerical task success rate because a sufficiently reliable aggregate success/failure record for the v1.6 checkpoint could not be reconstructed from the surviving experiment logs.
The inference video above therefore serves as qualitative evidence of the physical behavior of this checkpoint.
For a quantitative evaluation, future reporting should ideally include:
- number of physical trials;
- successful pickups;
- failures to reach the cube;
- grasp failures;
- success rate by tabletop region;
- success rate by cube orientation;
- in-distribution versus held-out positions.
Intended Use
This model is intended for:
- robotics research and education;
- experiments with ACT and imitation learning;
- analysis of visuomotor policies trained from relatively small real-robot datasets;
- tabletop object-manipulation demonstrations;
- reproducibility experiments using Hugging Face LeRobot.
Limitations
The policy was trained for a narrowly defined manipulation task and environment.
Performance may degrade with changes such as:
- substantially different camera viewpoints;
- different table geometry;
- significant lighting changes;
- different cube appearance or dimensions;
- robot calibration changes;
- positions outside the demonstrated workspace;
- occlusion of the cube;
- changes to robot dynamics or control timing.
The spatially distributed data-collection strategy increases variation within the training set but does not make the policy a general-purpose object-manipulation model.
The model should not be used as an unsupervised controller in safety-critical environments.
Reproducibility Notes
The most important configuration values required to reproduce the training setup are stored with the model in the LeRobot configuration files.
For exact reproduction, use:
- the same training dataset;
- the same ACT architecture;
- batch size 8;
- AdamW with learning rate
1e-5; - ResNet-18 ImageNet initialization;
- mean/std normalization;
- 100-step action chunks;
- seed 1000;
- the same 640 × 480 front-camera observation layout.
Because LeRobot APIs and configuration formats can evolve, reproducing the experiment with a newer LeRobot release may require adapting the command-line interface while preserving these model and optimization parameters.
Related Resources
Training dataset
leninangelov/act-training-data-picking-up-the-white-cube-v5.0
Architecture
Action Chunking with Transformers (ACT), introduced in:
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware arXiv:2304.13705
Although ACT was originally introduced for fine-grained manipulation, this model applies the policy architecture to a single-arm tabletop cube-picking task using LeRobot.
Authors
Lenin A. Villanueva Q. Thomy J. Villanueva Q.
Citation
If this model or its associated dataset is useful in your work, the following citation can be used:
@misc{villanueva2026actwhitecube,
author = {Lenin A. Villanueva Q. and Thomy J. Villanueva Q.},
title = {ACT Picking Up a White Cube v1.6},
year = {2026},
publisher = {Hugging Face},
note = {LeRobot ACT policy for real-robot tabletop manipulation}
}
License
Released under the Apache License 2.0.
- Downloads last month
- 84