Robotics
LeRobot
Safetensors
act
imitation-learning
action-chunking-transformer
manipulation
visual-robotics
robot-learning

ACT Picking Up a White Cube — v1.6

An Action Chunking with Transformers (ACT) policy trained with Hugging Face LeRobot to perform a tabletop manipulation task: visually locate, approach, grasp, and lift a white cube.

The policy was trained from demonstrations collected with a 6-DoF follower robot and a front RGB camera. The dataset collection strategy deliberately varied the cube's position and orientation across the robot's reachable tabletop workspace.

Authors: Lenin A. Villanueva Q. and Thomy J. Villanueva Q.

Model Summary

Property Value
Policy Action Chunking with Transformers (ACT)
Framework Hugging Face LeRobot
Task Pick up a white cube from a tabletop
Published checkpoint 160,000 training steps
Parameters 51,597,190 (~51.7M)
Vision backbone ResNet-18
Backbone initialization ImageNet-1K V1 pretrained weights
Visual observations 1 front RGB camera
Image resolution 640 × 480
Robot state dimension 6
Action dimension 6
Training dataset leninangelov/act-training-data-picking-up-the-white-cube-v5.0
Dataset episodes 54
Dataset frames 46,898
Dataset control/video rate 30 FPS
Training hardware NVIDIA GeForce RTX 3070 Laptop GPU
Main training device CUDA
License Apache-2.0

The model is task-specific and was trained through offline imitation learning from human-generated robot demonstrations.

Task

The objective of the policy is to execute a complete cube-picking behavior from visual and proprioceptive observations.

At inference time, the model receives:

  • one RGB image from the front camera;
  • the current 6-dimensional robot state.

It outputs a sequence of 6-dimensional robot actions representing:

  1. shoulder_pan.pos
  2. shoulder_lift.pos
  3. elbow_flex.pos
  4. wrist_flex.pos
  5. wrist_roll.pos
  6. gripper.pos

The policy therefore learns the visuomotor mapping from the observed configuration of the cube and robot to the joint/gripper commands required to approach and grasp the cube.

Dataset

The model was trained on:

leninangelov/act-training-data-picking-up-the-white-cube-v5.0

Dataset statistics

Property Value
Episodes 54
Frames 46,898
Tasks 1
Frame rate 30 FPS
Camera streams 1
Camera Front RGB
Resolution 640 × 480 × 3
Robot state 6 dimensions
Actions 6 dimensions
Video codec in dataset AV1
Audio None

Spatial design of the demonstrations

The demonstrations were not collected by repeatedly placing the cube at approximately the same location.

Instead, the reachable tabletop workspace was conceptually divided into six imaginary spatial regions. These regions were used as a collection protocol to encourage broader spatial coverage.

The collection strategy was designed around three principles:

  • Spatial coverage: demonstrations were distributed across the six regions of the table rather than concentrated around a single pickup location.
  • Approximate balance: the number of cases associated with the different regions was intentionally kept balanced so that one part of the workspace would not dominate the training distribution.
  • Within-region variation: cube placement was varied around each conceptual region instead of using a single fixed coordinate.
  • Orientation variation: the cube was presented at different orientations, adding variation independently of its position on the table.

The regions were collection aids rather than labels supplied to the ACT policy. The policy receives images and robot state; it is not explicitly told which region contains the cube.

This design was intended to reduce positional memorization and encourage the visual policy to use the observed cube location when selecting its trajectory.

Model Architecture

This model uses the ACT architecture implemented in LeRobot.

Inputs

Visual observation

observation.images.front
shape = [3, 480, 640]

A single RGB front-camera observation is processed by a ResNet-18 visual backbone initialized from ImageNet-1K V1 weights.

Robot state

observation.state
shape = [6]

The state contains the positions of the five arm joints plus the gripper.

Outputs

action
shape = [6]

Each action controls the same six robot dimensions represented in the state.

ACT configuration

Parameter Value
Observation steps 1
Action chunk size 100
Action steps 100
Transformer model dimension 512
Attention heads 8
Feed-forward dimension 3,200
Transformer encoder layers 4
Transformer decoder layers 1
Feed-forward activation ReLU
Dropout 0.1
Pre-norm False
Vision backbone ResNet-18
VAE Enabled
VAE latent dimension 32
VAE encoder layers 4
KL weight 10.0
Temporal ensembling Disabled
State normalization Mean / standard deviation
Action normalization Mean / standard deviation
Visual normalization Mean / standard deviation

The policy predicts an action chunk of 100 future robot actions from the current observation. With the dataset operating at 30 Hz, 100 control steps correspond nominally to approximately 3.33 seconds of action horizon.

Training

Published checkpoint

This repository represents the checkpoint saved after:

160,000 optimization steps

The surviving training log at the checkpoint reported approximately:

step:      160,000
epoch:     27.29
loss:      0.045
grad norm: 2.417
learning rate: 1.0e-5

Important note about the 200K configuration

The saved train_config.json contains:

steps: 200000

This represents the configured target of the resumed training run, not the checkpoint represented by this model version.

The v1.6 model itself corresponds to the 160,000-step checkpoint. The same training lineage was subsequently continued toward later checkpoints/model versions.

Optimization configuration

Parameter Value
Optimizer AdamW
Learning rate 1 × 10⁻⁵
Backbone learning rate 1 × 10⁻⁵
Weight decay 1 × 10⁻⁴
Adam β₁ 0.9
Adam β₂ 0.999
Adam ε 1 × 10⁻⁸
Gradient clipping 10.0
Batch size 8
Data-loader workers 4
LR scheduler None
Random seed 1000
Automatic mixed precision Disabled
cuDNN deterministic mode Disabled
Log frequency Every 200 steps
Checkpoint frequency Every 5,000 steps

Image-augmentation transformations were present in the LeRobot configuration but were disabled for this run.

Training environment

Training was performed on a personal laptop equipped with an:

NVIDIA GeForce RTX 3070 Laptop GPU

The project was developed in a Windows/WSL-based LeRobot workflow. The surviving training/resume logs for this checkpoint lineage show the later training stages running from a Windows Conda environment with CUDA enabled.

This is therefore an example of training a real-robot visuomotor Transformer policy on consumer laptop hardware, rather than a datacenter GPU.

Training Time

The model was trained through several resumed sessions rather than one uninterrupted process.

The surviving logs allow the following approximate reconstruction:

Training segment Logged process time
0 → 10K steps ~11 h 17 min
10K → 100K steps ~9 h 09 min
100K → 115K steps ~45 min
Resume from 110K → 160K ~2 h 38 min

The checkpoint at 115K encountered a disk-space error while saving optimizer state. Training was subsequently resumed from the previous valid checkpoint, causing part of the 110K–115K interval to be recomputed.

The recovered logs therefore account for approximately 23 h 49 min of active process time leading to the 160K checkpoint, including the repeated work caused by that interruption.

The first surviving training session started on April 4, 2026, and the 160K checkpoint was saved on April 5, 2026 at approximately 15:23, giving a calendar span of roughly 35 h 10 min including interruptions and inactive periods.

These numbers should be treated as a reconstruction from the original logs, rather than as a controlled GPU-training benchmark.

Inference Configuration

The policy configuration stored with the model provides the following inference characteristics:

Parameter Value
Device used by policy CUDA
Precision FP32
AMP Disabled
Front-camera tensor [3, 480, 640]
Robot state 6-D
Policy output 6-D actions
Observation history 1 step
Action chunk 100 steps
Executed action steps configuration 100
Temporal ensembling Disabled
State normalization Mean / standard deviation
Action normalization Mean / standard deviation

The saved LeRobot processor configuration performs the equivalent of:

RGB image + robot state
        │
        ▼
batch construction
        │
        ▼
move tensors to CUDA
        │
        ▼
mean/std normalization
        │
        ▼
ACT policy
        │
        ▼
100-step action chunk
        │
        ▼
action unnormalization
        │
        ▼
move action output to CPU
        │
        ▼
robot execution

The exact historical lerobot-record command used for the physical inference experiment was not preserved in the material currently available, so it is intentionally not reconstructed here from assumptions.

Inference Demo

The following video shows the trained policy operating the physical robot during inference.

The video is intended as a qualitative demonstration of the learned behavior. It should not be interpreted as a statistically controlled success-rate evaluation.

Evaluation

Physical inference was used to inspect the policy behavior on the real robotic setup.

This card does not report a numerical task success rate because a sufficiently reliable aggregate success/failure record for the v1.6 checkpoint could not be reconstructed from the surviving experiment logs.

The inference video above therefore serves as qualitative evidence of the physical behavior of this checkpoint.

For a quantitative evaluation, future reporting should ideally include:

  • number of physical trials;
  • successful pickups;
  • failures to reach the cube;
  • grasp failures;
  • success rate by tabletop region;
  • success rate by cube orientation;
  • in-distribution versus held-out positions.

Intended Use

This model is intended for:

  • robotics research and education;
  • experiments with ACT and imitation learning;
  • analysis of visuomotor policies trained from relatively small real-robot datasets;
  • tabletop object-manipulation demonstrations;
  • reproducibility experiments using Hugging Face LeRobot.

Limitations

The policy was trained for a narrowly defined manipulation task and environment.

Performance may degrade with changes such as:

  • substantially different camera viewpoints;
  • different table geometry;
  • significant lighting changes;
  • different cube appearance or dimensions;
  • robot calibration changes;
  • positions outside the demonstrated workspace;
  • occlusion of the cube;
  • changes to robot dynamics or control timing.

The spatially distributed data-collection strategy increases variation within the training set but does not make the policy a general-purpose object-manipulation model.

The model should not be used as an unsupervised controller in safety-critical environments.

Reproducibility Notes

The most important configuration values required to reproduce the training setup are stored with the model in the LeRobot configuration files.

For exact reproduction, use:

  • the same training dataset;
  • the same ACT architecture;
  • batch size 8;
  • AdamW with learning rate 1e-5;
  • ResNet-18 ImageNet initialization;
  • mean/std normalization;
  • 100-step action chunks;
  • seed 1000;
  • the same 640 × 480 front-camera observation layout.

Because LeRobot APIs and configuration formats can evolve, reproducing the experiment with a newer LeRobot release may require adapting the command-line interface while preserving these model and optimization parameters.

Related Resources

Training dataset

leninangelov/act-training-data-picking-up-the-white-cube-v5.0

Architecture

Action Chunking with Transformers (ACT), introduced in:

Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware arXiv:2304.13705

Although ACT was originally introduced for fine-grained manipulation, this model applies the policy architecture to a single-arm tabletop cube-picking task using LeRobot.

Authors

Lenin A. Villanueva Q. Thomy J. Villanueva Q.

Citation

If this model or its associated dataset is useful in your work, the following citation can be used:

@misc{villanueva2026actwhitecube,
  author       = {Lenin A. Villanueva Q. and Thomy J. Villanueva Q.},
  title        = {ACT Picking Up a White Cube v1.6},
  year         = {2026},
  publisher    = {Hugging Face},
  note         = {LeRobot ACT policy for real-robot tabletop manipulation}
}

License

Released under the Apache License 2.0.

Downloads last month
84
Safetensors
Model size
51.7M params
Tensor type
F32
·
Video Preview
loading

Dataset used to train leninangelov/act_picking_up_a_white_cube_v1.6

Paper for leninangelov/act_picking_up_a_white_cube_v1.6