SmolVLA fine-tuned on HARP VLA train37

This is the validated final LeRobot policy checkpoint from the joint 37-task HARP VLA fine-tuning run. It predicts absolute Franka joint-position targets plus gripper state (qpos_target_abs, action dimension 8).

Repository: zimplex/harp-vla-train37-smolvla
Final step: 100,000
Final training job: 131711
Final audit job: 131714

Dataset

Field Value
Public source zhouqh/harp_vla_data
Dataset card license MIT at the pinned revision
Pinned source revision 46691c4098ae2ca46a131745d1fe386d63a3fc33
Conversion LeRobot v3, qpos_target_abs, 20 Hz, PyAV video backend
Robot / renderer / schema Franka / NYX / vla_steps_v2_qpos_target
Coverage 37 source tasks, 36 language instructions, 300 episodes/task
Size 11,100 episodes; 4,373,192 frames
Camera payload head RGB + right-wrist RGB; recorded depth groups are empty by contract
Dataset completion marker SHA-256 f9d5121d9f7a210918e5d3dd394d915b00a413d1659538d01729e8f8f6b46db6
Converted data tree SHA-256 65dfc626682a26bfb64478d8b9da9b008aafe2d5536c97b52a21b2c320e8c055
Task manifest SHA-256 b0f46008dfe238c08be3c31f24d86ef80e155e7c5b9b8e75b99c8d73812aa8c6

Training protocol

All five policies used the same optimizer-step and example budget as the earlier HR-Bench train19 sweep. Every policy consumed exactly 6,400,000 training examples (6,400,000 for this model), approximately 1.46 passes over the 4,373,192-frame dataset. This is not equal-epoch training.

Field Value
Model / policy type smolvla
Initialization SmolVLM2-500M architecture/config/tokenizer; VLM weights were not loaded
Pinned base revision or digest 7b375e1b73b11138ff12fe22c8f2822d8fe03467
Nodes × GPUs/node 4 × 8 NVIDIA B200
Batch per process 2
Global batch 64
Optimizer steps 100,000
Save cadence every 5,000 steps
Total examples 6,400,000
Seed 1000
Image augmentation disabled
Mixed precision flag use_amp=false
Data workers 4 per process; prefetch factor 4
cuDNN deterministic false
W&B offline; project harp_vla_train37_gcp
Training budget policy same_optimizer_and_example_budget_as_hr_v2_train19

Important initialization detail: the run referenced the pinned SmolVLM2 repository for architecture/config/tokenizer/processor data, but the emitted policy config records load_vlm_weights=false and pretrained_path=null. It must not be described as a fine-tune of pretrained SmolVLM2 weights.

Pinned software runtime

Component Version / revision
LeRobot source 26ff40ddd784280efc133a8e5af1a76e5ac731c2 plus the recorded HARP entrypoint
Python 3.12
PyTorch / TorchVision 2.8.0 / 0.23.0
Diffusers / Datasets 0.35.2 / 4.4.1
PyAV 15.1.0

Optimizer and scheduler

{
  "optimizer": {
    "betas": [
      0.9,
      0.95
    ],
    "eps": 1e-08,
    "grad_clip_norm": 10,
    "lr": 0.0001,
    "type": "adamw",
    "weight_decay": 1e-10
  },
  "scheduler": {
    "decay_lr": 2.5e-06,
    "num_decay_steps": 30000,
    "num_warmup_steps": 1000,
    "peak_lr": 0.0001,
    "type": "cosine_decay_with_warmup"
  }
}

Policy architecture

Config field Value
n_obs_steps 1
chunk_size 50
n_action_steps 50
max_state_dim 32
max_action_dim 32
resize_imgs_with_padding [512, 512]
tokenizer_max_length 48
num_steps 10
freeze_vision_encoder true
train_expert_only true
train_state_proj true
attention_mode cross_attn
num_vlm_layers 16
expert_width_multiplier 0.75
normalization_mapping {"ACTION": "MEAN_STD", "STATE": "MEAN_STD", "VISUAL": "IDENTITY"}

Inputs and outputs

Feature Direction Type Shape
observation.image input VISUAL 3 × 256 × 256
observation.state input STATE 8
observation.wrist_image input VISUAL 3 × 256 × 256
action output ACTION 8

The exact, machine-readable training and policy configurations are included as train_config.json and config.json. They are the files emitted by the final checkpoint; their hashes were checked against the immutable completion marker.

Files and loading

The repository root contains the complete pretrained_model/ payload from the LeRobot checkpoint (1.12 GiB), including policy weights, processor state, train_config.json, and config.json. checkpoint_manifest.json records every uploaded checkpoint file's byte size and SHA-256; SHA256SUMS provides the same digests in standard text form. training_provenance.json records the immutable run, dataset, source, job, and audit contract.

Optimizer and RNG files from training_state/ are intentionally not published; this public repository is an inference checkpoint, not a resume bundle.

from lerobot.policies.factory import make_policy_from_pretrained

policy = make_policy_from_pretrained("zimplex/harp-vla-train37-smolvla", device="cuda")
policy.eval()

Use a LeRobot checkout compatible with the configuration included here.

Reproducibility and validation

Field Value
Immutable run ID harp_vla_train37_full_deadline_20261001T151243Z_2747a40
Original protocol source commit 2747a401bd29a070452aeab815d99847d7b507bb
Final attempt source commit da1a0456fe16bab22bdf3759b6577b335c683680
HARP race-safe entrypoint SHA-256 bf6ac26d9d4487ffe3258fd482a4c3561d5effaab9fc501442367a159631e88e
Pinned LeRobot trainer SHA-256 5ae8b0b8f3312f14f054364b56d21f39e3843d650162443286e7342bc489d8c8
Training environment marker SHA-256 dd775f736fa809107dc69b653d5c0eef4679e6958fc9b8950fe6aaa3835c2bab
Completion marker SHA-256 d971ca131972729ce6ca9744dd08fccf45f1fad5fa75329d9651f852663947a4
train_config.json SHA-256 f247056f3d2246ecc83de9d072a1d1e6f6c88e26ff2382d5d1a5ab673523a1d4
config.json SHA-256 f9524312b304998a9d357102a88fcb6e95f5d7f43fb930ab02dc6a4c054f6619

The final audit required Slurm COMPLETED/0:0, exact marker/config hashes, the pinned dataset and topology, and finite values in every final Safetensors tensor. This model passed with zero audit errors.

Limitations

  • This release contains training artifacts, not HARP benchmark evaluation scores. No downstream success rate is claimed by this model card.
  • The checkpoint is specific to the converted 20 Hz absolute-qpos convention and its feature/normalization schema.
  • No license is asserted by this card; users must comply with the licenses and terms of the base model/backbone, dataset, LeRobot, and other dependencies.
Downloads last month
18
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Dataset used to train zimplex/harp-vla-train37-smolvla