CUA-S1 Forge RLCD-v3

RLCD-v3 is a verifier-guided fine-tune of the CUA-S1 one-step form-action classifier. The objective uses exact expected reward for one-hot verifier labels, cross-entropy anchoring on the natural action distribution, and square-root inverse-frequency weighting on the reward term.

Evaluation

Evaluation used the deterministic synthetic smoke-600 dataset, with form signatures separated across splits. The held-out test partition contains 1,821 examples; the run used seed 20260924 on a Google Colab NVIDIA T4.

Test metric Baseline RLCD-v3 Change
Top-1 accuracy 51.6749% 52.3339% +0.659 pp
Negative log-likelihood 1.6714 1.6191 -0.0523
Expected calibration error 0.0355 0.0429 +0.0074
Macro action accuracy 47.9940% 48.7460% +0.752 pp
Action Baseline accuracy RLCD-v3 accuracy
Check 92.7536% 94.2029%
Click 0.0000% 0.0000%
Fill 1.2563% 4.3970%
Skip 97.9661% 96.3842%

Top-1 accuracy and NLL improved on this split. Calibration error increased, click accuracy remained zero, and skip accuracy declined.

Training configuration

  • Four epochs from the supervised baseline
  • Learning rate: 1e-4
  • Batch size: 128
  • Expected-reward coefficient: 0.10
  • Checkpoint selection: validation NLL, with balanced macro accuracy as a secondary criterion

Scope and reproducibility

This evaluation measures one-step candidate-action classification on synthetic form episodes. It does not measure autonomous navigation, live browser execution, or end-to-end task success. The test result is from one fixed seed and one held-out synthetic split.

The source repository and full report include the training implementation and evaluation details.

Downloads last month
68
Safetensors
Model size
706k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support