CUA-S1 Forge RLCD-v3
RLCD-v3 is a verifier-guided fine-tune of the CUA-S1 one-step form-action classifier. The objective uses exact expected reward for one-hot verifier labels, cross-entropy anchoring on the natural action distribution, and square-root inverse-frequency weighting on the reward term.
Evaluation
Evaluation used the deterministic synthetic smoke-600 dataset, with form signatures separated across splits. The held-out test partition contains 1,821 examples; the run used seed 20260924 on a Google Colab NVIDIA T4.
| Test metric | Baseline | RLCD-v3 | Change |
|---|---|---|---|
| Top-1 accuracy | 51.6749% | 52.3339% | +0.659 pp |
| Negative log-likelihood | 1.6714 | 1.6191 | -0.0523 |
| Expected calibration error | 0.0355 | 0.0429 | +0.0074 |
| Macro action accuracy | 47.9940% | 48.7460% | +0.752 pp |
| Action | Baseline accuracy | RLCD-v3 accuracy |
|---|---|---|
| Check | 92.7536% | 94.2029% |
| Click | 0.0000% | 0.0000% |
| Fill | 1.2563% | 4.3970% |
| Skip | 97.9661% | 96.3842% |
Top-1 accuracy and NLL improved on this split. Calibration error increased, click accuracy remained zero, and skip accuracy declined.
Training configuration
- Four epochs from the supervised baseline
- Learning rate:
1e-4 - Batch size:
128 - Expected-reward coefficient:
0.10 - Checkpoint selection: validation NLL, with balanced macro accuracy as a secondary criterion
Scope and reproducibility
This evaluation measures one-step candidate-action classification on synthetic form episodes. It does not measure autonomous navigation, live browser execution, or end-to-end task success. The test result is from one fixed seed and one held-out synthetic split.
The source repository and full report include the training implementation and evaluation details.
- Downloads last month
- 68