cua-s1.js weights
Browser-ready ONNX exports of Cua's cua-s1 form-filling decision models, for @ai-ecoverse/cua-s1.js and onnxruntime-web. The model, its training data and the planning rules are Cua's work (MIT); this repo only holds the converted checkpoint.
import * as ort from "onnxruntime-web/wasm";
import { loadCuaS1, extractEntities } from "@ai-ecoverse/cua-s1.js";
const model = await loadCuaS1("https://huggingface.co/ai-ecoverse/cua-s1.js/resolve/main/cua-s1-forms", { ort });
const plan = await model.plan("Northwind Clinic - New Patient Registration",
[{ role: "Edit", label: "Phone number", value: "", token: "phone" }],
extractEntities("Tel: (503) 555-0142\nWork phone: (503) 555-0110"));
Contents
| Folder | Source checkpoint | ONNX | Max |Δp| vs PyTorch |
|---|---|---|---|
cua-s1-4b-0.2 |
cua-ai/cua-s1-4b-0.2 @ 1681886 |
4.7 GB | 6.2e-02 over 60 tasks |
cua-s1-4b-0.2-multimodal |
cua-ai/cua-s1-4b-0.2 @ 1681886 |
4.7 GB | 1.0e-01 over 60 tasks |
cua-s1-forms |
cua-ai/cua-s1-forms @ f54adbf |
3.1 MB | 3.2e-06 over 1048 decisions |
Each folder has a manifest.json naming its source commit, the ONNX graphs' SHA-256s (checked by the loader) and
the parity measured at export time. The graphs and Cua's original JSON sidecar (checkpoint.json: architecture,
tensor signature, training metadata) live under r-<commit>/, so publishing a new checkpoint never changes a file
an older manifest points at.
Conversion
The checkpoint is loaded with cua_s1.model.load_checkpoint, which validates the SHA-256 tensor signature. It is
exported unmodified with the torch.export-based ONNX exporter (opset 18, dynamic batch, context, option and
option-token axes). It takes the byte tensors cua_s1's ByteCollator produces and returns per-option logits and
probabilities. onnxruntime matches PyTorch to within the max |Δp| in the table, on 1,048 decisions from Cua's own
synthetic episode generator, with no argmax flips.
model-shared-options.onnx runs the same modules for a batch whose rows all share one option list, which is every
form plan: the options are encoded once instead of once per element. Its inputs drop the batch axis from the option
tensors. It gives the same probabilities as model.onnx and is checked against the unmodified PyTorch forward on the
same episodes.
cua-s1 is a research checkpoint trained on synthetic forms. Read Cua's model card and SECURITY.md before relying on it.
cua-s1-4b
cua-s1-4b-0.2/ is cua-ai/cua-s1-4b-0.2's text adapter (a rank-16 LoRA) merged in fp32 into
Qwen/Qwen3.5-4B and exported with the onnxruntime-genai model builder: int8 weights, fp32
activations, WebGPU. The graph returns hidden states; Qwen3.5-4B ties its output layer to the embeddings, so the
option letters' logits are the final hidden state times head.safetensors (the 26 letter rows, fp32). Weights are
split into files of at most 32 MB.
import * as ort from "onnxruntime-web/webgpu";
import { elementDecisions, loadCuaS1FourB } from "@ai-ecoverse/cua-s1.js/4b";
const model = await loadCuaS1FourB("https://huggingface.co/ai-ecoverse/cua-s1.js/resolve/main/cua-s1-4b-0.2", { ort });
const r = await model.score(options, { app, taskFamily, goal, axTree });
elementDecisions(r.options); // per element, its likeliest action
Against Cua's own cua_s1.four_b.FourBModel (fp32 PyTorch) on 60 real Word, Excel and PowerPoint steps from
GUI-360's test split: max |Δp| 0.0615, 3 argmax flips. Every
element's likeliest action is the expected one on 51/60 of them with this export, as cua-bench-s1 scores it.
Licensing: the adapter and Qwen3.5-4B are both Apache-2.0.
cua-s1-4b
cua-s1-4b-0.2-multimodal/ is cua-ai/cua-s1-4b-0.2's screenshot (multimodal) adapter (a rank-16 LoRA) merged in fp32 into
Qwen/Qwen3.5-4B and exported with the onnxruntime-genai model builder: int8 weights, fp32
activations, WebGPU. The graph returns hidden states; Qwen3.5-4B ties its output layer to the embeddings, so the
option letters' logits are the final hidden state times head.safetensors (the 26 letter rows, fp32). Weights are
split into files of at most 32 MB.
The screenshot adapter also adapts the vision tower, so vision/ is Qwen3.5's ViT and patch merger from the same merge (668 MB, fp16 weights, fp32 compute). Its size-dependent inputs (position-table taps, 2D rotary angles) are computed by the caller, and the decoder takes its output as image_embeds at the <|image_pad|> tokens.
import * as ort from "onnxruntime-web/webgpu";
import { elementDecisions, loadCuaS1FourB } from "@ai-ecoverse/cua-s1.js/4b";
const model = await loadCuaS1FourB("https://huggingface.co/ai-ecoverse/cua-s1.js/resolve/main/cua-s1-4b-0.2-multimodal", { ort });
const r = await model.score(options, { app, taskFamily, goal, screenshot: imageData });
elementDecisions(r.options); // per element, its likeliest action
Against Cua's own cua_s1.four_b.FourBModel (fp32 PyTorch) on 60 real Word, Excel and PowerPoint steps from
GUI-360's test split: max |Δp| 0.1000, 0 argmax flips. Every
element's likeliest action is the expected one on 31/60 of them with this export, as cua-bench-s1 scores it.
Licensing: the adapter and Qwen3.5-4B are both Apache-2.0.