cua-s1.js weights

Browser-ready ONNX exports of Cua's cua-s1 form-filling decision models, for @ai-ecoverse/cua-s1.js and onnxruntime-web. The model, its training data and the planning rules are Cua's work (MIT); this repo only holds the converted checkpoint.

import * as ort from "onnxruntime-web/wasm";
import { loadCuaS1, extractEntities } from "@ai-ecoverse/cua-s1.js";

const model = await loadCuaS1("https://huggingface.co/ai-ecoverse/cua-s1.js/resolve/main/cua-s1-forms", { ort });
const plan = await model.plan("Northwind Clinic - New Patient Registration",
  [{ role: "Edit", label: "Phone number", value: "", token: "phone" }],
  extractEntities("Tel: (503) 555-0142\nWork phone: (503) 555-0110"));

Contents

Folder Source checkpoint ONNX Max |Δp| vs PyTorch
cua-s1-4b-0.2 cua-ai/cua-s1-4b-0.2 @ 1681886 4.7 GB 6.2e-02 over 60 tasks
cua-s1-4b-0.2-multimodal cua-ai/cua-s1-4b-0.2 @ 1681886 4.7 GB 1.0e-01 over 60 tasks
cua-s1-forms cua-ai/cua-s1-forms @ f54adbf 3.1 MB 3.2e-06 over 1048 decisions

Each folder has a manifest.json naming its source commit, the ONNX graphs' SHA-256s (checked by the loader) and the parity measured at export time. The graphs and Cua's original JSON sidecar (checkpoint.json: architecture, tensor signature, training metadata) live under r-<commit>/, so publishing a new checkpoint never changes a file an older manifest points at.

Conversion

The checkpoint is loaded with cua_s1.model.load_checkpoint, which validates the SHA-256 tensor signature. It is exported unmodified with the torch.export-based ONNX exporter (opset 18, dynamic batch, context, option and option-token axes). It takes the byte tensors cua_s1's ByteCollator produces and returns per-option logits and probabilities. onnxruntime matches PyTorch to within the max |Δp| in the table, on 1,048 decisions from Cua's own synthetic episode generator, with no argmax flips.

model-shared-options.onnx runs the same modules for a batch whose rows all share one option list, which is every form plan: the options are encoded once instead of once per element. Its inputs drop the batch axis from the option tensors. It gives the same probabilities as model.onnx and is checked against the unmodified PyTorch forward on the same episodes.

cua-s1 is a research checkpoint trained on synthetic forms. Read Cua's model card and SECURITY.md before relying on it.

cua-s1-4b

cua-s1-4b-0.2/ is cua-ai/cua-s1-4b-0.2's text adapter (a rank-16 LoRA) merged in fp32 into Qwen/Qwen3.5-4B and exported with the onnxruntime-genai model builder: int8 weights, fp32 activations, WebGPU. The graph returns hidden states; Qwen3.5-4B ties its output layer to the embeddings, so the option letters' logits are the final hidden state times head.safetensors (the 26 letter rows, fp32). Weights are split into files of at most 32 MB.

import * as ort from "onnxruntime-web/webgpu";
import { elementDecisions, loadCuaS1FourB } from "@ai-ecoverse/cua-s1.js/4b";

const model = await loadCuaS1FourB("https://huggingface.co/ai-ecoverse/cua-s1.js/resolve/main/cua-s1-4b-0.2", { ort });
const r = await model.score(options, { app, taskFamily, goal, axTree });
elementDecisions(r.options);   // per element, its likeliest action

Against Cua's own cua_s1.four_b.FourBModel (fp32 PyTorch) on 60 real Word, Excel and PowerPoint steps from GUI-360's test split: max |Δp| 0.0615, 3 argmax flips. Every element's likeliest action is the expected one on 51/60 of them with this export, as cua-bench-s1 scores it.

Licensing: the adapter and Qwen3.5-4B are both Apache-2.0.

cua-s1-4b

cua-s1-4b-0.2-multimodal/ is cua-ai/cua-s1-4b-0.2's screenshot (multimodal) adapter (a rank-16 LoRA) merged in fp32 into Qwen/Qwen3.5-4B and exported with the onnxruntime-genai model builder: int8 weights, fp32 activations, WebGPU. The graph returns hidden states; Qwen3.5-4B ties its output layer to the embeddings, so the option letters' logits are the final hidden state times head.safetensors (the 26 letter rows, fp32). Weights are split into files of at most 32 MB.

The screenshot adapter also adapts the vision tower, so vision/ is Qwen3.5's ViT and patch merger from the same merge (668 MB, fp16 weights, fp32 compute). Its size-dependent inputs (position-table taps, 2D rotary angles) are computed by the caller, and the decoder takes its output as image_embeds at the <|image_pad|> tokens.

import * as ort from "onnxruntime-web/webgpu";
import { elementDecisions, loadCuaS1FourB } from "@ai-ecoverse/cua-s1.js/4b";

const model = await loadCuaS1FourB("https://huggingface.co/ai-ecoverse/cua-s1.js/resolve/main/cua-s1-4b-0.2-multimodal", { ort });
const r = await model.score(options, { app, taskFamily, goal, screenshot: imageData });
elementDecisions(r.options);   // per element, its likeliest action

Against Cua's own cua_s1.four_b.FourBModel (fp32 PyTorch) on 60 real Word, Excel and PowerPoint steps from GUI-360's test split: max |Δp| 0.1000, 0 argmax flips. Every element's likeliest action is the expected one on 31/60 of them with this export, as cua-bench-s1 scores it.

Licensing: the adapter and Qwen3.5-4B are both Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ai-ecoverse/cua-s1.js

Finetuned
Qwen/Qwen3.5-4B
Adapter
(694)
this model