Supported models and hardware¶
This page covers what runs from main. Start with the
CPU decision walkthrough. The
rolling model tracker (#83)
records available paths separately from proposed integrations and hardware evidence.
Models that are being added are also tracked in issues labeled new model.
| Model | Worker | CPU | NVIDIA CUDA | Apple GPU | Requirements |
|---|---|---|---|---|---|
| LAYA, English checkpoint | External worker running the upstream Laya runtime, for text requests | Validated (#2); CPU demo | Unverified (#39) | Use the dedicated MPS worker below | Python 3.12, pinned CPU dependencies |
| LAYA, English checkpoint | Python MPS/CPU worker in this repository | Contract checks documented (#67) | No validation recorded for this path | PyTorch MPS validated on M1 Pro; M4/M5 checks documented (#30, #67) | Python 3.12, MPS dependency versions; native Metal execution remains planned |
| LAYA, English checkpoint | Native Rust worker | Processing/checkpoint checks; no CPU inference | Hopper sm_90a; validation scope |
Not supported | CUDA toolkit, TileLang for AOT generation, pinned checkpoint and rotary tables |
Cua-S1 4B 0.2, text adapter |
Reference worker on Transformers and PEFT | Unverified | Validated (#13) | Unverified | Python 3.12, the versions in requirements-text.txt |
Cua-S1 4B 0.2, text adapter |
Native Rust worker on the Qwen3.5 CUDA kernels | Not supported | Validated on compute capability 8.9 (#19, #52) | Not supported | Compute capability 8.0 or newer, the CUDA toolkit to build, weights merged with export_text_merged.py |
Cua-S1 4B 0.2, multimodal adapter |
Reference worker on Transformers and PEFT, src/frontend/cua_s1.py; no recipe yet |
Not supported | Validated (#17, #18) | Not supported | The state is one PNG or JPEG image; upstream's weights.lock.json next to the base weights |
| Open-Jev-27B-v1.1 | Native Rust/CUDA worker on the shared Qwen3.5/3.8 executor | Not supported | Validated on H200 (sm_90) for the 74 single-candidate workload | Not supported | Compute capability 8.0 or newer, CUDA toolkit to build, exported merged weights and trained head |
| CLM-v0.1-8B | External clm-serve recipe with a CPU stub embeddings server |
Stub-encoder contract checks only (#23); not real Qwen3-8B decisions | Real encoder unverified by the merged recipe | Unverified | Python, upstream CLM and head checkpoint; a real encoder requires a separate embeddings server |
- Validated: covered by the recipe on
mainor by the checks in the linked merged pull request. - Unverified: the worker accepts this device, but no recipe or merged pull request covers it.
- Planned: not implemented yet; the linked issue tracks it.
The Cua-S1 workers answer choice questions only.
LAYA's English worker and Open-Jev support choice, score, and noul text questions.
CLM's merged recipe exercises these answer shapes with stub embeddings; it does
not validate decision quality. MPS validation above is for a Python/PyTorch
worker, not a native Metal backend.
The architecture contracts describe the native target. Shared processing orchestration and dynamic batching remain planned. Native workers reuse serial admission. Qwen workers run independent single-prompt prefills with CPU heads; Laya packs questions within one request and runs its scorer/action head on CUDA. These target layers do not expand the validated model or hardware coverage above.