Skip to content

Supported models and hardware

This page covers what runs from main. Start with the CPU decision walkthrough. The rolling model tracker (#83) records available paths separately from proposed integrations and hardware evidence. Models that are being added are also tracked in issues labeled new model.

Model Worker CPU NVIDIA CUDA Apple GPU Requirements
LAYA, English checkpoint External worker running the upstream Laya runtime, for text requests Validated (#2); CPU demo Unverified (#39) Use the dedicated MPS worker below Python 3.12, pinned CPU dependencies
LAYA, English checkpoint Python MPS/CPU worker in this repository Contract checks documented (#67) No validation recorded for this path PyTorch MPS validated on M1 Pro; M4/M5 checks documented (#30, #67) Python 3.12, MPS dependency versions; native Metal execution remains planned
LAYA, English checkpoint Native Rust worker Processing/checkpoint checks; no CPU inference Hopper sm_90a; validation scope Not supported CUDA toolkit, TileLang for AOT generation, pinned checkpoint and rotary tables
Cua-S1 4B 0.2, text adapter Reference worker on Transformers and PEFT Unverified Validated (#13) Unverified Python 3.12, the versions in requirements-text.txt
Cua-S1 4B 0.2, text adapter Native Rust worker on the Qwen3.5 CUDA kernels Not supported Validated on compute capability 8.9 (#19, #52) Not supported Compute capability 8.0 or newer, the CUDA toolkit to build, weights merged with export_text_merged.py
Cua-S1 4B 0.2, multimodal adapter Reference worker on Transformers and PEFT, src/frontend/cua_s1.py; no recipe yet Not supported Validated (#17, #18) Not supported The state is one PNG or JPEG image; upstream's weights.lock.json next to the base weights
Open-Jev-27B-v1.1 Native Rust/CUDA worker on the shared Qwen3.5/3.8 executor Not supported Validated on H200 (sm_90) for the 74 single-candidate workload Not supported Compute capability 8.0 or newer, CUDA toolkit to build, exported merged weights and trained head
CLM-v0.1-8B External clm-serve recipe with a CPU stub embeddings server Stub-encoder contract checks only (#23); not real Qwen3-8B decisions Real encoder unverified by the merged recipe Unverified Python, upstream CLM and head checkpoint; a real encoder requires a separate embeddings server
  • Validated: covered by the recipe on main or by the checks in the linked merged pull request.
  • Unverified: the worker accepts this device, but no recipe or merged pull request covers it.
  • Planned: not implemented yet; the linked issue tracks it.

The Cua-S1 workers answer choice questions only. LAYA's English worker and Open-Jev support choice, score, and noul text questions. CLM's merged recipe exercises these answer shapes with stub embeddings; it does not validate decision quality. MPS validation above is for a Python/PyTorch worker, not a native Metal backend.

The architecture contracts describe the native target. Shared processing orchestration and dynamic batching remain planned. Native workers reuse serial admission. Qwen workers run independent single-prompt prefills with CPU heads; Laya packs questions within one request and runs its scorer/action head on CUDA. These target layers do not expand the validated model or hardware coverage above.