Open-Jev-27B-v1.1 native text worker¶
The worker owns request compilation, tokenization, candidate scoring and typed
responses in Rust. It uses the native CUDA prefill implementation introduced in
PR #19, shared with Cua-S1
under src/models/qwen3_5/native/.
Python is required only to prepare the merged checkpoint.
It supports choice (1–255 candidates), score (2–10 levels), and noul
(yes/no). Each candidate has an independent prompt; the last hidden state goes
through Open-Jev's trained FP32 scalar head. Noul uses logits [0, score].
The saved calibration temperature is applied before normalizing each complete
question. There is no autoregressive generation. Structured state and descriptions
use Open-Jev's sorted JSON rendering; question and candidate order is preserved.
Prepare the checkpoint¶
Use the reference dependencies from
Open-Jev @ 3308a15:
PyTorch 2.8 or newer, Transformers 5.10.2, PEFT 0.19.1, Accelerate 1.13.0,
and safetensors. An optional kernels installation must be compatible with that
Transformers release. Run these commands from the repository root:
hf download Qwen/Qwen3.8-27B \
--revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 \
--local-dir weights/Qwen3.8-27B
hf download ZefanCai/Open-Jev-27B-v1.1 \
--revision 28cf73067d5b337860bbef3c85b8b82ba8730956 \
--include 'package/checkpoint/*' --local-dir weights/Open-Jev-27B-v1.1
CUDA_VISIBLE_DEVICES='' python recipe/open_jev/export_merged.py \
--base weights/Qwen3.8-27B \
--checkpoint weights/Open-Jev-27B-v1.1/package/checkpoint \
--out weights/open-jev-27b-merged
CPU export needs roughly 110 GB of RAM and 52 GB of output storage. It merges
LoRA in BF16 and saves the trained head, temperature, and single-user chat
template in open_jev_export.json. The worker refuses a plain base checkpoint
or an incomplete export. The saved limit defaults to 4096 tokens per candidate;
--max-length may raise it to 16384. Oversize prompts fail before inference.
Build and serve¶
The CUDA kernels require compute capability 8.0 or newer. The current build
target below is Ada (89); pass your GPU's compute capability explicitly.
The CUDA shared library and both Rust workers must be rebuilt together because
the gated-attention entry point updates the library ABI to version 4 alongside
the shared CUDA Graph entry points.
src/backends/cuda/qwen3_5/build.sh target/release 89
cargo build --release --locked -p omni-open-jev-native -p omni-jev
OPEN_JEV_MODEL=weights/open-jev-27b-merged \
target/release/omni-open-jev-native
OPEN_JEV_HOST and OPEN_JEV_PORT default to 127.0.0.1 and 8000.
OPEN_JEV_CUDA_LIB overrides the default library next to the executable.
The worker loads all text weights onto visible CUDA device 0, performs a real
warmup inference, then exposes /health and /v1/systemone.
Use a reservation before any GPU command on hosts with a GPU scheduler.
In another terminal, start the existing Rust frontend:
OMNI_JEV_BIND=127.0.0.1:8080 OMNI_JEV_BACKEND_URL=http://127.0.0.1:8000 \
target/release/omni-jev
curl http://127.0.0.1:8080/v1/systemone \
-H 'Content-Type: application/json' --data-binary @recipe/open_jev/example-request.json
The worker accepts the model's base name Qwen/Qwen3.8-27B, open-jev,
jev-latest, and open-jev-27b-v1.1; the response model is the base name,
matching Open-Jev. Error wording and metadata differ from the reference service.
Requests are bounded to 4 MiB, 4096 questions, and 65536 candidate sequences.
Validation and optimization scope¶
Tests and fixtures live in the repository-level tests/ tree: Open-Jev's typed
contract and tokenizer cases are in
tests/open_jev/, and shared Qwen JSON, configuration
and CUDA reference tests are in tests/qwen3_5/.
The default suites below run on CPU without downloading model weights:
cargo test --locked -p omni-open-jev-native -p omni-qwen3-5-native
cargo test --locked -p omni-jev --test frontend
The frontend mock-worker API coverage is tracked in issue #46 and PR #58. Checkpoint tokenizer and CUDA kernel tests are opt-in; the latter require a GPU reservation:
# Inside a GPU reservation, after building the library:
CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \
cargo test --release --locked -p omni-qwen3-5-native --test kernels -- --ignored
CPU golden fixtures come from Open-Jev's request compiler and response formatter
at the revision above. Kernel tests compare attention and Gated DeltaNet with
float64 references and require exact BF16 equality between fused attention gating
and a separate gate pass. Residual RMSNorm and packed SiLU are checked against
rounded references, including odd widths and unaligned pointers. The shared
kernel tests retain PR #19's
CUA_S1_CUDA_LIB environment variable.
The worker reuses PR #19's fused norm, activation, QK/RoPE and chunked Gated DeltaNet operations. Attention's sigmoid gate is fused into its output epilogue, preserving both BF16 rounding points and removing one launch and one output read/write pass per full-attention layer (16 layers for this model). Residual RMSNorm keeps thread values in registers at widths 2560/5120. MLP SiLU uses 16-byte BF16 loads/stores when width, stride and pointers permit it, retaining both BF16 rounding points; other layouts use the scalar path.
This recipe leaves CUA_S1_GRAPH unset and runs one eager forward pass per
candidate. Set CUA_S1_GRAPH=1 on the worker to enable CUDA Graph replay. The
shared backend retains at most 64 graphs, keyed by exact candidate token length;
growing the scratch buffer clears them. Capturing a new length first runs an
eager forward to initialize its plans, then captures and replays the forward.
This adds cost for new lengths, so graph mode remains opt-in. Warm replay is
validated on the 74-case H200 workload: its mean HTTP latency is 2.03% below
eager execution after all workload lengths are warmed. Tokenization, transfers
and the CPU scalar head remain outside the graph. Prefix sharing, GEMM autotuning,
quantization and multimodal inference are not implemented. The
H200 validation reports full-checkpoint results for 74
single-candidate requests, including probability differences and timing
variability. It does not establish general accuracy parity or a speedup over
OpenJev-Fast; the author's B300 results use different hardware and workloads.