Open-Jev H200 validation¶
Raw HF Transformers comparison, 2026-10-03¶
The native Rust/CUDA worker delivers a 7.47× speedup over raw HF Transformers
by mean warm HTTP latency: 362.21→48.50 ms (86.61% lower) on one H200.
This comparison covers 74 real JevBench noul requests, one candidate each,
80–3399 tokens, BF16, max length 16384 and concurrency 1. Each backend reuses
one server for an excluded feasibility pass and two measured passes:
148 measured requests per backend.
| Configuration | Overall mean (ms) | Mean, pass 1 / pass 2 (ms) | P50, pass 1 / pass 2 (ms) | P95, pass 1 / pass 2 (ms) | Correct / 74 |
|---|---|---|---|---|---|
| Raw HF Transformers (unmerged LoRA; PyTorch fallback) | 362.209 | 362.238 / 362.180 | 324.822 / 321.321 | 651.999 / 652.911 | 64 |
| Native Rust/CUDA (cached RMSNorm + packed SiLU) | 48.503 | 48.471 / 48.535 | 25.797 / 24.895 | 215.696 / 218.882 | 64 |
| Original OpenJev-Fast | 50.936 | 51.097 / 50.775 | 24.651 / 24.622 | 248.121 / 251.193 | 63 |
Native and HF agree on all 74 thresholded decisions and both score 64/74;
their maximum probability difference is 0.020423. Fast scores 63/74 with
one different decision (hard-opus-a-temporal_numeric-09); maximum native/Fast
probability difference is 0.034353. Each backend's probabilities are exactly
unchanged across its feasibility and measured passes. These counts do not
establish statistical accuracy superiority or full numerical parity.
Raw HF baseline: the original Open-Jev server and DecisionModel, with
160 unmerged PEFT LoRA modules, the trained scalar head and saved temperature.
Full attention uses stock SDPA; runtime assertions verify all 48 linear-attention
layers use torch_chunk_gated_delta_rule, stock convolution and
Qwen3_5RMSNormGated. Because the shared environment contains optional kernels,
the baseline makes Transformers' is_flash_linear_attention_available and
is_causal_conv1d_available checks return false before importing model classes.
It uses no custom Fast model, torch.compile, CUDA Graph replay or prefix cache.
Native uses merged LoRA and eager execution (CUA_S1_GRAPH=0); original Fast
retains its custom kernels and CUDA Graph stack, reporting 30 retained graphs.
This comparison changes the complete backend; it does not isolate one optimization.
Timing includes localhost HTTP through the same frozen Rust frontend,
tokenization, worker execution and UTF8 response decoding. Client body
serialization and response JSON parsing are excluded. Downloads, preparation,
process-to-readiness, warmup, first inference after readiness and the complete
feasibility pass are excluded. No Nsight launcher or trace collection was used.
P50 is the median; P95 uses JevBench's sorted[int(0.95*N)-1] rank.
The observed mean pass ranges are disjoint; two-pass ranges are not confidence
intervals. Native's mean is 4.78% below Fast in this run, while Fast has a lower
median. Multi-candidate prefix sharing and broader JevBench coverage remain
unmeasured; native CUDA Graph replay is reported separately below. The author's
17.3 ms B300 result uses different hardware and workload.
Frozen controls and reproduction¶
- Exact H200 GPU 2, UUID
GPU-cbf66259-f4ab-0ede-1811-82037dde5924, NUMA 0, CPUs 0–15. The archived driver name isNVIDIA L20X, with 143771 MiB and SM90 / 132 SMs. The same device, affinity, requests, model and prepared environment are used for all three backends; no shared caches are dropped. - Frozen native worker/frontend:
202c0e163f868334a99d88407056ebe61dbb2dce; accepted packed-SiLU CUDA library SHA256:e033315d4c67127809e62f41d991a149ac8ee0ce81827997af063815fd00d435. PR source at measurement:ad1cb818d2ea001bbd8c84643e1a7241d0fc26c8. - Original Open-Jev
at
3308a15ccd7eea1df7a37d6ddc39b023b801ba16. Fast, JevBench, request JSON, base, adapter and temperature use the same pins recorded in the October 2 controls below and the native recipe. - Actual runtime: Torch 2.13.0+cu130, CUDA 13.0, Transformers 5.10.2, PEFT 0.19.1. Reuse prepared weights and SM90 extensions offline.
- Use
gpu run --gpu-ids 2 --wait 10m --timeout 45m --note <label> --, thennumactl --membind=0 --physcpubind=0-15. Run raw HF, native and Fast, preserving the 74-case request order, with one excluded feasibility pass and exactly two measured passes per configuration.
Raw commands, request/token IDs, timing rows, source snapshots and 48 verified
input hashes are archived locally in the benchmark worktree's
profile/jev-hf-transformers-comparison-20261003/, outside this PR.
A collector cleanup assertion rejected native's intentional SIGTERM exit after
all HF/native measurements were saved. Only the remaining Fast configuration
continued on a second reservation of the same GPU and affinity; no measured
passes were repeated or added. All task-owned processes exited and GPU 2 returned
to 0 MB used. The October 2 timings below are a separate experiment.
Native CUDA Graph replay, 2026-10-03¶
Warm graph replay reduces mean HTTP latency from 48.086 to 47.112 ms (2.03%)
on the same 74-case workload. The shared backend now retains up to 64 exact-length
graphs, enough for these 57 distinct lengths. Its previous eight-entry cache
regressed the mixed workload because evicted lengths require another eager
forward and graph capture. Graph mode remains opt-in with CUA_S1_GRAPH=1.
| Workload | Configuration | Mean (ms) | Mean, pass 1 / pass 2 (ms) |
|---|---|---|---|
| Fixed 107-token request | Eager | 19.829 | 19.835 / 19.823 |
| Fixed 107-token request | Graph, eight entries | 18.882 | 18.902 / 18.863 |
| Fixed 107-token request | Graph, 64 entries | 19.081 | 18.968 / 19.194 |
| 74 mixed-length requests | Eager | 48.086 | 48.062 / 48.110 |
| 74 mixed-length requests | Graph, eight entries | 95.289 | 95.425 / 95.153 |
| 74 mixed-length requests | Graph, 64 entries | 47.112 | 47.050 / 47.174 |
All 74 probabilities and decisions remain exactly unchanged (maximum delta 0.0). The 64-entry candidate also reduces the fixed-short mean by 3.77%. Both measured passes improve over eager for both workloads, satisfying the prespecified 3% short and 2% mixed mean gates. The mixed improvement only narrowly exceeds its gate; two passes are observed variability, not confidence intervals or evidence for a general workload winner. Observed device memory after mixed passes is 50,947 / 50,967 / 51,089 MB for eager / graph-eight / graph-64: 122 MB more for 64 entries than eight. These are scheduler samples, not peak-memory measurements.
This is a separate native A/B experiment, not a newly measured HF/Fast comparison.
It keeps H200 GPU 2, UUID, NUMA affinity, BF16, max length 16384, model export,
CUDA library, frontend and request order fixed. Both graph capacities are built
with identical rustc options and the frozen 202c0e1 dependency artifacts;
their source copies differ only in the cache limit. Eager uses the eight-entry
binary with graph mode disabled. Each configuration validates its first long
inference after readiness to allocate the workload's maximum scratch size, then
validates the short request. One server is reused for an excluded 32-request
short feasibility pass and two measured 32-request passes, followed by an excluded
74-case feasibility pass and two measured 74-case passes. HTTP timing has no
Nsight launcher and uses the same boundary as the raw HF comparison above.
New-length capture cost remains significant: excluded mixed feasibility means are 50.130 / 97.694 / 110.255 ms for eager / graph-eight / graph-64. These are single feasibility observations, not measured cold-latency comparisons. Scratch growth invalidates graphs, and more than 64 distinct lengths can still evict them. No model downloads, cache drops or clock changes occur. All owned processes exit and GPU 2 returns to 0 MB used.
A preceding experiment proves warm replay in Nsight Systems: one graph launch, zero recaptures and zero individual runtime kernel-launch calls per short request, with the same 834 kernels. Node-level graph tracing reports larger gaps despite lower unprofiled HTTP latency; use unprofiled measurements for the speedup. NVIDIA documents graph-node tracing overhead. Hardware counters and per-SM utilization remain unmeasured.
Raw plans, source copies, build commands, requests/responses, traces, memory
samples and verified input hashes are archived outside this PR in
profile/jev-cuda-graph-20261003-074613/ and
profile/jev-cuda-graph-cache64-20261003-075238/. Run the native worker with
CUA_S1_GRAPH=0/1 under the same reservation and affinity documented above;
warm the maximum workload length and all tested lengths before measured passes.
For new comparisons, declare the budget first and retain capture costs whenever
they occur inside measured requests.
Packed-SiLU A/B, 2026-10-02¶
A matched comparison of 74 real JevBench noul requests, each with one candidate,
measured packed MLP SiLU with cached residual RMSNorm fixed in both native variants.
All 74 native decisions and probabilities were exactly unchanged by packed SiLU;
the native worker scored 64/74 correct. OpenJev-Fast scored 63/74, with one different
decision (hard-opus-a-temporal_numeric-09) and maximum native/Fast probability
difference 0.03435. These counts do not establish statistical accuracy superiority.
Warm HTTP latency¶
| Configuration | Mean, pass 1 / pass 2 (ms) | P50, pass 1 / pass 2 (ms) |
|---|---|---|
| Native with cached residual RMSNorm; scalar SiLU | 49.563 /49.490 | 25.434 /25.539 |
| Native with cached residual RMSNorm and packed SiLU | 48.286 /48.307 | 24.925 /24.657 |
| Original OpenJev-Fast | 57.345 /50.908 | 25.026 /24.770 |
Packed SiLU reduces native mean 49.527→48.297 ms (2.483%). Its two-pass mean range is disjoint from baseline's. These are observed two-pass ranges, not confidence intervals. Fast's 6.436 ms mean spread prevents claiming a stable aggregate winner. This subset does not measure multi-candidate prefix sharing. The author's 17.3 ms B300 result is a different hardware/workload measurement.
HTTP includes the same Rust frontend, tokenization and worker execution; excludes client body serialization and response JSON parsing, includes UTF8 decoding. One server per configuration is reused. Real warmup, first request after readiness and one complete 74-case feasibility pass are excluded; two subsequent passes are measured. Nsight collection is inactive during HTTP passes, although CUPTI instrumentation may remain loaded. No shared caches are dropped or clocks changed.
Separate CUDA timelines¶
Each entry is the two-trace mean of 64 MLP SiLU launches per request, in milliseconds.
| Tokens | Scalar SiLU | Packed SiLU | Fast SiLU lookup |
|---|---|---|---|
| 107 | 0.680 | 0.305 | 0.359 |
| 936 | 5.643 | 1.746 | 1.513 |
| 3399 | 21.217 | 6.380 | 5.466 |
Long-request MLP SiLU decreases 69.93%, closing 94.19% of that measured kernel family's gap to Fast. Other kernel families vary between traces; their separate timeline totals are not HTTP latency. Fast uses padded/tree layouts and lookup tables, so this does not imply identical rows, fusion boundaries or arithmetic.
The prespecified acceptance gates were five GPU tests, unchanged native decisions, maximum probability delta ≤0.01, ≥25% long SiLU duration reduction and ≥2% warm HTTP mean reduction with disjoint observed ranges. All passed; no extra measured runs were added. Final integration also passed the six-test shared CUDA ABI 4 suite.
Frozen controls and reproduction¶
- Device: exact scheduler GPU 2, UUID
GPU-cbf66259-f4ab-0ede-1811-82037dde5924, NVIDIA H200, 143771 MiB (reported asNVIDIA L20Xin the archived device metadata); CUDA driver and Nsight identify SM90 / 132 SMs. NUMA 0, CPUs 0–15. - BF16, max length 16384, HTTP concurrency 1, native
CUA_S1_GRAPH=0. Fast retains its original graph/kernel stack. Build with nvcc 13.0.88, SM90,-O3 -std=c++17 -lineinfo; use the same cuBLASLt/runtime libraries for both native variants. - Frozen native worker/frontend:
202c0e163f868334a99d88407056ebe61dbb2dce. Both native libraries include the cached RMSNorm source; onlyelementwise.cudiffers for packed SiLU. The shared CUDA ABI remains 4. - JevBench
at
f8ce71361165846101d02ebc83ad44e47ae44fc3; select its 74noulcases, preserving request bodies and order. Frozen request JSON SHA256:0c756a7b0b4c1f1352225f2e01770b5b3a0646fe86d5cc6930ded9e292ef37df. - OpenJev-Fast
at
c52b8bb958c1f0d241d4eb7fce4ecd8d885bf1e4; original server/model/kernels, with prepared SM90 extensions, PyTorch 2.13cu130/Triton 3.7.1, Transformers 5.10.2, PEFT 0.19.1 and FLA 0.5.2. These differ from the author's B300 environment. - Model/export uses the pinned base, adapter and calibration in the native recipe; temperature 2.5343690298472983. Reuse prepared weights and extensions. Keep copies, downloads, compilation, warmup and process-to- readiness separate from measured execution.
Reserve the exact device with gpu run --gpu-ids 2 --timeout 45m --note <label> --,
then bind the command with numactl --membind=0 --physcpubind=0-15. Run baseline,
packed native and Fast once each, reusing each server for one excluded feasibility
and two measured passes. Separately trace two requests at each 107/936/3399 tokens
per configuration with Nsight Systems CUDA/node tracing: 18 traces total.
Raw bodies, timing rows, SHA256 manifests, reproduction commands,
18 .nsys-rep/SQLite pairs and plots are archived locally in the benchmark
worktree's profile/jev-single-candidate-silu-pack8-20261002/, outside this PR.
Nsight Compute counters are denied by the host policy. Geometry,
register/shared-memory metadata and CUDA timelines are available; achieved
occupancy, per-SM tails, stalls, Tensor Core utilization and bandwidth/cache
efficiency are unmeasured. No hardware-cause claim follows from timeline data alone.