Skip to content

Laya CUDA validation

The combination contains the selected RoPE tile, short QKV/output dispatch, down N64, short GEGLU selection and long GEGLU BN32/static registers. It preserves the original arithmetic and checkpoint precision. Failed/HOLD candidates are excluded; this does not claim zero GPU bubbles or maximum occupancy.

Frozen native engine comparison

These measurements precede the current-main integration. They used one frozen Rust CLI with the original and optimized CUDA configurations, on one shared H800, concurrency 1, FP16 token embeddings, BF16 projection matrices/activations and FP32 residuals/norms. Clocks were not locked. JSON parsing, tokenization, padding/upload, forward/heads, readback, decoding and response JSON write are included. HTTP, process/model startup and warmup/cold Graph construction are excluded.

Graph input Original p50 Combined p50 Latency reduction Speedup
choice, L48 2.5426 ms 1.5957 ms 37.24% 1.59×
score, L64 2.5819 ms 1.6101 ms 37.64% 1.60×
short, L48 2.5419 ms 1.5959 ms 37.22% 1.59×
medium, L176 2.8666 ms 1.9207 ms 33.00% 1.49×
long, L512 3.9926 ms 3.0001 ms 24.86% 1.33×

Each entry pools two 100-sample passes. The 12 processes measured 6000 requests across original/retained/combined variants, Graph/eager, in forward/reverse order. All raw engine/client timings and original file hashes are in measurements.json; private host paths are omitted. The original uses --original-rope and the reverse-patched exporter; both sides share the same runtime and Graph setting. There is no single workload-weighted percentage. The full pipeline was measured directly, not by adding kernel gains.

The frozen CLI binary SHA256 is 4d31e8593ff67127c5c908c3bdba26db393165036027e58a5554b70567a71c85; the combined library is eeb94478354b5b6f0b26c756263261dd61908a0aafd82f10b04b907c3d46cc36. Source kernels/runtime and exported library identity are retained in the evidence. The earlier three-way numeric check covered 210 responses, 36 hidden-state comparisons and six 17-request reuse sequences, bitwise equal to the original native configuration. This tests implementation equivalence for fixed inputs, not general model quality or all-input agreement with official PyTorch.

Integration checks

The current-main integration retains checkpoint inventory checks and imports the reviewed CPU parsing/decoding fixes. Local checks passed: workspace fmt/clippy, 83 Rust tests (10 fixture/checkpoint/GPU tests ignored), release build, eight CUDA build-entry tests, eight benchmark tests and strict docs build.

Linux laya-run, omni-laya and omni-jev were built from commit 10c14a34c870d9d4a6562e1dcf7ca787a716730f. Subsequent changes only update docs and normalize blank context lines in the baseline reproduction patch; execution source and dependencies are byte-identical. The build used Rust 1.98.1 and the same pinned CUDA library, checkpoint and fixtures as the engine comparison.

The new binaries passed actual H800 acceptance with those 12 fixtures:

  • Graph and eager CLI: 24 complete JSON responses exactly match the fixed official reference; all 24 raw-head results match the original native engine with exact FP32 bits and request-derived row widths.
  • Native worker and frontend proxy: 24 successful HTTP responses exactly match the same official JSON. Four malformed-JSON/content-type requests return 422/415, and four health checks report ready before and after rejection.
  • All four child processes exit with code 0; the worker and frontend shut down through SIGINT. Native process maps contain the registered CUDA library and no Python/Torch runtime. The run's GPU processes and personal lock are released.

JSON and raw-head comparisons use different references. The official reference comes from Laya 0.3.20's fast Graph path with selected RoPE and BF16 autocast. Four action values in long_3/truncated_3 already differed from that reference in the original native engine; they remain unchanged here. Therefore this check establishes native implementation equivalence and exact response JSON for the fixed inputs, not bitwise raw-head agreement with PyTorch or general model quality. The first runner incorrectly mixed these two references; its failure was retained and corrected by binding the original native raw outputs, without adding tolerance.

Complete responses, per-call output digests, test counts, binary/fixture/library identities and the original native raw reference are in integration_validation in measurements.json. These are functional checks; current HTTP latency and inference-failure injection were not measured. The timing table above remains historical engine evidence.

Build and reproduction commands are in the recipe. Full model weight/tokenizer oracle checks remain explicit opt-in tests, with their pinned artifact hashes; see the model contract.