Skip to content

CLM behind the frontend

Runs CLM's own clm-serve behind omni-jev, and checks that the frontend returns what the engine returned. CLM is the second model in #9: a frozen Qwen3-8B encoder behind an OpenAI-compatible /v1/embeddings endpoint, plus two projection heads and a cosine score.

No GPU and no 8B encoder are needed to run this. clm-serve is an HTTP client of the embeddings endpoint (src/clm/embedder.py), so stub_embedder.py can stand in for the encoder. What that exercises is the plumbing — request shape, question packing, the three answer types, the serving contract, and the frontend in front of it. It cannot tell you anything about CLM's decisions, because the vectors are not from Qwen.

Run it

Three terminals, all CPU:

# 1. a stand-in for the Qwen3-8B pooling server
python recipe/clm/stub_embedder.py --port 8090

# 2. CLM's own server, pointed at it (the checkpoint is 75 MB: the two heads, not the encoder)
# CLM_CKPT is the file itself; without it clm-serve looks in ~/.cache/clm and downloads.
CLM_CKPT=/tmp/CLM_v0.1-8B.pt clm-serve --port 8091 \
  --emb-url http://127.0.0.1:8090/v1/embeddings --emb-model qwen3-8b

# 3. the frontend from #2, pointed at CLM
OMNI_JEV_BIND=127.0.0.1:8080 OMNI_JEV_BACKEND_URL=http://127.0.0.1:8091 cargo run -p omni-jev --release
python recipe/compare_with_backend.py \
  --backend http://127.0.0.1:8091 --frontend http://127.0.0.1:8080 --model clm-latest

Installing CLM without the GPU stack, since vllm is a hard dependency of the package but is only needed for the encoder process:

pip install "numpy>=1.24" requests "fastapi>=0.100" "uvicorn>=0.23" torch
pip install --no-deps "contrastive-lm @ git+https://github.com/Contrastive-LM/CLM.git"

What passes

The contract lines up with no adapter: status, content type and all three answer types come back through the frontend unchanged, including CLM's X-CLM-Latency-Ms.

question answer
choice {"type":"choice","choice":"billing","confidence":…,"probabilities":{…}}
score {"type":"score","score":0.750,"confidence":…,"legend":{"0":…},"probabilities":{…}}
noul {"type":"noul","noul":0.9995}

What the comparison had to learn

compare_with_backend.py compared the whole response byte-for-byte, which holds for LAYA because its body is a pure function of the request. It does not hold for an engine that reuses encoder state across requests:

same state, three times:   noul=0.977197  usage.input_tokens=0
a state not seen before:   noul=0.280477  usage.input_tokens=26
that same state again:     noul=0.280477  usage.input_tokens=0

The decision is deterministic to six decimals; input_tokens counts only encoder cache misses. CLM's whole point is that candidate vectors are reusable across requests, so the field is moved by the feature that makes it interesting. Comparing the full body reports FAIL on a correct response, and whether it does depends on which call happened to warm the cache — so the same run can pass or fail on ordering.

The tool now compares status, content type and the answers subtree, and prints the usage difference instead of asserting on it. Strict equality is still the first test, so a backend whose body really is a pure function is unaffected:

PASS department: status 200 -> 200  (usage {'billing_units': 1, 'input_tokens': 33, 'output_tokens': 0} -> {'billing_units': 1, 'input_tokens': 0, 'output_tokens': 0})

Open contract question

Which response fields are allowed to differ between two otherwise identical requests? billing_units looks stable; input_tokens does not. If the project wants the strong form — the whole body identical — then input_tokens has to mean "tokens the request required" rather than "tokens this call paid for", which is a decision for the engine, not for the frontend.

With a real encoder

Replace step 1 with the upstream script and the answers become meaningful:

GPU=0 PORT=8090 UTIL=0.35 ./serve_qwen3_8b.sh     # from the CLM checkout; needs CUDA

Everything downstream is unchanged, which is the property this recipe is meant to demonstrate.