Flash serving cookbook
Based on the October 5, 2026 isolated campaign. Production wasn't changed. Commands and values below describe measurement and reproduction boundaries, not a qualified deployment.
1. Freeze the reproduction target
The reproduction download contains reproduction.json, the reviewed model-file hashes, measured-exl-settings.json and measured-exl-environment.json. The settings file is an exact subset of the tested config, using its real key names. Supply your own paths, network/authentication and the remaining pinned defaults; it isn't a complete launch config.
| Component | Tested identity |
|---|---|
| TabbyAPI | 53da7919d4e45c63f4acbcbbc00cbe0f60a1ce65 |
| Dual-5090 recipe and overlays | af2c4b841678a996632bf8e6bbd719aea66b823d |
| EXL tree, before recipe patches | 5783a9360b749a1d5bde6862dc35e7766e542ca1 (v1.5.2 source) |
| EXL base wheel | 1.5.0+cu128.torch2.9.0, Python 3.12 Linux x86-64 |
| EXL wheel SHA-256 | d8f5b483ca882c52d79a73c059b508c666663a68b4c60e420ca489e654e8104a |
| Model repository | r0b0tlab/Qwen3.8-Flash-Next-EXL3-2.50bpw |
| Model revision | 61a1a139ef2411ed6ed5717acfaff288511a1b08 |
| Strata comparison pin | 6f32ec070f23ced9f50e704d854d775da52591ab |
The model revision is not an EXL source commit. The runtime installed the wheel, prepared the recipe's rebase-dev-r3 tree from the pinned EXL dev source plus ported-vs-dev.patch, replaced the installed EXL tree, rebuilt its extension and installed the recipe overlays. A bare wheel install doesn't reproduce it. The retained setup script and log establish that sequence; this download doesn't include weights or the built runtime.
Read the pinned dual-5090 recipe, pinned TabbyAPI and pinned model artifact. These public identities were checked on October 6, 2026; historical performance belongs to the October 5 campaign. Don't substitute current upstream defaults. Record CUDA, driver, Torch, GPU topology, CPU, RAM and power limit before a new run.
Measured tuned EXL settings, expressed as an inventory rather than an unverified YAML schema:
| Setting | Tested value |
|---|---|
| Model artifact | EXL3 2.50bpw Flash Next |
| KV precision | Quant8 K / Quant8 V |
| GPU cache pool | 819,200 tokens |
| Maximum sequence | 262,144 tokens |
| Maximum batch | 5 |
| MTP draft depth | 4 |
| GPU placement | 28.5 / 31.5 GiB split values |
| Host KV tier | 32,768 MiB |
| Recurrent host tier | 8,192 MiB |
Do not copy inventory names into a server config without checking the pinned schema. Host pools and GPU split values are hardware-dependent. Fully warmed observed free VRAM was about 1,994 / 1,030MiB in the stated short/mixed cells; this isn't a guaranteed worst-case reserve.
2. Check the server in isolation
Before each arm, verify the exact preceding engine process and its descendants have stopped. Inspect the GPU process inventory. Refuse the next arm if an old engine remains.
nvidia-smi --query-gpu=index,memory.free,utilization.gpu,power.draw --format=csv
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csvUse the engine's actual ready endpoint. A model-list response can precede usable inference. Run a bounded smoke request before timing. Keep builds, downloads and correctness tests outside measured cells.
3. Fit the actual token budget
Use the server tokenizer and rendered template. Count system text, tool schema, messages and reasoning history. Reserve native end positions and generation tokens before admitting the request.
actual rendered prompt + requested total generation + native reserve <= context limit
Measured full-window arithmetic:
- EXL: 253952 + 8192 = 262144.
- Strata serial: 253943 + 8192 + 8 = 262143, one unused position.
- Stock/Bee: 253950 + 8192 + 2 = 262144.
Keep automatic truncation disabled for full-budget qualification. Rejection is a result. Record actual output and finish reason; an early EOS doesn't count as reaching the requested cap.
4. Preserve returning history
Replay exact assistant content, reasoning fields and template markers. Don't strip an empty thinking block because the next request has thinking disabled. Compare the actual rendered history across cold and returning requests.
Timing profile: temperature 0. Mixed sampled profile: temperature 1.0, top_p 0.95, top_k 20, min_p 0, repetition penalty 1, presence/frequency penalty 0. Medium thinking was tested separately where supported. The total 8192-token output budget includes reasoning; a separate reasoning budget is another feature to verify.
5. Capture a row readers can compare
{
"context_limit": 262144,
"actual_prompt_tokens": 253952,
"requested_output_tokens": 8192,
"actual_output_tokens": 8192,
"finish_reason": "length",
"completed_requests": 1,
"requested_concurrency": 1,
"observed_decode_overlap": 1,
"cache_state": "cold",
"wall_seconds": 59.94,
"wall_output_tok_s": 136.68,
"ttft_seconds": 38.53
}Example values are rounded from a measured tuned EXL C1 cell. Add source/config/model hashes, sampler, reasoning count, cache hits, memory minima, raw request/SSE references and failures to the internal record. Publish only reviewed, privacy-safe fields.
Wall throughput = completed actual output tokens / entire wave duration. Decode-only throughput uses a different interval. Don't put both under an unlabeled "tok/s" column. Queued clients don't establish simultaneous decode capacity.
6. Qualify the selected configuration
Test actual tools, tool-result follow-up, schema objects and thinking replay. Then short, 100K, 128K and full-window requests, cold/warm history, code and fictional operations, concurrency beyond five, churn, overload, cancellation and recovery.
Keep PASS, FAIL, PARTIAL, INVALID and NOT RUN separate. Refine concurrency around a measured plateau. A latency threshold may end a service-profile search without proving the maximum possible throughput.
This campaign's tuned EXL failed the overload fairness probe: a 4K request waited 226.34 seconds behind twelve long requests. No admission policy is supplied as a verified fix.
7. Recompute the public data locally
Unzip the download, then run:
python3 verify_data.py full-window.csvThis standard-library check requires nine completed C1 rows, verifies each 8,192-token output and recomputes output divided by wall time. It also checks each prompt budget against the profile's native reserve. It verifies the published arithmetic; the private raw archive isn't in the download, so this isn't an independent rerun of the campaign.
For an authorized experiment, collect the row fields in section 5 from actual API usage and client wall timing. Check tools and history before timing, then compare cold, exact returning-prefix and overload rows separately. No queue-policy fix or new GPU run was tested for this package.
8. Prepare a reversible canary
A production canary requires separate approval. Capture the existing production image, model hashes, config, consumer history and routing before changing anything. Qualify long/short admission and queue deadlines with the same overload workload. Recheck the actual rig's power, memory margin and consumer behavior.
Rollback plan: stop new experimental admission, drain or cancel its owned requests, restore the captured baseline target and configuration, then run the baseline consumer checks. An experimental rental's frozen EXL configuration isn't automatically the production baseline.
The production rig remains unchanged. Rental destruction and verified evidence preservation were completed separately.