This is a follow-up to the three-part series on running local LLMs against a structured, five-phase (AT1–AT5) Delphi migration benchmark, and to the follow-up comparing qwen3.6:35b-a3b against KAT-Coder-V2.5-Dev on the same weights class. This post holds the model constant and swaps the inference engine instead: same GGUF-class model, same benchmark harness, Ollama against a purpose-built CUDA engine called NInfer.
A correction up front, because it affects a number this series already published. The KAT-Coder follow-up stated that both qwen3.6:35b-a3b and KAT-Coder-V2.5-Dev held “100% GPU utilization with no CPU offloading” at NumCtx=131072. New validation data gathered for this post shows qwen3.6:35b-a3b running at ~91% GPU under that configuration on the current hardware — not the fully GPU-resident state previously reported. We're flagging this because the series' own standard is to publish a correction the moment a later measurement contradicts an earlier one rather than let a wrong number stand uncorrected.
What Changed: The Engine, Not the Model
Everything so far in this series ran on Ollama. This round adds NInfer, a C++/CUDA inference engine built specifically for the RTX 5090 (Windows forknatpate/ninfer-windows, portable zip, no install). It does not use GGUF — it has its own model format (.ninfer), and the model under test here is Qwen3.6-35B-A3B converted to that format at “groupwise-int” quantization (20.6 GiB on disk, KV-cache in bf16, 150k context). The Ollama side of the comparison is the same qwen3.6:35b-a3b Q4_K_M this series has used throughout.
The benchmark harness itself did not change to make this comparison possible: the existing run_benchmark.ps1 gained one additional code path, -Engine ninfer, which talks to NInfer's OpenAI-compatible API. The Ollama path is untouched. Same test cases, same prompts, same scoring functions. Settings held constant across both engines: NumCtx 131072 for AT1–AT3, 4096 for AT4, temperature 0, max 8192 output tokens, thinking enabled on both sides.
One asymmetry worth stating plainly: the Ollama scores in this comparison are the existing result files from earlier runs (AT1: 2026-05-20, AT2/AT3: 2026-08-03, AT4: 2026-05-18) — they were not re-measured for this post. Only the NInfer side is new. This is a deliberate reuse of already-validated numbers, not a re-run under identical wall-clock conditions; keep that in mind when reading the timing comparison below.
The Numbers: Ollama vs. NInfer, Same Model, Same Harness
| Suite | n (O/N) | Ollama (all) | NInfer (all) | n (common) | Ollama (common) | NInfer (common) | Unit |
|---|---|---|---|---|---|---|---|
| AT1 | 28/28 | 41.6 | 54.6 | 26 | 44.6 | 54.9 | Hit % |
| AT1-DE | 30/28 | 38.5 | 42.6 | 28 | 40.6 | 42.6 | Hit % |
| AT2 | 101/112 | 64.9 | 72.1 | 93 | 69.9 | 74.4 | Hit % |
| AT3 | 30/30 | 68.4 | 78.7 | 30 | 68.4 | 78.7 | Structure+content % |
| AT4 | 90/90 | 0.839 | 0.846 | 90 | 0.839 | 0.846 | Score 0–1 |
| AT5 | 73/73 | 0.996 | 1.000 | 73 | 0.996 | 1.000 | Score 0–1 |
| AT6 | 60/60 | 0.772 | 0.800 | 60 | 0.772 | 0.800 | Score 0–1 |
Why “n (common)” exists: NInfer's 150k context rejects two source files outright — one 453,190 characters, the other 463,994 characters — with an explicit context_length_exceeded error, costing it 10 tasks across AT1/AT2. Ollama does not reject the same prompts. It silently truncates them tonum_ctx: for all five AT1 tasks against the larger file, the returned prompt_tokens is exactly 131072 — the num_ctx value, not the file's actual token count. This means every prior AT-series result this blog has published for oversized files was scored against a truncated version of the source, not the file the reader would assume was analyzed. NInfer's hard rejection is, in this specific respect, the more honest failure mode. The “n (common)” columns above restrict both engines to tasks both could actually complete on the full, untruncated input — that is the fairer of the two comparisons in this table.
On quality: every suite improves or ties under NInfer, most visibly AT1 (+13 points) and AT3 (+10.3 points). AT4/AT5/AT6 — already near ceiling — move by a point or two, consistent with there being little room left to gain.
Throughput: The Actual Headline
| Suite | Ollama total wall time | NInfer total wall time | Speedup |
|---|---|---|---|
| AT1 | 912 s | 192 s | 4.75× |
| AT2 | 4,076 s | 977 s | 4.17× |
| AT3 | 622 s | 137 s | 4.54× |
| AT4 | 403 s | 99 s | 4.07× |
| AT5 | 668 s | 53 s | 12.6× |
| AT6 | 422 s | 94 s | 4.49× |
And at the token level, mean (median) tokens/sec per suite:
| Suite | Ollama tok/s | NInfer tok/s |
|---|---|---|
| AT1 | 133 (137) | 412 (491) |
| AT2 | 131 (135) | 449 (490) |
| AT3 | 155 (159) | 694 (696) |
| AT4 | 135 (135) | 541 (531) |
A fairness note that cuts in NInfer's favor, not against it: the Ollama figure is a pure decode rate (eval_count / eval_duration, prefill excluded). The NInfer figure is output tokens divided by total wall time — prefill, network round-trip, everything included. NInfer's real decode-only rate is therefore higher than the number shown here; the ~4–5× gap above is, if anything, an underestimate of the engine's raw speed advantage. Isolated single-request samples against NInfer reached 460–620 tok/s.
Parallelism does not help much on either side at this hardware tier. Running 8 or 16 concurrent requests against NInfer barely increases aggregate throughput — a single request already saturates the RTX 5090 for a model this size. If your workload is “one interactive user,” this is not a limitation; if you were hoping to batch many requests per GPU, it is worth knowing before you plan around it.
Bonus: The VT Pre-Checks (VT1–VT9)
Beyond the AT-series, the harness runs nine smaller pre-checks (instruction discipline, JSON-schema fidelity, reproducibility, context isolation, tool-name alignment, Delphi-migration-risk keyword detection, tool-use-over-guessing, thinking on/off, and large-scale tool naming). Every VT that has a pass/fail verdict passes identically on both engines. Two results stand out beyond pass/fail:
- VT9 Phase 1 (tool naming against 198 MCP tools): Ollama scores PARTIAL (exact 28 / semantic 85 / different 81 / missing 1), NInfer scores PASS (exact 51 / semantic 122 / different 25 / missing 0) — a real quality gap on this specific task, not just a speed difference.
- Wall-clock on VT1–VT7: NInfer finishes in roughly a tenth to a hundredth of Ollama's time (e.g. VT1: 13.5 s → 0.1 s, VT7: 12.8 s → 0.2 s) — consistent with the AT-series speedups, and a reminder that for short, low-token-count prompts the fixed overhead Ollama carries matters more than raw decode speed.
Follow-Up: Qwen3.8-27B — Thinking Disabled, Plus INT8/252k Context
The two threads left open above — the dense Qwen3.8-27B numbers and the INT8/256k-context option — are now measured, not estimated. The earlier *_ninfer_qwen3.8-27b-nvfp4.json files (run with thinking) are superseded and were not used below: at temperature 0, thinking made Qwen3.8 run to the 8192-token output limit without ever answering in 14 of 30 AT3 tasks. Both new NInfer runs disable thinking (enable_thinking=false); without it, AT3 answers average 238 output tokens instead.
Setup. Ollama side: qwen3.8:27b-q4_K_M, existing results from 2026-08-29/30, not re-measured for this post. That run already used a large context (prompts up to 188k tokens, no silent truncation observed) and /no_think in the prompt, so it did comparatively little thinking either. Whether Ollama offloaded to CPU in that run was not checked — unlike the qwen3.6:35b-a3b correction earlier in this post, no claim is made either way here.
NInfer side: v0.9.0, model qwen3_8_27b_nvfp4 (.ninfer, NVFP4), both runs without thinking, temperature 0:
- Run “bf16”: KV-cache in bf16, 150k context (default startup script).
- Run “INT8”:
--kv-dtype int8,--max-context 252928, truncation threshold 212992 to match the Ollama run.
Results (Ollama / NInfer bf16 / NInfer INT8; time = full suite):
| Suite | Ollama | NInfer bf16 | NInfer INT8 | Unit | Ollama time | bf16 time | INT8 time |
|---|---|---|---|---|---|---|---|
| AT1 | 63.6 | 65.7 (28/30) | 62.8 (30/30) | Hit % | 4,309 s | 467 s | 449 s |
| AT1-DE | 49.7 | 38.7 (28/30) | 36.2 (30/30) | Hit % | 1,616 s | 348 s | 320 s |
| AT2 | 79.8 (113 scored) | 72.8 (112 scored) | 71.7 (120 scored) | Hit % | 18,116 s | 1,865 s | 2,026 s |
| AT3 | 85.3 | 80.9 | 82.3 | Structure+content % | 3,335 s | 53 s | 50 s |
| AT4 | 0.802 | 0.826 | 0.820 | Score 0–1 | 467 s | 63 s | 56 s |
| AT5 | 1.000 | 1.000 | 0.996 | Score 0–1 | 207 s | 46 s | 37 s |
| AT6 | 0.806 | 0.817 | 0.817 | Score 0–1 | 279 s | 41 s | 33 s |
AT1–AT4 combined: Ollama 26,227 s vs. NInfer ≈2,450 s — roughly 10× faster, on top of an already-fast model class.
On AT1 and AT2, the “all tasks” row above hides three different denominators, and that's worth spelling out rather than glossing over: AT1 and AT2 both have 120 (AT2) or 30 (AT1) nominal tasks, but each engine/run scored a different subset for a different reason. NInfer INT8 answered everything (120/120, 30/30). NInfer bf16 rejected 8 of AT2's tasks outright with context_length_exceeded — the two oversized files from earlier in this post — leaving 112/120. Ollama scored only 113/120 on AT2, but for a different reason: seven specific calls returned empty responses with prompt_tokens=0 and no error text — not context-size-related, since none of those are oversized files, and the exact cause (likely a timeout or abort during the 2026-08-30 run) can no longer be determined from the stored results. Because the three “all tasks” numbers aren't comparing the same task set, the paired comparison restricted to the 98 tasks all three completed is the cleaner AT2 figure: Ollama 81.8, NInfer bf16 73.4, NInfer INT8 71.7. The equivalent restriction for AT1 (26 shared tasks) gives 65.9 / 66.1 / 64.8.
VT pre-checks: VT1–VT8 pass/fail identically to Ollama on both NInfer runs, 3–100× faster (e.g. VT7: 48.4 s → 0.5 s). VT9 Phase 1 (exact tool-name matches out of 198 MCP tools): Ollama 65, NInfer bf16 76, NInfer INT8 77. VT9v2 has no built-in quality scoring — it's a timing check only (411 s vs. 15 s / 12 s) and should not be read as a quality result.
Token rates (mean per suite, AT1/AT2/AT3/AT4): Ollama 33/29/10/68, NInfer bf16 48/72/128/110, NInfer INT8 82/89/155/121 tok/s. The same caveat as the Qwen3.6 numbers above applies, more sharply here: Ollama's figure is decode-only; NInfer's is output tokens over total wall time including prefill — and with thinking disabled, answers are short, so prefill dominates NInfer's wall time more than in the Qwen3.6 runs. Use the suite wall-clock times above for the speed comparison, not these per-token rates.
Three takeaways:
- Speed: NInfer is roughly 9–10× faster than Ollama on Qwen3.8 at the same task volume.
- Quality is mixed, unlike Qwen3.6. AT1, AT3–AT6, and the VT suite are tied within noise. AT2 is the exception — NInfer trails Ollama by about 8 points (paired comparison), and AT1-DE by about 13 points, NInfer's weakest result in this whole post. This is the opposite pattern from the Qwen3.6 comparison earlier, where NInfer matched or beat Ollama on every metric. Possible causes — NVFP4 vs. Q4_K_M quantization, a chat-template mismatch, or something specific to German-language prompts — have not been investigated; treat this as an open question, not a diagnosed regression.
- INT8 KV-cache at 252k context costs essentially nothing and buys real headroom. Quality moves by at most ±1.5 points against bf16 (sometimes better, sometimes worse), speed is even or slightly faster, and — the actual point of testing it — no task is rejected anymore. The bf16/150k run hit
context_length_exceededon 12 tasks (AT1: 2, AT1-DE: 2, AT2: 8); the INT8/252k run answered all of them, including prompts up to 188,232 tokens. The model vendor's own README says this model “should stay at 150k” — the INT8 KV-cache evidently buys the full advertised context back without a quality cost worth mentioning.
For the headline comparison in this section, INT8 vs. Ollama is the one to read; bf16 is included as a secondary data point specifically to isolate what INT8 changes (quality: negligible difference; context ceiling: the whole story).
What This Means in Practice
The quality delta for Qwen3.6 should not be over-read: it is the same model, converted to a different runtime format and quantization scheme (groupwise-int vs. Q4_K_M), so some of the AT1/AT3 gain plausibly comes from quantization, not the engine per se — this post does not have a same-quantization, engine-only isolation to separate those two variables, and that is a gap worth closing before drawing firm conclusions about NInfer's format specifically. What is unambiguous is the throughput number: 4–10× faster wall-clock on identical benchmark suites, on the same GPU, across two different model families, with no quality regression on most metrics measured. For a pipeline currently paying Ollama's decode rate, that is a substantial, reproducible win, and the integration cost is low — OllamaMCP 1.6.0 already routes any model name prefixed ninfer/ to NInfer's OpenAI-compatible endpoint, so existing callers (email triage, chat) switch by changing a model name, not by rewriting integration code. A new NInferMCP module starts, stops, and checks the status of the NInfer server headlessly, without a console or web UI dependency.
The truncation finding deserves to travel independently of the speed story: if you have been benchmarking or running any local model against source files near or above your configured num_ctx, check whether your engine is silently truncating rather than telling you. Ollama's behavior here — accepting the oversized prompt and returning prompt_tokens capped at exactly num_ctx with no error — is easy to miss and directly invalidates any benchmark result computed on the truncated remainder. It should be looked for in any pipeline handling large Delphi units, not just this one.
What this post does not claim: that NInfer's .ninfer quantization format is categorically better than GGUF Q4_K_M — that requires an engine-only comparison at matched quantization, which does not exist yet. The dense Qwen3.8-27B result above is measured, not preliminary, but it opens its own open question rather than closing one: why NInfer trails Ollama specifically on AT2 and AT1-DE for this model, when it matched or beat Ollama on every metric for qwen3.6:35b-a3b. That is a genuine gap in this post, not a hedge — worth a targeted follow-up rather than a guess here.
Note on the Use of AI: This article was created with the "assistance" of generative AI. The content, technical statements, and conclusions have been reviewed and revised by the author. The author bears full responsibility for the publication.


