Monday, September 28, 2026

Local LLMs for Delphi Migration — Follow-Up: Swapping the Inference Engine, Not the Model

This is a follow-up to the three-part series on running local LLMs against a structured, five-phase (AT1–AT5) Delphi migration benchmark, and to the follow-up comparing qwen3.6:35b-a3b against KAT-Coder-V2.5-Dev on the same weights class. This post holds the model constant and swaps the inference engine instead: same GGUF-class model, same benchmark harness, Ollama against a purpose-built CUDA engine called NInfer.



A correction up front, because it affects a number this series already published. The KAT-Coder follow-up stated that both qwen3.6:35b-a3b and KAT-Coder-V2.5-Dev held “100% GPU utilization with no CPU offloading” at NumCtx=131072. New validation data gathered for this post shows qwen3.6:35b-a3b running at ~91% GPU under that configuration on the current hardware — not the fully GPU-resident state previously reported. We're flagging this because the series' own standard is to publish a correction the moment a later measurement contradicts an earlier one rather than let a wrong number stand uncorrected.


What Changed: The Engine, Not the Model

Everything so far in this series ran on Ollama. This round adds NInfer, a C++/CUDA inference engine built specifically for the RTX 5090 (Windows forknatpate/ninfer-windows, portable zip, no install). It does not use GGUF — it has its own model format (.ninfer), and the model under test here is Qwen3.6-35B-A3B converted to that format at “groupwise-int” quantization (20.6 GiB on disk, KV-cache in bf16, 150k context). The Ollama side of the comparison is the same qwen3.6:35b-a3b Q4_K_M this series has used throughout.

The benchmark harness itself did not change to make this comparison possible: the existing run_benchmark.ps1 gained one additional code path, -Engine ninfer, which talks to NInfer's OpenAI-compatible API. The Ollama path is untouched. Same test cases, same prompts, same scoring functions. Settings held constant across both engines: NumCtx 131072 for AT1–AT3, 4096 for AT4, temperature 0, max 8192 output tokens, thinking enabled on both sides.

One asymmetry worth stating plainly: the Ollama scores in this comparison are the existing result files from earlier runs (AT1: 2026-05-20, AT2/AT3: 2026-08-03, AT4: 2026-05-18) — they were not re-measured for this post. Only the NInfer side is new. This is a deliberate reuse of already-validated numbers, not a re-run under identical wall-clock conditions; keep that in mind when reading the timing comparison below.


The Numbers: Ollama vs. NInfer, Same Model, Same Harness

Suite n (O/N) Ollama (all) NInfer (all) n (common) Ollama (common) NInfer (common) Unit
AT128/2841.654.62644.654.9Hit %
AT1-DE30/2838.542.62840.642.6Hit %
AT2101/11264.972.19369.974.4Hit %
AT330/3068.478.73068.478.7Structure+content %
AT490/900.8390.846900.8390.846Score 0–1
AT573/730.9961.000730.9961.000Score 0–1
AT660/600.7720.800600.7720.800Score 0–1

Why “n (common)” exists: NInfer's 150k context rejects two source files outright — one 453,190 characters, the other 463,994 characters — with an explicit context_length_exceeded error, costing it 10 tasks across AT1/AT2. Ollama does not reject the same prompts. It silently truncates them tonum_ctx: for all five AT1 tasks against the larger file, the returned prompt_tokens is exactly 131072 — the num_ctx value, not the file's actual token count. This means every prior AT-series result this blog has published for oversized files was scored against a truncated version of the source, not the file the reader would assume was analyzed. NInfer's hard rejection is, in this specific respect, the more honest failure mode. The “n (common)” columns above restrict both engines to tasks both could actually complete on the full, untruncated input — that is the fairer of the two comparisons in this table.

On quality: every suite improves or ties under NInfer, most visibly AT1 (+13 points) and AT3 (+10.3 points). AT4/AT5/AT6 — already near ceiling — move by a point or two, consistent with there being little room left to gain.


Throughput: The Actual Headline

SuiteOllama total wall timeNInfer total wall timeSpeedup
AT1912 s192 s4.75×
AT24,076 s977 s4.17×
AT3622 s137 s4.54×
AT4403 s99 s4.07×
AT5668 s53 s12.6×
AT6422 s94 s4.49×

And at the token level, mean (median) tokens/sec per suite:

SuiteOllama tok/sNInfer tok/s
AT1133 (137)412 (491)
AT2131 (135)449 (490)
AT3155 (159)694 (696)
AT4135 (135)541 (531)

A fairness note that cuts in NInfer's favor, not against it: the Ollama figure is a pure decode rate (eval_count / eval_duration, prefill excluded). The NInfer figure is output tokens divided by total wall time — prefill, network round-trip, everything included. NInfer's real decode-only rate is therefore higher than the number shown here; the ~4–5× gap above is, if anything, an underestimate of the engine's raw speed advantage. Isolated single-request samples against NInfer reached 460–620 tok/s.

Parallelism does not help much on either side at this hardware tier. Running 8 or 16 concurrent requests against NInfer barely increases aggregate throughput — a single request already saturates the RTX 5090 for a model this size. If your workload is “one interactive user,” this is not a limitation; if you were hoping to batch many requests per GPU, it is worth knowing before you plan around it.


Bonus: The VT Pre-Checks (VT1–VT9)

Beyond the AT-series, the harness runs nine smaller pre-checks (instruction discipline, JSON-schema fidelity, reproducibility, context isolation, tool-name alignment, Delphi-migration-risk keyword detection, tool-use-over-guessing, thinking on/off, and large-scale tool naming). Every VT that has a pass/fail verdict passes identically on both engines. Two results stand out beyond pass/fail:

  • VT9 Phase 1 (tool naming against 198 MCP tools): Ollama scores PARTIAL (exact 28 / semantic 85 / different 81 / missing 1), NInfer scores PASS (exact 51 / semantic 122 / different 25 / missing 0) — a real quality gap on this specific task, not just a speed difference.
  • Wall-clock on VT1–VT7: NInfer finishes in roughly a tenth to a hundredth of Ollama's time (e.g. VT1: 13.5 s → 0.1 s, VT7: 12.8 s → 0.2 s) — consistent with the AT-series speedups, and a reminder that for short, low-token-count prompts the fixed overhead Ollama carries matters more than raw decode speed.

Follow-Up: Qwen3.8-27B — Thinking Disabled, Plus INT8/252k Context

The two threads left open above — the dense Qwen3.8-27B numbers and the INT8/256k-context option — are now measured, not estimated. The earlier *_ninfer_qwen3.8-27b-nvfp4.json files (run with thinking) are superseded and were not used below: at temperature 0, thinking made Qwen3.8 run to the 8192-token output limit without ever answering in 14 of 30 AT3 tasks. Both new NInfer runs disable thinking (enable_thinking=false); without it, AT3 answers average 238 output tokens instead.

Setup. Ollama side: qwen3.8:27b-q4_K_M, existing results from 2026-08-29/30, not re-measured for this post. That run already used a large context (prompts up to 188k tokens, no silent truncation observed) and /no_think in the prompt, so it did comparatively little thinking either. Whether Ollama offloaded to CPU in that run was not checked — unlike the qwen3.6:35b-a3b correction earlier in this post, no claim is made either way here.

NInfer side: v0.9.0, model qwen3_8_27b_nvfp4 (.ninfer, NVFP4), both runs without thinking, temperature 0:

  • Run “bf16”: KV-cache in bf16, 150k context (default startup script).
  • Run “INT8”: --kv-dtype int8, --max-context 252928, truncation threshold 212992 to match the Ollama run.

Results (Ollama / NInfer bf16 / NInfer INT8; time = full suite):

Suite Ollama NInfer bf16 NInfer INT8 Unit Ollama time bf16 time INT8 time
AT163.665.7 (28/30)62.8 (30/30)Hit %4,309 s467 s449 s
AT1-DE49.738.7 (28/30)36.2 (30/30)Hit %1,616 s348 s320 s
AT279.8 (113 scored)72.8 (112 scored)71.7 (120 scored)Hit %18,116 s1,865 s2,026 s
AT385.380.982.3Structure+content %3,335 s53 s50 s
AT40.8020.8260.820Score 0–1467 s63 s56 s
AT51.0001.0000.996Score 0–1207 s46 s37 s
AT60.8060.8170.817Score 0–1279 s41 s33 s

AT1–AT4 combined: Ollama 26,227 s vs. NInfer ≈2,450 s — roughly 10× faster, on top of an already-fast model class.

On AT1 and AT2, the “all tasks” row above hides three different denominators, and that's worth spelling out rather than glossing over: AT1 and AT2 both have 120 (AT2) or 30 (AT1) nominal tasks, but each engine/run scored a different subset for a different reason. NInfer INT8 answered everything (120/120, 30/30). NInfer bf16 rejected 8 of AT2's tasks outright with context_length_exceeded — the two oversized files from earlier in this post — leaving 112/120. Ollama scored only 113/120 on AT2, but for a different reason: seven specific calls returned empty responses with prompt_tokens=0 and no error text — not context-size-related, since none of those are oversized files, and the exact cause (likely a timeout or abort during the 2026-08-30 run) can no longer be determined from the stored results. Because the three “all tasks” numbers aren't comparing the same task set, the paired comparison restricted to the 98 tasks all three completed is the cleaner AT2 figure: Ollama 81.8, NInfer bf16 73.4, NInfer INT8 71.7. The equivalent restriction for AT1 (26 shared tasks) gives 65.9 / 66.1 / 64.8.

VT pre-checks: VT1–VT8 pass/fail identically to Ollama on both NInfer runs, 3–100× faster (e.g. VT7: 48.4 s → 0.5 s). VT9 Phase 1 (exact tool-name matches out of 198 MCP tools): Ollama 65, NInfer bf16 76, NInfer INT8 77. VT9v2 has no built-in quality scoring — it's a timing check only (411 s vs. 15 s / 12 s) and should not be read as a quality result.

Token rates (mean per suite, AT1/AT2/AT3/AT4): Ollama 33/29/10/68, NInfer bf16 48/72/128/110, NInfer INT8 82/89/155/121 tok/s. The same caveat as the Qwen3.6 numbers above applies, more sharply here: Ollama's figure is decode-only; NInfer's is output tokens over total wall time including prefill — and with thinking disabled, answers are short, so prefill dominates NInfer's wall time more than in the Qwen3.6 runs. Use the suite wall-clock times above for the speed comparison, not these per-token rates.

Three takeaways:

  1. Speed: NInfer is roughly 9–10× faster than Ollama on Qwen3.8 at the same task volume.
  2. Quality is mixed, unlike Qwen3.6. AT1, AT3–AT6, and the VT suite are tied within noise. AT2 is the exception — NInfer trails Ollama by about 8 points (paired comparison), and AT1-DE by about 13 points, NInfer's weakest result in this whole post. This is the opposite pattern from the Qwen3.6 comparison earlier, where NInfer matched or beat Ollama on every metric. Possible causes — NVFP4 vs. Q4_K_M quantization, a chat-template mismatch, or something specific to German-language prompts — have not been investigated; treat this as an open question, not a diagnosed regression.
  3. INT8 KV-cache at 252k context costs essentially nothing and buys real headroom. Quality moves by at most ±1.5 points against bf16 (sometimes better, sometimes worse), speed is even or slightly faster, and — the actual point of testing it — no task is rejected anymore. The bf16/150k run hit context_length_exceeded on 12 tasks (AT1: 2, AT1-DE: 2, AT2: 8); the INT8/252k run answered all of them, including prompts up to 188,232 tokens. The model vendor's own README says this model “should stay at 150k” — the INT8 KV-cache evidently buys the full advertised context back without a quality cost worth mentioning.

For the headline comparison in this section, INT8 vs. Ollama is the one to read; bf16 is included as a secondary data point specifically to isolate what INT8 changes (quality: negligible difference; context ceiling: the whole story).


What This Means in Practice

The quality delta for Qwen3.6 should not be over-read: it is the same model, converted to a different runtime format and quantization scheme (groupwise-int vs. Q4_K_M), so some of the AT1/AT3 gain plausibly comes from quantization, not the engine per se — this post does not have a same-quantization, engine-only isolation to separate those two variables, and that is a gap worth closing before drawing firm conclusions about NInfer's format specifically. What is unambiguous is the throughput number: 4–10× faster wall-clock on identical benchmark suites, on the same GPU, across two different model families, with no quality regression on most metrics measured. For a pipeline currently paying Ollama's decode rate, that is a substantial, reproducible win, and the integration cost is low — OllamaMCP 1.6.0 already routes any model name prefixed ninfer/ to NInfer's OpenAI-compatible endpoint, so existing callers (email triage, chat) switch by changing a model name, not by rewriting integration code. A new NInferMCP module starts, stops, and checks the status of the NInfer server headlessly, without a console or web UI dependency.

The truncation finding deserves to travel independently of the speed story: if you have been benchmarking or running any local model against source files near or above your configured num_ctx, check whether your engine is silently truncating rather than telling you. Ollama's behavior here — accepting the oversized prompt and returning prompt_tokens capped at exactly num_ctx with no error — is easy to miss and directly invalidates any benchmark result computed on the truncated remainder. It should be looked for in any pipeline handling large Delphi units, not just this one.

What this post does not claim: that NInfer's .ninfer quantization format is categorically better than GGUF Q4_K_M — that requires an engine-only comparison at matched quantization, which does not exist yet. The dense Qwen3.8-27B result above is measured, not preliminary, but it opens its own open question rather than closing one: why NInfer trails Ollama specifically on AT2 and AT1-DE for this model, when it matched or beat Ollama on every metric for qwen3.6:35b-a3b. That is a genuine gap in this post, not a hedge — worth a targeted follow-up rather than a guess here.


Note on the Use of AI: This article was created with the "assistance" of generative AI. The content, technical statements, and conclusions have been reviewed and revised by the author. The author bears full responsibility for the publication.


Thursday, September 3, 2026

DelphiHg: A Native Mercurial Client in Delphi — Built in a Day, With a Little Help from AI

This post is not only about Mercurial. It is about what is possible today with Delphi — especially when combined with AI.


A Short Word on Python

If you know me, you probably know that I don’t like Python. Or, to put it more precisely: I hate Python. For two reasons.

The first one is nostalgic, and I admit it freely: because of a spectacularly dumb management decision, Turbo Pascal — the number one teaching language at schools and universities — was replaced by Python.

The second reason is that everyone seems to believe Python is fast and wonderful, especially for AI. But let’s be honest here: it is the libraries written in C and C++ that make everything fast. Python is the wrapper around them.

So when I needed a Mercurial interface, going through Python was not where I wanted to end up.


What I Actually Needed

I needed a Mercurial interface for three different consumers at once:

  • my backend AI MCP server,
  • my own Ollama console, the one that works with my environment,
  • and a Delphi IDE plugin.

One native interface, three consumers. That is the kind of problem where “just shell out to hg.exe” stops being a good answer.


Asking the Obvious Question

So I opened a new chat with ChatGPT and asked exactly one question — nothing else, no context, no priming:

“What do you know about the sources of Mercurial?”

And the answer came back:

If by “the sources of Mercurial” you really mean the internal source code of the version control system Mercurial (hg): it is fully open and remarkably interesting to analyse.

The current version is 7.2.4, released on 11 August 2026. The project is still actively developed.

1. Technically

Mercurial today essentially consists of three layers:

  • Python — still the largest part: business logic, commands, repository abstractions, merge, network protocols, and so on.
  • Rust — more and more performance-critical parts are being moved there.
  • C — older native accelerators and platform-specific parts.

The official documentation explicitly describes Mercurial as written in Python, C and Rust.

…and quite a bit more. But then came the sentence that made me sit up:

And I suspect what your question may be leading up to: Mercurial would in fact be a rather good candidate for deriving a native Delphi VCS core from it, or for rebuilding it compatibly.

Wow. “He” knows me a little too well — to infer that from the question above.


How Long Would It Take?

A somewhat longer conversation later, I asked for an effort estimate:

ScopeEffort
Read POC: changelog + reading individual files10–15 agent hours
Full read support for repository history25–40 h
Read + working copy/status / working diff35–55 h total
After that: simple local write+30–50 h
Robust local write I would trust real repositories to+60–90 h
Near-complete Mercurial client incl. merge/push/pull/clone+150–250 h or more

The Usual Next Step: Turn It Into a Spec

Then came the usual request — turn all of this into a Markdown file — plus a few requirements of my own:

All the information you found should be included. But it must be made clear that:

  1. We are only building read support.
  2. We build against version 7.0.1 first, but it has to stay updatable later on.
  3. We need what I’ll call a “linkage”: when a version jump happens, we need to know exactly which information sources each unit was built from, so we don’t have to guess afterwards where something needs adjusting when TortoiseHg or Mercurial ships a new release.
  4. We may retrofit the write part later — first stage only add, revert and commit.
  5. We want to add the full write part at some point after that.
  6. All of this is written as a Delphi 13.1 64-bit DLL that has to be thread-safe.
  7. We have to respect Mercurial/TortoiseHg locking (switchable, if needed).
  8. The implementation happens in small tasks, all of which are created in the PlanMCP server up front.
  9. Parallel work should actually run in parallel — for that we need a matrix stating which task blocks the next step and which tasks can run concurrently (in sub-agents).
  10. Which sources we need to clone locally so we have a fast reference to look things up in.

I imported the resulting Markdown file into my PlanMCP server, had 68 tasks generated from it — and bingo: one day later, DelphiHg was able to read any Mercurial repository.

That “one day” is not a figure of speech. The tasks were created on 31 August 2026 at 13:59. The first one was closed eleven minutes later. By the end of that same day, HG-R-010 through HG-R-056 were done: DLL skeleton, types, binary reader, I/O, revlog index, revlog reconstruction, all three codecs, delta, hash, requirements, store paths, changelog, manifest, filelog, DAG, dirstate v1 and v2, ignore, status. That is the complete read path.

The full span was three calendar days. 1 September added the ABI, diff, tags, copy metadata, the test suites, the release gate and the performance baseline. 2 September was optimisation, and nothing else. Final state: 89 changesets, 32 units in src\, 483 tests, all green.


The Numbers — and Why Three Different Ones Exist

After the read path stood, we evaluated the runtime behaviour and measured. Before any number: there are three different comparisons here, and they produce wildly different factors. Quoting only one of them misleads.

QuestionComparisonFactor
“What does a command line invocation cost me?”hg.exe vs. HgCli.exe, both as a process5× to 15×
“What does an operation inside my application cost me?”hg.exe vs. DLL hot5× to 2637×
“How good is the code?”hg hot vs. DLL hot — both loaded as a library in a running process1.0× to 10.8×

Only the third one says anything about the implementation. The 866× and 2637× from the second are honestly measured, but what they mostly show is Python’s interpreter startup — a process start compared against no process start. I am not going to quote those alone, and neither should anyone else.

“Hot” means the process is already running, the library is loaded, and the command is executed repeatedly. On the Mercurial side via hg debugshell -c, where the mercurial module is loaded — which, incidentally, is exactly the situation TortoiseHg creates, since it loads Mercurial as a Python module into its own process.

The table that belongs on the outside

A real repository: 2,922 revisions, 7,814 files in tip, 1.08 GiB store, zstd. Both sides hot, both measured the same evening, 5 repetitions each (cat -r tip 3), median.

Operationhg hotDelphiHg (DLL hot)Ratio
open (open repository)1.909 ms0.176 ms10.8× faster
manifest -r tip (7,814 entries)6.162 ms6.022 mslevel
files -r tip6.378 ms4.705 ms1.36× faster
log (interpret 2,923 changesets)85.560 ms26.419 ms3.24× faster
cat -r tip (7,814 files, 1.7 GB)8,702.9 ms5,490.6 ms1.58× faster

No measured operation is slower than Mercurial. Before the final optimisation round, manifest -r tip was the exception at 22.24 against 5.56 ms.

The machine: Windows 11, AMD Ryzen Threadripper PRO 7955WX, running in a VMware VM with 8 logical processors, Delphi XE16/Win64 Release, Mercurial 7.0.1, FastMM5. Every series carries a control measurement — 60 million rounds of xorshift64, no memory, no file, not one line of production code — calibrated at roughly 74 ms on this machine. That evening it ran at 81–83 ms, so the machine was about 10% slower than on calibration day. Which counts against DelphiHg, not for it.

Method, briefly: one warm-up run that does not count, then 3, 5 or 9 counted runs; median, not mean; the repository is reopened for every repetition (on Mercurial’s side too, a fresh hg.repository() each time), otherwise you are only measuring caches; before/after comparisons run balanced-alternating, so a drift in the machine shows up. And: if the individual values of two versions overlap, the result counts as unproven — even when the medians differ.


It Started 2.81× Slower

This is the part that usually gets left out. The first working version was not fast. The performance baseline on 1 September measured it honestly — reading all files of a revision:

StateDelphiHg cat -r tiphg netRatio
First working version163 ms58 ms2.81× SLOWER
after HG-R-07149 ms68 ms0.72
after HG-R-08232.7 ms79.9 ms0.41

The baseline had three findings: cat was about three times slower than Mercurial’s C extensions, concurrency barely scaled at all, and the process-start advantage was large but said nothing about the code.

And this is where Fable 5.1 showed up, at exactly the right moment — right after that evaluation of the runtime behaviour. The entire first optimisation round, HG-R-071 through HG-R-085, is its work. Everything in this section and the next one is what it found.

The single biggest item was not exotic at all: ManifestOf reconstructed the same manifest again for every single file — with 253 files, 252 times for nothing. That was 128 of 177 ms.

The rest of the round was similar in character. Delphi’s TZDecompressionStream decompresses every chunk twice because it CopyFrom(…, 0) asks for Source.Size, the stream does not know its size, and decompresses everything just to answer. Inline revlogs were read twice. Every delta series started at the snapshot instead of the remembered predecessor — turning roughly 24,000 delta applications across 400 manifests into about 400, worth −70.2% on that measurement.

One negative result worth keeping: SHA-1 ran at 272 MB/s through THashSHA1, so we moved it to Windows CNG. That got us to 807 MB/s — a factor 3, where 6 was expected. SHA-NI is present on this CPU, and Windows uses it for SHA-256 (2,046 against 791 MB/s), but Windows does not accelerate SHA-1 with it. Single-threaded, 9–29% of the gain arrives. Eight-threaded, nothing arrives. Which brings us to the actual story.


The Bottleneck Was Never Our Code

Three optimisation steps in sequence, all measured on eight threads sharing one repository instance:

  • HG-R-075 removed allocations → −37%
  • HG-R-076 removed computation → 0%
  • HG-R-078 removed allocations again → 0%

The difference between the first and the third was size: 110 KB per chunk in the first case, a few hundred bytes in the third.

Then came the number that settled it. The eight-threaded read measurement took 123.1 ms. A control measurement that does nothing but allocate and free memory took 124.1 ms. The read path was exactly as expensive as doing no work at all — every byte we read was free, and we were paying entirely for the allocator.

The allocation histogram killed the obvious theory immediately: not a single large allocation (>260 KB). The assumption “we produce big blocks” was simply wrong. 78–98% are small (<2.6 KB), and the most frequent bucket is 32–63 bytes. Scaled from 1 to 8 threads:

1 thread8 threadsFactor
Medium blocks (4–68 KB)12.61 ms127.10 ms10.1
Small blocks (24–88 bytes)1.08 ms110.01 ms102
Control (computation only)79 ms81 ms1.0

The reason: FastMM4 splits its 56 locks by size class, not by thread. If all eight threads request the same size — and that is precisely what a read path does when it keeps building the same structures over and over — everything squeezes through one lock.

Before touching a single line of production code, we measured the ceiling: a throwaway version with a per-thread bump pointer that never frees anything and could never ship, to find out what was even available. It gave −90%. Only then did the real work start. Switching to FastMM5, which keeps several arenas per size class and switches instead of waiting:

1 thread8 threadsSlowdown
One instance per thread, FastMM429.25 ms200.22 ms6.8
One instance per thread, FastMM529.43 ms51.04 ms1.7
Control (computation only)79 ms81 ms1.0

(Those are slowdowns from 1 to 8 threads, not speedups — smaller is better.)

Eight threads sharing a single instance went from 80.01 to 4.42 ms, −94.5% — below the ceiling the throwaway version had established (7.22 ms). Measured across the whole project, that number travelled from 879.8 ms in the first version to 4.42 ms.

The cost is honest: single-threaded, FastMM5 costs nothing at all (396.4 vs. 396.2 MiB peak working set). Under eight-way parallel load it costs +29% memory for −43% time.

While we were there, we also checked a piece of folklore. “TRTLCriticalSection is faster than TMonitor” — measured alternating, with the locking primitive as the only variable: 191.5 against 189.8 ms, and 68.8 against 67.8 ms. No difference; if anything TMonitor is ahead. Though the honest reading is “irrelevant at this point”, not “equally fast” — there are only 4,000–8,000 lock operations in 190 ms. A lock-free version was tried and crashed reproducibly (“revision 5 is 0 bytes long”): one thread read the revision of one state and the text of another.


The Last Round Disproved Its Own Premise

This is my favourite part, and it is the reason I am publishing the method along with the numbers. It is also where the models change hands: the first optimisation round above was Fable 5.1, this last one — HG-R-086, changesets 82 to 89 — was Opus 5.

A comparison table showed manifest -r tip costing us 22.24 ms against Mercurial’s 5.56 ms — 4.0× slower, the one place where DelphiHg lost. The diagnosis seemed obvious: our manifest parser cuts its own copy of the path for every entry, so 7,814 allocations, while Mercurial’s _lazymanifest holds views into one shared buffer.

Both halves of that statement were wrong.

First: manifest -r tip does not measure our parser. Broken into stages, each measurement doing exactly one step more than the previous:

StageMedianof which new
Open + read tip changeset0.46 ms0.46 ms
+ manifest fulltext reconstructed18.72 ms18.26 ms
+ parsed (= manifest -r tip)19.59 ms—
The parser alone, on an existing full text2.04 ms2.04 ms

The parser is 10%. Reconstruction is 93%.

Second: Mercurial’s 5.56 ms reconstructs nothing at all. ctx.manifest() Reads the full text straight out of .hg\wcache\manifestfulltextcache. Go around that file and do the same work, and Mercurial needs 8.6 to 10.6 ms. The honest comparison — reconstruction against reconstruction — is 18.26 against 8.6–10.6 ms, a factor of 1.8 to 2.1. And at the parser, the deficit we set out to fix is not measurable at all: 2.04 against ~1.5 ms, with the Mercurial values scattering between 13.3 and 19.4 ms. DelphiHg additionally verifies the sort order there, which _lazymanifest does not.

We had accused the wrong component, and we had compared against a cache hit.

What we actually changed: we now read Mercurial’s fulltext cache too (never write it — DelphiHg is strictly read-only, and Mercurial and TortoiseHg fill the file anyway). Because it is a foreign file that is allowed to be stale, every fulltext it hands us is verified against the node — length against the revlog index first, then the hash over all bytes. On failure, it silently walks the delta chain; a stale cache is not a data error. Mercurial does not verify at this point, which is fair enough: it wrote the file itself.

And we stopped copying the delta chain link by link. The obvious implementation allocates a full new text per link — at 651 links and 536 KB that is 349 MB of copying to produce half a megabyte. The intermediate state is now a list of segments that only record where their bytes come from; the result is written once, at the end.

Measurementbeforeafter
manifest -r tip (normal case, cache present)19.59 ms3.67 ms5.3×
manifest -r tip without the cache20.02 ms14.61 ms1.37×
— of which: applying deltas11.66 ms7.87 ms−32%
147 manifests, no cache hits1,323.71 ms988.94 ms−25.0%
2,923 manifests, full series15,738.59 ms12,849.29 ms−18.3%

Two things went wrong here, and both are more interesting than the wins.

The expected jump never arrived. Writing 536 KB instead of 349 MB should have been worth a factor 5 to 10. It was worth 32%. A built-in diagnostic counter explained it: the segment list grows to 2,333 segments, and at 651 links with around 1,200 segments on average, that is roughly 780,000 segment operations ≈ 7.8 ms — against 7.87 ms measured. The time was never in copying bytes. It was in managing the list. Pairwise folding, the way Mercurial’s mpatch does it, would turn O(links × segments) into O(segments × log links), a factor of 35 on that item. We deliberately did not do it, because the item only occurs on long chains without a cache hit.

The first version cost 218 MiB. Folding across the entire chain pushed the peak working set of cat -r tip from 395.7 to 613.4 MiB — in exchange for 5.9% time. It now folds in sections, holding at most 4 MB of delta at a time. What that cap costs is unevenly distributed: on the manifest, nothing (the limit never triggers); on cat -r tip, +245 ms, which is nearly the entire gain there.


The Correction That Went Against Us

One more, because our benchmark series has been criticised for not having enough of these.

hg cat -r tip on a big repo measures 29,119 ms as a command. That number contains 1.7 GB of pipe output — the command writes every file’s content to stdout, while our measurement only sums the lengths. The honest comparison for the same work is the hot figure: 8,703 ms.

Our lead on that operation therefore shrank from 5.3× to 1.58×. The comparison had been skewed in our favour, we found, and the corrected number is the one in the table above.


What Was Not Measured

  • TortoiseHg — not measured at all. It appears in our documentation only as a use case. Not a single number in this post comes from TortoiseHg.
  • status — implemented and tested byte-exact against hg statusbut there is no performance measurement. From code analysis, we know it does more than Mercurial does (a stat per tracked file plus a full tree walk, descending into ignored directories, content comparison via revlog reconstruction rather than hashing the file on disk). That is read, not measured.
  • Working-directory diff — not measured. What is measured is diff between two revisions: 32.08 ms on the synthetic repository.
  • Concurrency on a real repository — every parallel number above comes from the synthetic repository with 253 files.
  • Lock acquisition under contention as a timing — functionally tested (a foreign hg commit succeeds while DelphiHg reads and holds files open, read consistency is preserved, a held lock is detected and can be overridden on request), but never timed.
  • Other operating systems, CPUs, or Delphi versions — none.
  • Writing operations — DelphiHg is strictly read-only. There are none.

And one caveat that deserves its own paragraph, because it is the one most likely to mislead: on the synthetic repository, the findings of the last round do not exist. 253 files, short delta chains, no populated full-text cache — the DLL column there did not move at all. Everything described above only became visible on a grown repository with a 651-link delta chain. Anyone repeating this on a toy repository will measure nothing and conclude we made it up.

Further caveats worth stating plainly: the machine is a VM; the control measurement ran 10% above calibration all evening; one repository is one profile (zstd, 2,922 revisions, 7,814 files — treemanifest is not read at all and was not part of any measurement); the fulltext cache holds four manifests and is not always there, so the 5.3× applies to the most common case, not every case; and where individual values overlap, we report “unproven” rather than a percentage. In this data set, an unproven result counts for exactly as much as a percentage.

We also nearly published a measurement error. The first DLL run after the last optimisation came out “essentially unchanged” (21.55 instead of 22.24 ms). The cause: HgCli.exe loads the DLL from its own directory, and the copy sitting there was a day old. We had measured the state before all the changes.


What This Means in Practice

DelphiHg reads any Mercurial repository — natively, from a Delphi 13.1 64-bit DLL, thread-safe, with no Python in the process and no hg.exe being spawned. Three consumers use it through the same interface: the backend AI MCP server, the Ollama console, and the IDE plugin.

The performance answer is less spectacular than the big factors suggest and more useful than I expected: measured as a library against a library, DelphiHg is between level and 10.8× faster, and no measured operation is slower. The 866× number is real, but it mostly measures Python starting up, and I would rather publish the 1.58× that survived a correction against us than the 5.3× that did not.

The part I would take away if I only got one thing: the bottleneck was the memory manager, not our code. Three optimisation rounds went into the read path before a single measurement pointed at the allocator, and the number that finally pointed there was an eight-threaded read costing exactly as much as doing nothing but allocating and freeing. Everything after that was cheap — the FastMM5 switch is a one-line change worth −94.5%.

And the methodological one: the last round started from a task that was wrong. A parser was accused of a 4× deficit; the parser turned out to be 10% of the measurement, and the number it was being compared against was a cache hit rather than a reconstruction. Nobody had been careless. The measurement was simply asking a different question than everyone assumed. That is worth a warm-up run, a median instead of a mean, and a rule that overlapping values count as unproven — all three earned their keep here.

Turbo Pascal is still gone from the universities. But the read path of a version control system written in Python, C, and Rust now has a native Delphi implementation that matches or beats it on every operation we measured — built in three days, against a spec that started with one question to a chatbot.

So far, all of this is read-only. But the 35 follow-up tasks are already running, of course: the goal is a full tool that not only reads any Mercurial repository, but also carries the add and commit operations and the other functions hg.exe offers. Going by the experience so far — one day, maybe two.

And the longest part of the development so far was not the code at all. It was the speed measurements: for every run I had to stop all four to six Claude instances first, so that the CPU was “undisturbed” for the tests.


Note on the Use of AI: This article was created with the "assistance" of generative AI. The content, technical statements, and conclusions have been reviewed and revised by the author. The author bears full responsibility for the publication.