Skip to content
Engineering

Trying a new engine on the Strix Halo: 41 minutes to first token became 3.5

Christopher Maher

Christopher Maher

12 min read

I sent a 250,000-token prompt to my Strix Halo and watched the clock. On the llama.cpp setup I had been running for weeks, the first token of the answer shows up after about 41 minutes. On Gufo, a new engine I had never run before, it showed up after three and a half. My first reaction was not excitement. It was to ask whether we were sure, because a number like that usually means the benchmark is broken. This post is how I found out, what broke next, and how the box ended up serving my coding agent on the new engine. It all runs under LLMKube, the Kubernetes operator I build to turn heterogeneous hardware (NVIDIA, AMD, Apple Silicon) into one LLM inference platform. The coding agents are just the tenant that pushes it hardest.

Almost none of the hard parts are mine. Gufo is the work of francescobozzo and fedeizzo, MIT licensed, and the reason this post exists. The prompt-cache fix in the second half of this story is theirs too: joryirving reported the bug, fedeizzo reproduced it and fixed it (with a related change from myml), and the fix shipped in a release before I had finished writing it up. unsloth made the GGUF quants and the MTP draft sidecar that both engines read. What I added is the Kubernetes shape, a pinned image, and a lot of measuring.

The box was serving at a quarter of its window

The Strix Halo is an AMD Ryzen AI Max+ 395 with 128 GB of unified memory and one integrated GPU. For most of September it served one model: Qwen3.8 Flash Next, a 125B mixture-of-experts model with about 6B active per token. The stack was upstream llama.cpp on Vulkan, a Q3 quant, and an MTP draft head, and I wrote it up as a lab build. It worked well. It also ran at 65,536 tokens of context, while the model supports 262,144 and my coding agent was configured to think it had all of it.

So the first job was boring: prove the same stack could hold the full window. It could. Same quality gate, same speed at short prompts, and it found a hidden phrase at 131K and 250K every time. The catch was depth. A cold 250K prompt took around 40 minutes before the first token. Agent sessions mostly replay a cached prefix, so that is not every turn. But any cache miss deep into a session would be brutal.

That is the number I wanted to move, and it is why I went looking at other engines instead of other flags.

Gufo read the files I already had

Gufo stood out for boring reasons first. It is MIT licensed, it is a native HIP engine built for AMD, and it reads exactly the files I already had mirrored: unsloth's Q4 GGUF shards plus the shared MTP sidecar. No new quant, no conversion step. Its own published numbers on this chip looked too good to be true, which is the other reason I wanted to measure it myself.

I built a runtime image for it in llmkube-runtimes, pinned Gufo by commit, and ran it through the same quality gate as the incumbent. It passed everything. The hidden phrase turned up at both depths, my 12-question battery came back 12 of 12 with thinking on and off, and LiveCodeBench scored 83 of 100 twice against the incumbent's 82 and 83.

The first roadblock was small. Gufo checks the model name in every request and answers anything it does not recognize with model_not_found. llama.cpp ignores that field, so my old setup had never cared what my agent sent. One flag fixed it:

--served-model-name qwen38-flash-next

The second roadblock was mine, in LLMKube. Gufo has its own command line, so it runs under the operator's generic runtime, where you bring the image and the arguments. It turned out the generic runtime could not use a Model to stage weights at all. The built-in runtimes get the download, the cache and the read-only /models mount for free. A generic service got none of it, so my first deployment mounted a hand-staged volume read-only and pointed modelRef at an unrelated Model only because the field is required. I filed that as #1961 and fixed it in #1962: an opt-in spec.stageModel that gives a generic service the same staging and hands the container the paths through environment variables. It ships in 0.10.2.

I asked whether we were sure, and the answer was yes

The first comparison was Gufo against the llama.cpp configuration I happened to be running. That is not a fair fight. I had tuned that config for 64K, and llama.cpp has knobs that matter at depth. If I was going to say "11x" out loud, I wanted llama.cpp at its best, not at my defaults.

So I swept it. Q3 and Q4 quants, batch sizes from 512 up to 2048, the same harness and the same box, then took the best result at each depth. A fifth config, Q4 at the largest batch size, never got going: the GPU ran out of memory for command submission and the device was lost on load. The best llama.cpp config did get faster at short prompts. At depth it barely moved.

PromptGufo, Q4llama.cpp Vulkan, best of 4 tunedTime to first token
8K1,343 tok/s361 tok/s6 s vs 22 s
32K1,286 tok/s307 tok/s25 s vs 104 s
131K1,213 tok/s137 tok/s108 s vs 952 s
250K1,184 tok/s101 tok/s3.5 min vs 41 min

Cold prefill, nothing cached. The 8K and 32K rows are medians of three runs. The deep rows are single runs, because one 250K llama.cpp run takes 41 minutes. That last row is 11.7x, and the tuned sweep is what lets me say it with a straight face.

Lab card: Qwen3.8 Flash Next at full 262K context on one AMD Ryzen AI Max+ 395 with 128 GB. Time to first token for a 250K-token prompt drops from 41 minutes to 3.5 minutes, 11.7 times faster. Cold prefill in tokens per second, Gufo Q4 versus the best of four tuned llama.cpp Vulkan configs: 1,343 vs 361 at 8K, 1,286 vs 307 at 32K, 1,213 vs 137 at 131K, 1,184 vs 101 at 250K. Checks: needle found at 131K and 250K, two agents at once, battery 12 of 12, LiveCodeBench 83 of 100.

Two honest caveats. Gufo is about 7% slower at decode on a single stream in my agent-style test, 33.5 tok/s against 36.2. And this is Gufo on Q4 against llama.cpp's best of Q3 and Q4, so it is an engine comparison at each engine's best, not a same-file race. The picture flips the moment two agents share the box, though. My llama.cpp server had one slot, so a second agent waited its turn. Gufo served both at once, at 26.5 and 28.1 tok/s each.

I will trade 7% of single-stream decode for an 11.7x shorter wait on a deep cache miss, every time.

Then every agent turn got slower than the last

The last gate before a swap is a soak: three hours of growing coding sessions against the endpoint, up to 240K of conversation, logging every turn. Gufo finished it with zero errors and zero stalls. It also finished only 102 turns, and the time to first token climbed with every one of them. At the deep end of a session the agent was waiting over three minutes per turn.

The log said why. The engine's cached-token count sat at 2,132 while the prompts grew toward 240K. Every turn was re-reading the whole conversation from scratch, which is exactly the cold-prefill cost I had just spent a day measuring.

My first theory was that Gufo's prompt cache simply did not work across turns. So I ran a small A/B: the same four-turn conversation on both engines, with thinking on and off. That theory died fast. With thinking off, Gufo reused the entire previous conversation on every turn, just like llama.cpp. With thinking on, it reused the first prompt and nothing after it. llama.cpp kept reusing either way.

Cached tokens per turnTurn 2Turn 3Turn 4
Gufo, thinking off1,6413,2514,860
Gufo, thinking on1,5761,5761,576
llama.cpp, thinking on1,5773,1154,649

Thinking was on because nobody turned it off. My soak harness does not send a thinking setting, and neither does my coding agent, so both get the server default. The mechanism is a nice piece of detective work. With thinking on, the model's reply contains reasoning that the client never sends back, so the next prompt stops matching what the engine cached partway through the last turn. Most of this model's layers are recurrent, and recurrent state cannot be rewound to an arbitrary position. llama.cpp keeps checkpoints and backs up to the last good one. Gufo, at that version, could only restore a snapshot that matched completely, so it fell all the way back to turn one.

Turning thinking off would have fixed the cache and quietly made the agent dumber. I did not want to buy speed that way.

Someone had already filed it, and the maintainer had already fixed it

When I went to Gufo's issue tracker, it was already there: gufo-org/gufo#335, filed by joryirving, describing long sessions re-prefilling around 100K tokens every turn. fedeizzo had reproduced it at 5K and 15K tokens and already had a stack of fixes open.

Their repros were short, and my problem was long. So the useful thing I could do was test the fix at the full window on this exact model. I built Gufo from main plus the open fix stack, with the same flags, and reran everything. The four-turn A/B now matched llama.cpp to within one token. I also wanted to know the restore was exact and not just fast, so I ran six conversations twice each at temperature 0: once restored from the cache, once with the cache disabled. The reasoning and the answer came out byte-identical in all six.

Then the stack merged, and I did it all again on Gufo main: the A/B, the exact-restore check, and the full three-hour soak. I posted the long-context numbers on the issue. fedeizzo replied with thanks for helping pin it down and a note that more cache improvements were on the way. Gufo v0.5.0, the first release with the fix, came out the same day the fixes merged.

Forty-eight times faster on the turns that hurt

Same harness, same box, same three hours, thinking on. The only change is the cache fix. The "after" run used Gufo main at 93af45d4, right after the fix merged, which is the same server code that shipped as v0.5.0.

Conversation sizeBefore the fixAfter the fix
0-20K7.0 s2.8 s
20-60K28.8 s3.1 s
60-120K67.2 s3.4 s
120-180K128.8 s3.8 s
180-240K199.2 s4.1 s
Turns in 3 hours102379

Those are median times to first token, per band. At the deep end that is 199.2 seconds down to 4.1, which I round down to 48x. The time to first token now barely grows with the conversation, which is what a working cache looks like. Decode did not change, 35 to 39 tok/s median either way.

Lab card: long agent sessions, no more re-prefill. Qwen3.8 Flash Next at 262K context with thinking on, multi-turn coding sessions on one Strix Halo. Time to first token at 180 to 240K of conversation drops from 199 seconds to 4.1 seconds, 48 times faster. Median time to first token by conversation size, Gufo v0.3.0 versus with the cache fix: 7.0 vs 2.8 seconds at 0 to 20K, 28.8 vs 3.1 at 20 to 60K, 67.2 vs 3.4 at 60 to 120K, 128.8 vs 3.8 at 120 to 180K, 199.2 vs 4.1 at 180 to 240K. 379 turns in 3 hours, up from 102, with 0 errors and 0 stalls.

One thing I checked before trusting it. Two of the 379 turns ended in a repeating tail when they hit the token limit. The soak runs at temperature 0, and greedy decoding with thinking on can loop. Since the restore check came out byte-identical, a cache defect does not explain them, and two in 379 is not a meaningful difference from zero in 102.

It is serving my agent now

The swap went in today. The image is Gufo v0.5.0, built on llmkube-runtimes main from #54 with GitHub build provenance, and pinned by digest. The InferenceService kept its name and port, so the Service and the gateway route did not change, and the coding agent now gets the full 262,144 tokens it was always configured for. I ran the four-turn A/B against the live endpoint afterward, thinking on, and it reused 1,576, then 3,114, then 4,648 tokens. Same as the lab.

The live deployment still mounts the hand-staged volume, because it runs LLMKube 0.10.1. When 0.10.2 ships, the same service moves to stageModel: true and a normal Model, and the volume goes away. The previous llama.cpp InferenceService is one apply away if I ever need to roll back.

If you have a Strix Halo

This is what I would tell myself at the start of the week.

  1. Measure your current engine at the depth you actually use. At 8K, llama.cpp and Gufo are 16 seconds apart. At 250K they are about 37 minutes apart, and you will not see that from a short benchmark.
  2. Tune the incumbent before you crown a challenger. Sweep quant and batch size, take the best per depth, and compare against that. Watch memory while you do it: the biggest batch size on Q4 lost the device on this box.
  3. Use Gufo v0.5.0 or later. Earlier versions lose the prompt cache on every turn when thinking is on, and most agents leave it on without telling you.
  4. Set --served-model-name to the id your clients send. Gufo rejects any other name, where llama.cpp quietly accepted anything.
  5. Probe readiness on /v1/models, not a TCP port. The port opens before the model has finished loading.
  6. Before trusting any engine with agent traffic, run a multi-turn cache check with your agent's real settings. Send a growing conversation and read the cached-token count each turn. If it stops growing, every turn is a cold start.
  7. On LLMKube 0.10.2 or later, use runtime: generic with stageModel: true. Before that, mount pre-staged weights read-only.

The full step-by-step build, with the Model, the InferenceService and every flag, is the lab page: Qwen3.8-Flash-Next on Gufo on one Strix Halo. It sits next to the llama.cpp Vulkan build it replaces, if you want to compare the two.

What is pinned, and where the receipts are

The runtime image, by digest, with Gufo v0.5.0 at commit 23cacbb9:

ghcr.io/defilantech/llmkube-gufo-rocm-gfx1151@sha256:3055ba7de07281f679f7974b8bc9ec7544dc84e7734fc757080db5deb5971c0f

Every number in this post comes from the raw benchmark files in my lab notes, measured on this one box with one harness. If you run Gufo or llama.cpp on a Strix Halo and your numbers disagree with mine, I would like to know. The Discord is open. Next up in the lab: DeepSeek V4.1 on the three DGX Sparks.

LLMKube LLMKube

Kubernetes for Local LLMs. Deploy, manage, and scale AI inference workloads with production-grade orchestration.

© 2026 Defilan Technologies LLC

Community

Built for the Kubernetes and AI communities

LLMKube is not affiliated with or endorsed by the Cloud Native Computing Foundation or the Kubernetes project. Kubernetes® is a registered trademark of The Linux Foundation.