Skip to content
Releases

LLMKube 0.10.0: making a Mac safe to put on the cluster

Christopher Maher

Christopher Maher

11 min read

LLMKube is the Kubernetes operator I build to turn a mixed pile of hardware (NVIDIA, AMD and Apple Silicon) into one inference platform. The Mac part has always felt a little like a magic trick. A small agent runs on the laptop, starts the engine, and the rest of the cluster calls it through a normal Service. Over the weekend I got a 27B model to 227 tokens a second that way. Then, before asking anyone at KubeCon to plug their own Mac in, I asked a boring question. Who can talk to the models on mine? The honest answer was anyone on my Wi-Fi. And anyone who could create an InferenceService could do quite a bit more than talk. LLMKube 0.10.0 is the release that fixes that. This is how it went, in order.

Almost nothing here is a new idea. Pinned certificates, per-tenant tokens and allowlists are old, boring security. The work was fitting them into how the agent already registers a Mac with the cluster, without making anyone change how they call a model. The TensorFold engine that made the weekend fast is Ash's work, and it ships in this release as a first-class runtime.

I started with an audit, and it was not flattering

I had an agent read the Mac code path end to end, then I checked every claim against the code by hand. Four findings mattered. None of them needed a clever attacker. They needed someone who could write an InferenceService in the namespace the agent watches, or someone on the same network.

  • Name takeover. The agent created or updated the Service named after the InferenceService without checking who owned it. Name a model after an existing Service and the agent would rewrite it and point its traffic at the Mac.
  • Model files from anywhere. The agent opened local model paths as given. The operator in the cluster had an allowlist for this; the agent on the Mac had none.
  • Any engine flag. Extra arguments went to the engine verbatim. Some engine flags serve a directory over HTTP, write logs to any path, or move the listener.
  • Open engines. Every engine listened on every interface with no authentication, admin endpoints included.

Adoption is still early, so I chose to fix all of this in the open, as hardening, and to publish one security advisory covering the four once the release is out. The first three are about what a model spec is allowed to ask for. The fourth is about who can reach the Mac at all.

Two guards were small. The third one fought back.

Ownership was the easy one. The agent now only touches a Service or EndpointSlice it labelled itself, and refuses the start otherwise, with a Warning event that says why. Model paths were next. They now have to resolve inside roots the operator lists on the agent, and the model store is always one of them. Symlinks are followed, because the Hugging Face cache is a forest of them. A .. anywhere in the path is refused, and so is a dangling link.

Engine flags took most of the week. My first plan was a denylist: refuse the dangerous flags and allow the rest. Every review round found another way to spell a refused flag. Python argument parsers accept any unique prefix. vLLM treats an underscore like a dash, reads --flag.key as a nested option, and rewrites some -O tokens into other flags. llama.cpp has its own quoting rules for comma-separated paths. After the fourth round I stopped trying to guess.

The fix was to flip it. Every flag each engine accepts is now pinned from that engine's own --help, typed as a switch, a value, or a path, and parsed the way that engine parses its argv. Anything the table does not know is refused. Paths still have to land inside the roots.

Engine (pinned version)Flags in the table
llama-server 0.5.0411
vllm-swift 0.4.2316
TensorFold v0.3.4.130
mlx-server8

There is an escape hatch, --allow-unsafe-extra-args, for a Mac whose model authors you fully trust. It relaxes most of the rules. It never relaxes the flags that move the listener or join a multi-node group.

The live test found a bug no unit test could

Every change got tested on my MacBook against the production cluster before I opened the pull request. The refusals all fired correctly. Then I noticed two of them firing every five seconds.

A refused service never gets an endpoint, so the operator kept marking it as waiting for the agent. That overwrote the agent's refusal in the status, which looked like a change to the agent, which refused again. The pair looped 17 times in 80 seconds. It turned out an old memory check had the same fight and nobody had seen it. The operator now leaves the agent's refusal reasons alone. The unit tests for each side were green the whole time, because the bug only exists when both run.

Locking the engines down meant building a way back in

The obvious fix for open engines is to bind them to 127.0.0.1. That is one line per engine. The hard part is that the cluster was reaching them by IP and port, and now it could not.

I wanted a Mac-backed model to behave exactly like a pod-backed one: same Service name, same port, and a NetworkPolicy that means the same thing. So the operator now runs a tiny relay pod for each Mac model and points the model's Service at it. The relay talks to one TLS listener the agent opens on the Mac, and the agent forwards to the engine on loopback.

client pod
  -> Service <model>          (unchanged name, port and ClusterIP)
  -> relay pod               (pins the Mac's certificate, presents a token)
  -> agent ingress :9443     (on the Mac, TLS 1.3)
  -> engine on 127.0.0.1

The agent makes its own key and never lets it leave the Mac. It publishes a fingerprint of its certificate in the cluster, and the relay refuses any certificate that does not match. The relay presents a token from a per-namespace Secret. The ingress checks, in order, that it serves the namespace, that the token is right for it, and that the path is on the engine's allowlist. Admin endpoints like llama-server's /slots are refused even for a valid caller.

The upgrade had to be boring too. The operator only switches a model to the relay once the agent on that Mac advertises its ingress. An old agent keeps working exactly as before. Roll an agent back and the operator hands the Service back to it.

Three things almost broke on the way to my Mac

The first came from writing the docs. My Foreman coding agents can run natively on the Mac and find a model's port by reading its EndpointSlice. After the switch, that slice belongs to the relay pod, so the lookup would have returned the wrong port. Nothing live used that path, but the next person to set it up would have hit it. The agent now publishes the engine's loopback port for it.

The second one would have been a real outage. My lab deploys pre-release builds with an Ansible playbook, and it only built a fresh image for the operator. The relay runs in the router-proxy image, and the released one did not know the new relay mode. The relays would have crash-looped, and the moment my Mac advertised its ingress, three models would have gone dark. I caught it while checking image tags before installing the agent.

The third was noise. Right after creating the relays, the operator logged warnings about updates that conflicted with Kubernetes' own writes. They cleared on their own, but they would scare anyone reading kubectl describe. It now skips updates that change nothing and retries quietly. TensorFold has no metrics endpoint, so the relay now answers that scrape with an empty page instead of a failure.

It works, and the cost is too small to measure

I ran the release candidate on my cluster and my MacBook before tagging it. From another machine on the LAN, the engine ports refuse connections and the ingress answers 401 without a valid token. Through the Services, chat and streaming work as before. Time to first token barely moved.

Time to first token, p50 (20 requests each)Direct to the Mac (before)Through the relay (0.10.0)
Qwen3.8-27B, TensorFold80.7 ms73.0 ms
Qwen3.8-27B Q4_K_M, llama-server79.8 ms84.4 ms
Qwen3.8-27B Q5_K_S, llama-server86.8 ms91.7 ms

Those differences go both ways, and they are inside the noise at this sample size. I also broke things on purpose.

What I didWhat happened
Rotated the relay token under live traffic290 of 290 requests succeeded
Probed again after the old token expired60 of 60 succeeded
Restarted the Mac agent on a new buildAll three models back in 19 s, no relay restart
Deleted a relay podServing again in 8 s

The relay pod test is the one honest weak spot. There is one relay per model, so losing it costs a few seconds. That is the same as a one-replica pod-backed model, and a second replica is on my list.

The rest of 0.10.0 is last weekend's work

TensorFold is now a runtime you can pick per model with runtime: tensorfold. Every Mac engine is selectable per InferenceService again, which a regression had quietly broken since June. A dead engine now fails in well under a second with its own log lines, instead of blocking the agent for two minutes. The full story of those is in the TensorFold post.

Upgrading: do these in order

Breaking releases don't happen often in LLMKube, but this is one of them. Most of the list is about not being surprised.

  1. Upgrade the Helm chart before any Mac agent. A new agent with an old operator leaves Mac models without endpoints. Keep the router-proxy image at the operator's version; the chart defaults already do. The operator can now create Secrets, for the relay token.
  2. List your model directories on the agent. Any local model outside the model store needs its directory in --allowed-model-roots. If your store links into the Hugging Face cache, add that too.
  3. Check your extraArgs. Flags the pinned tables do not know are refused, and the event names the flag. On llama.cpp 0.5.0, --no-mmap is one. --allow-unsafe-extra-args exists for fully trusted Macs.
  4. Upgrade the agent and open one port. Engines now listen on loopback only. Open the ingress port (9443 by default) in the macOS firewall, not the engine ports. Anything that called <mac-ip>:<engine-port> directly should call the Service instead.
  5. If you run Ollama, unset OLLAMA_HOST=0.0.0.0. The agent warns at startup if it can still reach Ollama or oMLX on the host IP.
  6. Rotate the relay token by updating the Secret in place. The agent accepts the old and the new token for ten minutes. Deleting the Secret revokes the token at once.
  7. If you cannot upgrade the operator yet, run the new agent with --legacy-direct-endpoints. It restores the old behavior, warns you about it, and goes away in a future release.

To see the switch happen, watch the events on a Mac model:

kubectl describe inferenceservice <name>
# RelayCreated, then ServiceAdopted, and the Service keeps its ClusterIP
kubectl get deploy,endpointslice | grep <name>

Where the receipts are

The full network model, the allowlist tables and the flag reference live in the macOS agent README. The pull requests, in the order they landed:

  • #1932: the agent never overwrites or deletes a Service it does not own.
  • #1939: allowed roots for local model paths.
  • #1941: the typed extraArgs allowlist, and the status loop fix.
  • #1944: the relay and the operator side.
  • #1945: loopback engines and the authenticated ingress.
  • v0.10.0: the release and its changelog.

If you put a Mac on your cluster with this and something surprises you, I want to hear about it. The Discord is open.

LLMKube LLMKube

Kubernetes for Local LLMs. Deploy, manage, and scale AI inference workloads with production-grade orchestration.

© 2026 Defilan Technologies LLC

Community

Built for the Kubernetes and AI communities

LLMKube is not affiliated with or endorsed by the Cloud Native Computing Foundation or the Kubernetes project. Kubernetes® is a registered trademark of The Linux Foundation.