- Sources: primary, discussion
- Summary: Apple's Virtualization.framework reports a conservative Metal capability profile to a macOS guest, so llama.cpp selects slower kernels than the host GPU can run. The post describes a process-scoped shim that changes two reported answers, the Apple family enum and the threadgroup memory limit, moving TinyLlama prompt processing from 432 to 4,787 tokens per second on one M1 Ultra host and reaching 98.25 percent of the bare-metal result. It publishes the build scripts, raw logs, and checksums, and states that the method relies on private, version-sensitive guest Metal behavior, that MLX-LM showed no gain, and that validation covers one host and guest pair only.
- Why it matters: The GPU penalty for running model inference inside a macOS VM comes from a reported capability profile rather than from the virtualized hardware.
- Follow-up: Whether the shim survives macOS updates and whether it reproduces on other host and guest pairs.
send feedback on this story