• Sources: Simon Willison, HN discussion
  • Summary: Qwen 3.8 27B is released with open weights under the Apache 2 license. Simon Willison reports tool calling and vision working locally from a 17GB Q4_K_M quantized LM Studio build, and reports that the xhigh reasoning default over-thinks, with 21 minutes and 22,276 reasoning tokens spent on one SVG and a request for a circle answered with an animated geometric study. He measures 15 to 30 tokens per second and writes that performance is the only thing holding the model back from being a daily driver, and that it will be hard to win him away from hosted API models. The same post reports a throughput gain of roughly 72 percent from running with Multi-Token Prediction under llama serve --spec-type draft-mtp, benchmarked against the LM Studio default GGUF.
  • Why it matters: An Apache 2 model with working tool calling and vision runs on hardware engineers already own, and the throughput reported without Multi-Token Prediction is what keeps it short of replacing a hosted API for agent work.

send feedback on this story