• Sources: Thinking Machines, Hugging Face weights, HN discussion
  • Summary: Thinking Machines Lab, the company founded by former OpenAI CTO Mira Murati, released Inkling on 2026-07-15, its first model after about 18 months of stealth work. Inkling is a Mixture-of-Experts transformer with 975B total parameters and 41B active (256 routed plus 2 shared experts, 6 routed active per token), a context window up to 1M tokens, and native text, image, and audio input, pretrained on 45 trillion tokens of text, images, audio, and video. Full weights ship on Hugging Face under Apache-2.0 with an NVFP4 checkpoint for NVIDIA Blackwell, hosted APIs on Together, Fireworks, Modal, Databricks, and Baseten, and a preview Inkling-Small at 276B total / 12B active. The company states Inkling is not the strongest model overall and positions it on breadth, customization, and controllable thinking effort.
  • Comments: HN commenters call it the strongest Western open-weights model and welcome a long-context multimodal option, while several note it is roughly 30% larger than GLM 5.2 without clearly beating it and looks weaker at coding than at instruction following.
  • Why it matters: A frontier-scale open-weights multimodal model under a permissive license gives teams a customizable, self-hostable alternative to closed APIs and to the leading Chinese open models.
  • Follow-up: Watch independent benchmark reproduction of the vendor and blinded-eval figures, and the full Inkling-Small release.

send feedback on this story