Nearly two years after Mira Murati left OpenAI to start her own lab, the industry finally has something concrete to judge her by. Thinking Machines Lab released Inkling on July 15, its first foundation model, and put the full weights on Hugging Face under a permissive Apache 2.0 license. The model is now also live on OpenRouter, where developers can test it against the rest of the open-weight field rather than take the lab's word for it.
The numbers back up an unusually bold claim: Inkling is the strongest open-weights model built by a Western lab to date. On MCP Atlas, a benchmark that measures how reliably an AI agent completes real-world tasks using the Model Context Protocol, Inkling scores 74.1%, nearly 30 points above Nvidia's Nemotron 3 Ultra, making it the best-performing Western open-weights model on agentic tool use. That is a wide enough gap to matter for anyone building agents that need to reliably call external tools.
Strong on Agents, Still Behind China
The catch is that "best in the West" is not the same as "best, period." Chinese models GLM 5.2 and Kimi K2.6 still lead on several key benchmarks. On coding specifically, Z.ai's GLM 5.2 scores 82.7% on Terminal Bench 2.1, a benchmark measuring autonomous AI coding agents in a real terminal environment, against Inkling's 63.8%. Kimi K2.6 also leads on Humanity's Last Exam, a test of PhD-level scientific reasoning. Separate figures reported by VentureBeat show Kimi K2.6 ahead on GPQA Diamond, BrowseComp, and Humanity's Last Exam with tools, while Inkling proves more resilient on general chat instruction following, scoring 79.8% on IFBench compared to Kimi K2.6's 76.0%.
Thinking Machines is not pretending otherwise. In its own launch materials, the company was blunt about where Inkling stands in the broader field:
"Inkling is not the strongest overall model available today, open or closed."
Against its closest domestic rival, however, the gap tilts firmly in Inkling's favor. Inkling consistently outperforms Nemotron 3 Ultra across reasoning and coding, posting 97.1% on AIME 2026 and 77.6% on SWEBench Verified versus Nemotron's 94.2% and 70.7%, and it leads significantly in agentic workflows, scoring 74.1% on MCP Atlas against Nemotron's 44.7%. Thinking Machines also touts a token-efficiency edge: on one internal benchmark, the company says Inkling uses a third as many tokens as Nvidia's Nemotron 3 Ultra to hit the same coding performance.
The Price-to-Performance Question
Efficiency claims aside, Inkling is not cheap to run. Listed pricing on OpenRouter puts the model at $1 per million input tokens and $4.05 per million output tokens, while independent tracking from Artificial Analysis shows the higher "xhigh" reasoning-effort configuration costing $1.87 per 1M input tokens, which the firm labels expensive against an average of $0.43, and $4.68 per 1M output tokens against an average of $1.20. That is markedly pricier than the free-flowing Chinese open-weight models many Western developers have already defaulted to.
Accuracy is also a lingering concern. According to analysis from Artificial Analysis cited by The Decoder, Inkling is currently the most powerful U.S. open-weights model, outperforming competitors like Kimi K2.6 and DeepSeek v4 Flash max on agentic tasks while also demonstrating high token efficiency, but despite its strong benchmark performance, the model shows notable weaknesses in factual accuracy, with a hallucination rate of 63 percent, and comes at a higher cost than comparable Chinese models. For teams weighing Inkling against a nearly free Chinese alternative, that combination of higher price and higher hallucination risk complicates the pitch.
Related: Nvidia, Meta, Microsoft Urge Washington to Protect Open-Source AI
What Inkling Actually Is
Under the hood, Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41 billion active parameters out of 975 billion total, designed for general-purpose reasoning, coding, agentic and tool-use systems, retrieval-augmented generation, instruction following, and multilingual conversational applications, with native image and audio understanding. The model was trained on 45 trillion tokens across multiple modalities and supports a context window of just over 1 million tokens, with a maximum output of 262,144 tokens. It was also the first model from the start-up to be trained on Nvidia GB300 NVL72 systems, following a partnership agreement between the two companies earlier this year.
Rather than compete purely on raw intelligence, Thinking Machines has framed Inkling around flexibility. As the company put it in its release notes, Inkling reasons natively over text, images, and audio, and balances cost with performance through efficient and controllable thinking effort, having been trained to be a broad, balanced foundation model that is strong across many domains and flexible enough to adapt. Developers can fine-tune it directly through Tinker, the company's customization platform, and the company is also launching, in preview, Inkling-Small, a lighter-weight model with 12 billion active parameters that supposedly achieves strong performance with even lower cost and latency.
The Long Wait, Explained
The release closes an unusually long quiet period for a startup that raised money at extraordinary speed. When OpenAI's board fired Sam Altman in November 2023, Murati, then CTO, was named interim CEO; Altman was reinstated five days later, Murati returned to CTO, then left for good roughly 10 months after that, before founding Thinking Machines Lab in February 2025. Thinking Machines raised $2 billion at a $12 billion valuation in July 2025, then reportedly sought a $50 billion raise in November before those talks fell apart by January 2026. Investors in the seed round reportedly included Andreessen Horowitz and Nvidia.
For now, Inkling gives Western developers a genuine alternative to routing agentic workloads through Chinese open-weight models, even if it doesn't unseat them outright. Whether that's enough to justify the premium price tag will likely depend on how heavily a given workload leans on tool use and agentic reliability, where Inkling's MCP Atlas lead is most pronounced, versus raw coding or scientific reasoning, where GLM 5.2 and Kimi K2.6 still hold the edge.