Drooid Logo
Back to story perspectives

Full Breakdown

Inkling Launch Marks First Open-Weight Frontier Model from Thinking Machines Lab

7/16/2026, 6:00:55 AM

Core Release Details

On July 15 2026, Thinking Machines Lab—a San Francisco AI startup founded in February 2025 by former OpenAI chief technology officer Mira Murati—made its inaugural model, Inkling, publicly available. Inkling is an open-weight, mixture-of-experts (MoE) transformer comprising 975 billion total parameters with ?41 billion active per token. The model was trained from scratch on 45 trillion multimodal tokens (text, images, audio, video) and supports a one-million-token context window. Weights are released under the Apache 2.0 license and can be downloaded directly or accessed via the company’s Tinker fine-tuning platform.

Background & Context

Thinking Machines Lab raised a $2 billion seed round at a $12 billion valuation from investors including Nvidia, AMD, Cisco, Andreessen Horowitz and Jane Street. The firm positions Inkling as a counterpoint to the “one-size-fits-all” closed models from OpenAI, Anthropic and Google, arguing that enterprises lose domain-specific expertise when they rely on centrally hosted APIs. The launch follows a broader industry shift, with Microsoft CEO Satya Nadella and Hugging Face CEO Clem Delangue warning that proprietary AI can double-pay enterprises—once in subscription fees and again through the inadvertent transfer of proprietary knowledge.

Technical Specs & Benchmarks

Inkling’s MoE design mirrors DeepSeek-V3, routing each token through six experts. The model can run at native 16-bit precision on roughly eight Nvidia B300 accelerators (or sixteen H200s) and also offers an NVFP4-quantized checkpoint that halves the hardware requirement. Reported benchmark scores include 77.6 % on SWE-bench Verified, 97.1 % on AIME 2026, 87.2 % on GPQA Diamond, and 73.5 % on MMMU Pro. Compared with Nvidia’s Nemotron 3 Ultra, Inkling allegedly achieves comparable coding performance on Terminal Bench 2.1 while using roughly one-third the generated tokens.

Why It Matters for Enterprises

Inkling is marketed as a “base” rather than a finished product. Its open-weight nature lets firms download, modify, and fine-tune the model on proprietary data via Tinker, potentially reducing reliance on costly API subscriptions and preserving confidential business logic. The model also features a “thinking-effort” dial that lets users trade depth for speed, and it provides calibrated responses that flag uncertainty instead of guessing—attributes aimed at high-stakes enterprise applications.

Official Statements & Responses

Thinking Machines emphasized that Inkling is “not the strongest model available today, closed or open,” but framed its strength as flexibility and cost-efficiency. The company highlighted a partnership with Nvidia that supplied GB300 NVL72 systems and a planned gigawatt of Vera Rubin computing capacity for future training. Revenue, the firm said, will derive from Tinker services—training, fine-tuning, and hosting—rather than metered model access.

Criticism & Opposition

Analysts note that Inkling trails leading Chinese open models such as Zhipu’s GLM 5.2 and Moonshot’s Kimi K2.6 on benchmarks like Humanity’s Last Exam and SWE-Bench. Critics also point out that the model’s post-training data incorporated outputs from those very Chinese models, raising concerns about distillation practices and the originality of Inkling’s capabilities.

Conflicting Reports & Gaps

Sources differ on Inkling’s relative performance: some claim it “matches” Nemotron 3 Ultra on specific coding tasks, while others state it “lags” behind leading Chinese models overall. The total cost of training Inkling has not been disclosed, and a reported $50 billion fundraising round was said to have stalled, leaving the company’s long-term financing unclear.

Verbatim Quotes

  • “it is not the strongest model available today, closed or open.” — Thinking Machines Lab
  • “Our interactions with each other are very rich,” she told Bloomberg last month. “There’s a lot of information in our interactions – when we’re silent, when we’re thinking, when we’re interrupting one another. Interaction models are able to capture all of these nuances.” — Mira Murati, CEO, Thinking Machines Lab
  • “the chain of thought became more concise over time, dropping grammatical overhead while remaining comprehensible and leaving the final response unaffected.” — Thinking Machines Lab blog
  • “cost against performance,” — Thinking Machines Lab statement
  • “On one benchmark, the company says, Inkling uses a third as many tokens as Nvidia’s Nemotron 3 Ultra — its latest generation open-weight model — to hit the same coding performance.” — Thinking Machines Lab technical release

What’s Next

Thinking Machines previewed Inkling-Small, a 276-billion-parameter MoE variant optimized for lower latency and cost, with weights slated for release after testing. The company also plans to expand API availability to third-party services such as TogetherAI, Fireworks, Modal, Databricks and Baseten, and to continue developing self-contained post-training pipelines for future models.