organisation · org/thinking-machines-lab

Thinking Machines Lab

Also called Thinking Machines, Inkling

The debut model from the lab founded by OpenAI's former chief technology officer is built on a Chinese architecture, trained partly on another Chinese lab's output, and given away under Apache. thinkingmachines/inkling, listed 17 July 2026, is 975B total, 41B active — a 66-layer decoder-only transformer routing each token to 6 of 256 experts, plus 2 shared experts active on every tokensource, accessed 2026-08-28, published under Apache License 2.0source, accessed 2026-08-28. Wikipedia's account of the release records that it incorporated architecture from the Chinese model DeepSeek-V3 and synthetic data from Moonshot AI's Kimi K2.5source, accessed 2026-08-28. So a DeepSeek architecture and Moonshot AI training data sit inside the first model shipped by an American company that raised US$2 billion at a US$12 billion valuation (July 2025), led by Andreessen Horowitz with Nvidia, AMD, Cisco and Jane Street participatingsource, accessed 2026-08-28.

Everything the lab has listed arrived inside thirteen days. Five rows — thinkingmachines/inkling, its batch and free variants, thinkingmachines/inkling-small and that model's free variant — carry listing dates between 17 and 30 July 2026, and nothing has followed. Inkling Small is a 276-billion-parameter distillation with 12 billion active. Every model this lab has published is downloadable, and every one is also available at no charge through the router: two models, five rows, no closed weights and no paywall that cannot be walked around.

The three ways to call the same model do not rank the way the names suggest. thinkingmachines/inkling:free serves the full window at no cost. The standard row serves the same 1048576tokensopenrouter-models, last checked 2026-09-14. The batch row halves it to 524288tokensopenrouter-models, last checked 2026-09-14 and heads at $1.00per million tokensopenrouter-models, last checked 2026-09-14 for input against $1.00per million tokensopenrouter-models, last checked 2026-09-14 on the standard row — a batch listing above the row it batches, where the convention is a discount. Neither figure is necessarily this lab's, though: each is the top listed provider's rate for its row, and two rows are not obliged to be headed by the same provider. Batch pricing is a convention and not a guarantee, but an inversion visible only across two separately ranked rows is a fact about who was listed first on each, not a decision the lab made.

Facts

founded
February 2025, by Mira Murati, formerly OpenAI's chief technology officersource, accessed 2026-08-28
headquarters
2300 Harrison Street, Mission District, San Franciscosource, accessed 2026-08-28
funding
US$2 billion at a US$12 billion valuation (July 2025), led by Andreessen Horowitz with Nvidia, AMD, Cisco and Jane Street participatingsource, accessed 2026-08-28
flagship license
Apache License 2.0source, accessed 2026-08-28
flagship parameters
975B total, 41B active — a 66-layer decoder-only transformer routing each token to 6 of 256 experts, plus 2 shared experts active on every tokensource, accessed 2026-08-28
flagship lineage
incorporated architecture from the Chinese model DeepSeek-V3 and synthetic data from Moonshot AI's Kimi K2.5source, accessed 2026-08-28

Timeline

  1. Inkling Small listed — a 276B distillation, also with a free rowsource
  2. Inkling listed, with free and batch rows alongside the standard onesource
  3. Tinker released — an API for fine-tuning open-weight models on the company's infrastructuresource