TL;DR
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Thinking Machines Lab released the full weights for its first foundation model, Inkling, on July 15 under the Apache 2.0 license before offering a closed API. The release gives organizations more control, but high hardware demands, unpublished training data and a reported use policy limit what ownership means in practice.
Thinking Machines Lab, founded by former OpenAI technology chief Mira Murati, released the full weights for its first foundation model, Inkling, on July 15 under Apache 2.0. Publishing the weights before a closed API gives customers the ability to download, modify and operate the model themselves, making the release a test of whether ownership can compete with rented access to leading AI systems.
Inkling is a mixture-of-experts model with 975 billion total parameters and 41 billion active parameters per operation. Thinking Machines says it has a one-million-token context window and was pretrained on 45 trillion tokens spanning text, images, audio and video. It accepts text, image and audio inputs and produces text.
The lab published BF16 and NVFP4 checkpoints through Hugging Face, with initial support in Transformers, vLLM, SGLang and llama.cpp. The Apache license permits modification and commercial use, but the release does not include the training dataset or full training pipeline, so it is more accurately described as open-weight than fully open-source.
Thinking Machines also said Inkling is not the strongest available model, whether compared with open or closed systems. Its vendor-reported results include 97.1% on AIME 2026 and 87.2% on GPQA Diamond, while results on several software-engineering and agentic tasks trail named rivals. Those figures await independent replication, and some reportedly came from a prerelease checkpoint.
Model Ownership Becomes the Product
Model weights are the numerical parameters learned during training. Possession of them lets an organization run the model on its own infrastructure, alter it through further training and avoid dependence on a provider that can change prices, access rules or product availability. For governments and regulated businesses, that creates a possible sovereignty and continuity benefit.
The order of release offers a clue to the lab’s commercial strategy. By putting weights ahead of API access, Thinking Machines appears to be treating deployment control and fine-tuning as core products rather than using open weights as a delayed version of a proprietary service. Its adjustable 0.2-to-0.99 reasoning-effort setting also targets operating cost and latency, though the claimed efficiency gains have not been independently verified.
As an affiliate, we earn on qualifying purchases.
Open Models Shift Toward Deployment
Open-weight releases have often arrived after closed models or with restrictions that narrowed commercial use. Inkling reverses that sequence and arrives with deployment software available on release day. Thorsten Meyer AI described the launch as a strategic open release, while stressing that access to weights does not disclose how the training corpus was assembled.
The model also enters a field where Chinese open-weight systems, including GLM-5.2 and Kimi K2.6, have posted stronger results on some reasoning, agentic and multimodal evaluations. The source material reports that Inkling’s post-training used synthetic data from Kimi K2.5, underscoring the technical overlap among competing model developers.
“Inkling is not the strongest model available today, closed or open.”
— Thinking Machines Lab announcement
hardware for running large AI models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Licensing and Performance Need Verification
Several limits remain unresolved. The source material reports a separate Model Acceptable Use Policy covering the original parameters and modified versions, including restrictions involving surveillance, deception and automated decisions affecting rights. That policy was not independently verified in the supplied reporting, and its relationship with Apache 2.0 may affect commercial adoption.
The practical reach of the release is also uncertain. The supplied estimates put BF16 operation at roughly two terabytes of aggregate VRAM and NVFP4 at 600 gigabytes or more, keeping the flagship beyond ordinary workstations. Claims about benchmark leadership, token savings and compressed reasoning traces remain vendor-reported findings until outside evaluators reproduce them.
As an affiliate, we earn on qualifying purchases.
Independent Tests Will Define Inkling
Researchers and prospective users can now test Inkling’s published checkpoints against GLM-5.2, Kimi K2.6 and proprietary services on their own workloads. Early scrutiny is likely to focus on benchmark reproducibility, inference costs, licensing terms and whether the reasoning-effort control produces reliable savings at scale.
Thinking Machines also plans to publish weights for Inkling-Small, a 276-billion-parameter model with 12 billion active parameters, after testing is complete. That smaller release may prove more accessible, but the lab has not provided a confirmed publication date.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does releasing AI model weights allow users to do?
Users can download and operate the model on infrastructure they control, modify it and use it commercially under the stated license. That reduces dependence on a hosted API provider, subject to hardware and policy limits.
Is Inkling fully open-source?
No. The weights are available under Apache 2.0, but the training data and complete training pipeline were not published. Inkling is best described as an open-weight model.
Is Inkling the most capable open model?
Thinking Machines does not claim that it is. Inkling reports strong results on selected evaluations, but trails competitors on several coding, agentic and multimodal tests. The published numbers still require independent replication.
Can Inkling run on a personal workstation?
The flagship is unlikely to run effectively on typical consumer hardware. Even its compressed NVFP4 checkpoint reportedly needs at least 600 gigabytes of VRAM. The planned Inkling-Small release may offer a more practical option.
Why is the release order meaningful?
Publishing complete weights before a closed API makes user control part of the initial product rather than a later concession. It signals that Thinking Machines expects self-hosting and customization to influence purchasing decisions alongside raw benchmark scores.
Source: Thorsten Meyer AI
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
