In plain terms
With a closed model you send your request to the provider's servers and get an answer back; the model itself never leaves the provider. With an open-weight model you receive the model file, the set of numbers that is the model, and run it wherever you like: on a laptop, in your data centre or in your own cloud account.
Why it matters
It gives control. Data stays in your environment, the model cannot be changed or withdrawn from under you, it can be fine-tuned in depth, and at steady high volume it can cost less. It also gives you the work: hardware, operations, security, updates and the safeguards a provider would otherwise supply. The most capable models at any moment are usually closed; open-weight models follow some months behind and are good enough for a large share of tasks. Read the licence: some forbid commercial use or attach conditions.
Example
A hospital wants to summarise patient records and is not permitted to send them to an outside service. It downloads an open-weight model, runs it on two GPU servers in its own data centre and fine-tunes it on anonymised discharge letters. No patient data leaves the building. The price is an engineer's time to operate it, and a model less capable than the closed frontier.
Most often confused with
Open-weight Model vs. Open source
The two are routinely mixed up. Open weights let you use and adapt the finished model; they do not let you see what it was trained on, or rebuild it. Open source in the full sense would require the training code and enough information about the data to do that, under a licence without restrictions on use. Many models advertised as open source are open-weight only.
Under the hood
What is released: the weight files, the architecture and tokenizer, usually code for running the model, and a model card. What is usually withheld: the training data and the full training pipeline. Licences range from permissive ones (Apache 2.0, MIT) to custom terms that restrict use, scale or field. The Open Source Initiative published a definition of open-source AI in 2024 that most open-weight releases do not meet. Well-known families include Llama, Mistral, Gemma, Qwen, DeepSeek and Phi. They are run with local tools such as Ollama and llama.cpp, with serving frameworks such as vLLM, or hosted by cloud providers. Quantised versions reduce memory needs at some cost in quality. Two further points: once released, weights cannot be recalled and safeguards can be removed by fine-tuning, which is the centre of the policy debate; and model files should be downloaded only from trusted sources, as with any software.