How it works
The weights are the billions of numbers a model learns in training. When they are published, usually on Hugging Face, you can run the model on your own hardware, host it with any provider, inspect it and fine-tune it. Families from Meta (Llama), Google (Gemma), Mistral, Alibaba (Qwen), DeepSeek, Microsoft (Phi) and OpenAI (gpt-oss) come in sizes from a few billion parameters, which run on a laptop, to hundreds of billions or more, which need a cluster of GPUs.
Open weight is not the same as open source: the training data and code are rarely released, and licences range from permissive (Apache 2.0, MIT) to custom ones with conditions for very large companies or on certain uses. Quantisation stores each weight in 4 or 8 bits so models fit in less memory, at some cost in quality. Tools such as Ollama, llama.cpp and vLLM run them, and hosts such as OpenRouter serve them per token.
Related terms
More in AI and LLMs
Run it yourself