Good for a fast, on-device assistant for short replies and simple instructions. Too small for long documents or complex reasoning, and it will make things up under pressure.
Intended for short, low-latency assistant tasks on constrained hardware such as phones or edge devices. Not intended for long-context work or tasks requiring strong reasoning.
Base weights and instruction tuning data are Qwen's own general web, code, and multilingual corpora, redistributed here with quantised GGUF builds added for local inference.
Apache-2.0. Free for commercial and research use with attribution retained in redistributions.
from 3 reviews of 7ae5576
ollama run Qwen/Qwen2.5-0.5B-Instruct| Build | File | VRAM |
|---|---|---|
| Q4_K_M | 397 MB | 0.9 GB |
| Q8_0 | 531 MB | 1.1 GB |
Set a hardware profile in the browse sidebar to see which builds fit.
No hosted providers listed yet.