Qwen2.5 0.5B Instruct
Any deviceAlibaba Qwen · GGUF by bartowski
Tiny and fast. Good for quick rewrites and classification on any device.
- Download
- 398 MB
- Parameters
- 0.5B
- Quantization
- Q4_K_M
- Context
- 32K tokens
Licence: Apache-2.0
All 10 models.
Alibaba Qwen · GGUF by bartowski
Tiny and fast. Good for quick rewrites and classification on any device.
Licence: Apache-2.0
Meta · GGUF by bartowski
Fast general assistant with a long context window. Comfortable on phones.
Licence: Llama 3.2 Community License
Hugging Face · GGUF by bartowski
Small model trained for instruction following and summarisation.
Licence: Apache-2.0
Google · GGUF by bartowski
Strong writing quality for its size. Best on devices with 4 GB or more free.
Licence: Gemma Terms of Use
Alibaba Qwen · GGUF by bartowski
Balanced general model. The recommended open-source choice on a phone.
Licence: Qwen Research License
Alibaba Qwen · GGUF by bartowski
Tuned for code and shell questions. Pairs well with the Shell helper persona.
Licence: Qwen Research License
Meta · GGUF by bartowski
Long-context assistant. Needs about 3 GB of free memory.
Licence: Llama 3.2 Community License
Microsoft · GGUF by bartowski
Reasoning-leaning small model with a permissive licence.
Licence: MIT
Mistral AI · GGUF by bartowski
Mac-class model. Needs roughly 6 GB of free memory at Q4.
Licence: Apache-2.0
Alibaba Qwen · GGUF by bartowski
The strongest general model in the catalog. Best on a Mac with 16 GB or more.
Licence: Apache-2.0
Comfortable on any iPhone, iPad or Mac that runs Sotto.
Wants about 3 GB of free memory. Fine on a recent iPhone or any Mac.
Wants roughly 6 GB of free memory. Really a Mac model.
Memory is the real limit, not storage. The figure Sotto checks before loading a model is the weights plus the KV cache — so if a model will not load, either pick a smaller quantization or lower the context length in Settings › Models.
Every download goes to huggingface.co over HTTPS, and that host is pinned in the app: a catalog entry pointing anywhere else is refused at decode time. Sotto sends no identifier and no account with the request. Model weights are third-party works under the licences listed above.