Menu

Model catalog

Ten models, sized and licensed up front

Sotto ships a curated list rather than a live search, so every entry has a verified file size and a stated licence before you spend a gigabyte on it. You can also import any GGUF file you already have by dropping it on the window.

Catalog updated 2026-09-02 · 10 entries

Runs on

All 10 models.

  • Qwen2.5 0.5B Instruct

    Any device

    Alibaba Qwen · GGUF by bartowski

    Tiny and fast. Good for quick rewrites and classification on any device.

    Download
    398 MB
    Parameters
    0.5B
    Quantization
    Q4_K_M
    Context
    32K tokens

    Licence: Apache-2.0

  • Llama 3.2 1B Instruct

    Any device

    Meta · GGUF by bartowski

    Fast general assistant with a long context window. Comfortable on phones.

    Download
    912 MB
    Parameters
    1B
    Quantization
    Q5_K_M
    Context
    128K tokens

    Licence: Llama 3.2 Community License

  • SmolLM2 1.7B Instruct

    Any device

    Hugging Face · GGUF by bartowski

    Small model trained for instruction following and summarisation.

    Download
    1.1 GB
    Parameters
    1.7B
    Quantization
    Q4_K_M
    Context
    8K tokens

    Licence: Apache-2.0

  • Gemma 2 2B it

    Any device

    Google · GGUF by bartowski

    Strong writing quality for its size. Best on devices with 4 GB or more free.

    Download
    1.7 GB
    Parameters
    2B
    Quantization
    Q4_K_M
    Context
    8K tokens

    Licence: Gemma Terms of Use

  • Qwen2.5 3B Instruct

    Roomy phone or Mac

    Alibaba Qwen · GGUF by bartowski

    Balanced general model. The recommended open-source choice on a phone.

    Download
    1.9 GB
    Parameters
    3B
    Quantization
    Q4_K_M
    Context
    32K tokens

    Licence: Qwen Research License

  • Qwen2.5 Coder 3B Instruct

    Roomy phone or Mac

    Alibaba Qwen · GGUF by bartowski

    Tuned for code and shell questions. Pairs well with the Shell helper persona.

    Download
    1.9 GB
    Parameters
    3B
    Quantization
    Q4_K_M
    Context
    32K tokens

    Licence: Qwen Research License

  • Llama 3.2 3B Instruct

    Roomy phone or Mac

    Meta · GGUF by bartowski

    Long-context assistant. Needs about 3 GB of free memory.

    Download
    2.3 GB
    Parameters
    3B
    Quantization
    Q5_K_M
    Context
    128K tokens

    Licence: Llama 3.2 Community License

  • Phi-3.5 mini Instruct

    Roomy phone or Mac

    Microsoft · GGUF by bartowski

    Reasoning-leaning small model with a permissive licence.

    Download
    2.4 GB
    Parameters
    3.8B
    Quantization
    Q4_K_M
    Context
    128K tokens

    Licence: MIT

  • Mistral 7B Instruct v0.3

    Mac

    Mistral AI · GGUF by bartowski

    Mac-class model. Needs roughly 6 GB of free memory at Q4.

    Download
    4.4 GB
    Parameters
    7B
    Quantization
    Q4_K_M
    Context
    32K tokens

    Licence: Apache-2.0

  • Qwen2.5 7B Instruct

    Mac

    Alibaba Qwen · GGUF by bartowski

    The strongest general model in the catalog. Best on a Mac with 16 GB or more.

    Download
    4.7 GB
    Parameters
    7B
    Quantization
    Q4_K_M
    Context
    32K tokens

    Licence: Apache-2.0

Any device

Comfortable on any iPhone, iPad or Mac that runs Sotto.

Roomy phone or Mac

Wants about 3 GB of free memory. Fine on a recent iPhone or any Mac.

Mac

Wants roughly 6 GB of free memory. Really a Mac model.

Memory is the real limit, not storage. The figure Sotto checks before loading a model is the weights plus the KV cache — so if a model will not load, either pick a smaller quantization or lower the context length in Settings › Models.

Every download goes to huggingface.co over HTTPS, and that host is pinned in the app: a catalog entry pointing anywhere else is refused at decode time. Sotto sends no identifier and no account with the request. Model weights are third-party works under the licences listed above.