// model_catalog

29 models that run on Little Brother

Every model is quantized for Apple Silicon, tested by us, and runs fully on-device. Grouped by where it runs, then by family. Memory figures are on-device peaks — weights are download size estimates.

iPhone & iPad

18 models

Models that fit within iOS's ~5 GB per-app memory budget. “Tight” models still run, but sit close to the limit on some devices.

Qwen 6

Qwen2.5 0.5B Instruct (4-bit)

Ultra-light chat and quick PDF Q&A; long answers cut off

RAM peak
1.5–2.3 GB
Weights
0.27 GB
Context
32,768 tok
Min Mac RAM
8 GB
Alibaba (Qwen) Apache 2.0

Qwen3 0.6B (4-bit)

Fast reasoning, PDF Q&A, and doc lookup at a tiny size

RAM peak
1.5–2.3 GB
Weights
0.33 GB
Context
40,960 tok
Min Mac RAM
8 GB
Reasoning
Alibaba (Qwen) Apache 2.0

Qwen2.5 1.5B Instruct (4-bit)

Solid everyday chat and PDF work; doc lookup can miss specifics

RAM peak
2.0–2.8 GB
Weights
0.82 GB
Context
32,768 tok
Min Mac RAM
8 GB
Alibaba (Qwen) Apache 2.0

Qwen3 1.7B (4-bit)

Reasoning plus accurate doc lookup in a light model

RAM peak
2.1–2.9 GB
Weights
0.92 GB
Context
40,960 tok
Min Mac RAM
8 GB
Reasoning
Alibaba (Qwen) Apache 2.0

Qwen3.5 2B (4-bit)

Current-gen Qwen chat with reasoning; forgot cross-chat facts in tests

RAM peak
2.8–3.6 GB
Weights
1.63 GB
Context
262,144 tok
Min Mac RAM
8 GB
Reasoning
Alibaba (Qwen) Apache 2.0

Qwen3 4B Instruct 2507 (4-bit)

Top all-rounder — reasoning, web search, PDF tools; passed every test

RAM peak
3.3–4.1 GB
Weights
2.12 GB
Context
262,144 tok
Min Mac RAM
8 GB
Reasoning Tools Web search Documents
Alibaba (Qwen) Apache 2.0

LFM (Liquid AI) 1

LFM2.5 1.2B Instruct (4-bit)

Fast, dependable chat and PDF/doc Q&A (Liquid AI)

RAM peak
1.8–2.6 GB
Weights
0.62 GB
Context
128,000 tok
Min Mac RAM
8 GB

Llama 2

Llama 3.2 1B Instruct (4-bit)

Light general chat and PDF summaries; weak doc lookup in tests

RAM peak
1.9–2.7 GB
Weights
0.66 GB
Context
131,072 tok
Min Mac RAM
8 GB

Llama 3.2 3B Instruct (4-bit)

Reliable general chat and doc Q&A; PDF summaries occasionally thin

RAM peak
2.9–3.7 GB
Weights
1.7 GB
Context
131,072 tok
Min Mac RAM
8 GB
Tools

Gemma 3

Gemma 3 1B Instruct QAT (4-bit)

Quality chat at a tiny footprint (QAT); mixed doc lookup in tests

RAM peak
1.9–2.7 GB
Weights
0.72 GB
Context
32,768 tok
Min Mac RAM
8 GB
Google DeepMind Gemma Terms of Use

Gemma 3n E2B Instruct (4-bit)

Efficient chat with strong doc lookup; long answers sometimes cut off

RAM peak
3.6–4.4 GB
Weights
2.37 GB
Context
32,768 tok
Min Mac RAM
8 GB
Google DeepMind Gemma Terms of Use

Gemma 3 4B Instruct QAT (4-bit)

Higher-quality Gemma chat and PDF summaries (QAT-tuned)

RAM peak
4.0–4.8 GB
Weights
2.83 GB
Context
131,072 tok
Min Mac RAM
8 GB
iOS: tight
Google DeepMind Gemma Terms of Use

SmolLM 1

SmolLM3 3B (4-bit)

Reasoning, PDFs, and doc lookup — passed every on-device test

RAM peak
2.9–3.7 GB
Weights
1.67 GB
Context
65,536 tok
Min Mac RAM
8 GB
Reasoning Tools Documents
Hugging Face Apache 2.0

Phi 2

Phi-3.5 Mini Instruct (4-bit)

Default model — balanced chat and PDF Q&A; long answers cut off in tests

RAM peak
3.2–4.0 GB
Weights
2.0 GB
Context
131,072 tok
Min Mac RAM
8 GB
Microsoft MIT

Phi-4 Mini Instruct (4-bit)

Short chat and simple doc questions; unstable on iOS and long PDF tasks

RAM peak
3.3–4.1 GB
Weights
2.16 GB
Context
131,072 tok
Min Mac RAM
8 GB
iOS: tight Tools Documents
Microsoft MIT

Nemotron 2

Nemotron 3 Nano 4B (4-bit)

Reasoning, web search, and doc Q&A — near-perfect in tests

RAM peak
3.2–4.0 GB
Weights
2.24 GB
Context
262,144 tok
Min Mac RAM
8 GB
Reasoning Tools Web search Documents

Llama Nemotron Nano 4B v1.1 (4-bit)

Reasoning and math; skipped web search and failed PDF summaries in tests

RAM peak
3.5–4.3 GB
Weights
2.54 GB
Context
131,072 tok
Min Mac RAM
8 GB
iOS: tight Reasoning Tools Web search Documents

Ministral 1

Ministral 3 3B Instruct (4-bit)

Compact instruction-following chat and doc Q&A

RAM peak
3.8–4.6 GB
Weights
2.59 GB
Context
262,144 tok
Min Mac RAM
8 GB
iOS: tight Tools Documents
Mistral AI Apache 2.0

Mac

29 models

Every catalog model runs on a Mac with enough memory — the card shows the minimum RAM each one needs.

Qwen 9

Qwen2.5 0.5B Instruct (4-bit)

Ultra-light chat and quick PDF Q&A; long answers cut off

RAM peak
1.5–2.3 GB
Weights
0.27 GB
Context
32,768 tok
Min Mac RAM
8 GB
Alibaba (Qwen) Apache 2.0

Qwen3 0.6B (4-bit)

Fast reasoning, PDF Q&A, and doc lookup at a tiny size

RAM peak
1.5–2.3 GB
Weights
0.33 GB
Context
40,960 tok
Min Mac RAM
8 GB
Reasoning
Alibaba (Qwen) Apache 2.0

Qwen2.5 1.5B Instruct (4-bit)

Solid everyday chat and PDF work; doc lookup can miss specifics

RAM peak
2.0–2.8 GB
Weights
0.82 GB
Context
32,768 tok
Min Mac RAM
8 GB
Alibaba (Qwen) Apache 2.0

Qwen3 1.7B (4-bit)

Reasoning plus accurate doc lookup in a light model

RAM peak
2.1–2.9 GB
Weights
0.92 GB
Context
40,960 tok
Min Mac RAM
8 GB
Reasoning
Alibaba (Qwen) Apache 2.0

Qwen3.5 2B (4-bit)

Current-gen Qwen chat with reasoning; forgot cross-chat facts in tests

RAM peak
2.8–3.6 GB
Weights
1.63 GB
Context
262,144 tok
Min Mac RAM
8 GB
Reasoning
Alibaba (Qwen) Apache 2.0

Qwen3 4B Instruct 2507 (4-bit)

Top all-rounder — reasoning, web search, PDF tools; passed every test

RAM peak
3.3–4.1 GB
Weights
2.12 GB
Context
262,144 tok
Min Mac RAM
8 GB
Reasoning Tools Web search Documents
Alibaba (Qwen) Apache 2.0

Qwen3.5 4B (4-bit)

Strong reasoning, web search, and PDF tools — Mac only (OOM on iOS)

RAM peak
4.1–4.9 GB
Weights
2.85 GB
Context
262,144 tok
Min Mac RAM
8 GB
Mac only Reasoning Tools Web search Documents
Alibaba (Qwen) Apache 2.0

Qwen2.5 7B Instruct (4-bit)

High-accuracy PDF tools, web search, and doc lookup (needs 16 GB Mac)

RAM peak
5.2–6.0 GB
Weights
4.0 GB
Context
32,768 tok
Min Mac RAM
16 GB
Mac only Tools Web search Documents
Alibaba (Qwen) Apache 2.0

Qwen3 8B (4-bit)

Strongest reasoning in class — passed every on-device test

RAM peak
5.5–6.3 GB
Weights
4.31 GB
Context
40,960 tok
Min Mac RAM
16 GB
Mac only Reasoning Tools Web search Documents
Alibaba (Qwen) Apache 2.0

LFM (Liquid AI) 2

LFM2.5 1.2B Instruct (4-bit)

Fast, dependable chat and PDF/doc Q&A (Liquid AI)

RAM peak
1.8–2.6 GB
Weights
0.62 GB
Context
128,000 tok
Min Mac RAM
8 GB

LFM2 8B A1B (3-bit)

Fast mixture-of-experts chat and PDFs on Mac; doc lookup missed specifics

RAM peak
5.1–5.9 GB
Weights
3.89 GB
Context
128,000 tok
Min Mac RAM
16 GB
Mac only Tools Documents

Llama 3

Llama 3.2 1B Instruct (4-bit)

Light general chat and PDF summaries; weak doc lookup in tests

RAM peak
1.9–2.7 GB
Weights
0.66 GB
Context
131,072 tok
Min Mac RAM
8 GB

Llama 3.2 3B Instruct (4-bit)

Reliable general chat and doc Q&A; PDF summaries occasionally thin

RAM peak
2.9–3.7 GB
Weights
1.7 GB
Context
131,072 tok
Min Mac RAM
8 GB
Tools

Meta Llama 3.1 8B Instruct (4-bit)

Dependable general chat and doc Q&A on Mac

RAM peak
5.4–6.2 GB
Weights
4.22 GB
Context
131,072 tok
Min Mac RAM
16 GB
Mac only Tools Documents

Gemma 3

Gemma 3 1B Instruct QAT (4-bit)

Quality chat at a tiny footprint (QAT); mixed doc lookup in tests

RAM peak
1.9–2.7 GB
Weights
0.72 GB
Context
32,768 tok
Min Mac RAM
8 GB
Google DeepMind Gemma Terms of Use

Gemma 3n E2B Instruct (4-bit)

Efficient chat with strong doc lookup; long answers sometimes cut off

RAM peak
3.6–4.4 GB
Weights
2.37 GB
Context
32,768 tok
Min Mac RAM
8 GB
Google DeepMind Gemma Terms of Use

Gemma 3 4B Instruct QAT (4-bit)

Higher-quality Gemma chat and PDF summaries (QAT-tuned)

RAM peak
4.0–4.8 GB
Weights
2.83 GB
Context
131,072 tok
Min Mac RAM
8 GB
iOS: tight
Google DeepMind Gemma Terms of Use

SmolLM 1

SmolLM3 3B (4-bit)

Reasoning, PDFs, and doc lookup — passed every on-device test

RAM peak
2.9–3.7 GB
Weights
1.67 GB
Context
65,536 tok
Min Mac RAM
8 GB
Reasoning Tools Documents
Hugging Face Apache 2.0

Phi 3

Phi-3.5 Mini Instruct (4-bit)

Default model — balanced chat and PDF Q&A; long answers cut off in tests

RAM peak
3.2–4.0 GB
Weights
2.0 GB
Context
131,072 tok
Min Mac RAM
8 GB
Microsoft MIT

Phi-4 Mini Instruct (4-bit)

Short chat and simple doc questions; unstable on iOS and long PDF tasks

RAM peak
3.3–4.1 GB
Weights
2.16 GB
Context
131,072 tok
Min Mac RAM
8 GB
iOS: tight Tools Documents
Microsoft MIT

Phi-4 (3-bit)

High-accuracy reasoning and math on Mac

RAM peak
7.0–7.8 GB
Weights
5.98 GB
Context
16,384 tok
Min Mac RAM
16 GB
Mac only
Microsoft MIT

Nemotron 7

Nemotron 3 Nano 4B (4-bit)

Reasoning, web search, and doc Q&A — near-perfect in tests

RAM peak
3.2–4.0 GB
Weights
2.24 GB
Context
262,144 tok
Min Mac RAM
8 GB
Reasoning Tools Web search Documents

Llama Nemotron Nano 4B v1.1 (4-bit)

Reasoning and math; skipped web search and failed PDF summaries in tests

RAM peak
3.5–4.3 GB
Weights
2.54 GB
Context
131,072 tok
Min Mac RAM
8 GB
iOS: tight Reasoning Tools Web search Documents

Llama Nemotron 8B UltraLong 1M (4-bit)

Long-context reasoning, web search, and doc Q&A on Mac

RAM peak
5.5–6.3 GB
Weights
4.52 GB
Context
1,073,152 tok
Min Mac RAM
16 GB
Mac only Reasoning Tools Web search Documents

Nemotron Nano 9B v2 (4-bit)

Reasoning, web search, and docs — passed every on-device test

RAM peak
6.0–6.8 GB
Weights
5.0 GB
Context
131,072 tok
Min Mac RAM
16 GB
Mac only Reasoning Tools Web search Documents

Nemotron 3 Nano 30B A3B (4-bit)

MoE flagship — ~3B active/token, strong reasoning on Mac

RAM peak
18.8–19.6 GB
Weights
17.8 GB
Context
262,144 tok
Min Mac RAM
32 GB
Mac only Reasoning Tools Web search Documents

Llama Nemotron 70B Instruct (4-bit)

Top Llama Nemotron instruct quality on high-RAM Macs

RAM peak
40.7–41.5 GB
Weights
39.7 GB
Context
131,072 tok
Min Mac RAM
64 GB
Mac only Reasoning Tools Web search Documents

Nemotron 3 Super 120B A12B (4-bit)

Largest Nemotron 3 — ~12B active/token for M3 Ultra class Macs

RAM peak
121.7–122.5 GB
Weights
120.7 GB
Context
262,144 tok
Min Mac RAM
128 GB
Mac only Reasoning Tools Web search Documents

Ministral 1

Ministral 3 3B Instruct (4-bit)

Compact instruction-following chat and doc Q&A

RAM peak
3.8–4.6 GB
Weights
2.59 GB
Context
262,144 tok
Min Mac RAM
8 GB
iOS: tight Tools Documents
Mistral AI Apache 2.0

// sovereign_vault_protocol

Enter the Sovereign Vault

Little Brother is coming soon. No account required. Nothing leaves your device.

Coming Soon