$ inference --device=local --cloud=none
Little Brother runs open LLMs entirely on your Mac and iPhone. You own your data, your software, and your hardware. Private by physics, not by policy.
// 100% on-device
Little Brother runs curated MLX models directly on Apple Silicon. Your prompts, your documents, and every token generated stay on your device — there is no server to trust because there is no server at all.
// 29 curated models
Qwen, Llama, Phi, Gemma, Nemotron, SmolLM3 and more — each with on-device memory estimates so you know what runs well on your Mac, iPad, or iPhone.
Browse all models →
// local RAG
Attach PDFs and ask questions. Little Brother indexes them locally with on-device embeddings and picks the right retrieval strategy — whole document, chunked, or summarized — without your files ever leaving the machine.
// live web search
When a question needs fresh information, models can call a live web search tool and answer with clickable citations — the search happens on demand, and only the query leaves your device.
How web search works →
// cross-model memory
Switch from a 1B model to an 8B model mid-conversation and keep the thread. Conversation memory is shared across every model in the catalog, so your context travels with you.
// sovereign_vault_protocol
Little Brother is coming soon. No account required. Nothing leaves your device.