The capability manifest lists each model's vision, PDF, OCR, and tool-use flags. Routing picks the model that both handles the input type and fits your device's memory budget — quantization is preferred when the tier budget is tight.
On-device intelligence.
Chat, code, and agents — all running locally, nothing leaves your device.
Explore the complete local-first workspace. Toggle between modules to see how they run on-device, and switch devices to preview native rendering on iPhone, iPad, and Mac.
Talk to instruct-tuned LLMs with the polish of a modern chat app. Streaming thoughts, structured tool calls, and local history. Zero cloud required.
Cirrus plugs straight into your GitHub repositories. Describe a task, inspect an inline diff card with hunk headers, and open a PR with a single tap.
Ashe runs background tasks ("Hands") locally. Watch tool calls, checks, and models trigger in real time. Ashe always asks for permission before modifying files.
Mist merges vector search and web results locally. Search public databases, extract text, and compile cited answers without leaking your query.
Prompt-to-image running directly on Apple Silicon. Mirage leverages palettized diffusion models to render crisp images without consuming cloud tokens.
Dictate, transcribe recordings, and run voice-first workflows locally. Overture uses Whisper in GGML format for lightning-fast transcription.
Point your camera, capture, and extract structured data. Stratus routes images through multimodal chat models and fast OCR kernels.
Translate documents, notes, or messages side-by-side. Zephyr utilizes chat LLMs with translation templates for context-aware, idiomatic translations.
The capability manifest lists each model's vision, PDF, OCR, and tool-use flags. Routing picks the model that both handles the input type and fits your device's memory budget — quantization is preferred when the tier budget is tight.
ModelRouter.classify() takes a repoId, filename, tags, libraryName, and pipelineTag and returns a Classification with .module and .confidence.
A quick tour of MLX vs GGUF on Apple Silicon, how we rank quants, and why a 3B-Q4 often beats a 7B in day-to-day chat.
Because Nimbus8 runs entirely on-device, memory is the ultimate budget. Slide below to select your device RAM and see which models run optimally. We account for both raw weights and active KV cache overhead to prevent Out-Of-Memory (OOM) crashes.
Highly optimized ultra-lightweight 2026 chat model.
Google's mid-2026 lightweight model with deep reasoning.
Excellent reasoning and multilingual capabilities for compact devices.
Premium balanced language assistant, fits well with active KV cache overhead.
State-of-the-art mid-2026 chat model with visual capability.
Robust general assistant, runs comfortably on 8 GB devices.
Top-tier coding model. Demands 12 GB+ RAM to accommodate 1.8 GB KV cache.
Distilled reasoning model with offline chain-of-thought capabilities.
Flagship local model. Requires 16 GB+ RAM for full 8k context window.
Download Nimbus8 today for iPhone, iPad, and Mac. Gale is free forever, and Nimbus8 Pro unlocks the full local studio with a one-time Apple In-App Purchase.
Available now for iOS 26+, iPadOS 26+, and macOS 26+ on Apple silicon Macs.