Coming Soon

Private AI.
No cloud.
Your choice.

CappedAI runs open-source AI privately on your phone or drive — no internet required, no subscription, no server that sees your data.

Free app to start. Expand with knowledge packs that accumulate live data into permanent local knowledge. Or get a drive shipped pre-loaded and ready to go.

🧪 iOS App Coming Soon
💾 Get notified when drives ship
✓ You're on the list.
Airplane mode on. WiFi off. CappedAI still responds.
Three ways in

Start free. Go deeper when you're ready.

The open-source models are free. CappedAI gives you a private, offline way to run them — on whatever hardware you already have.

📱
Free App
Private AI on your phone
Download and go. No credit card, no account, no cloud.

  • Private AI chat, fully offline
  • Curated starter models included
  • Conversation history on device
  • iOS — coming soon
  • Knowledge packs
  • Living knowledge packs
  • Full model library
  • Text-to-image generation
🧪 iOS App — Coming Soon
Hardware
💾
CappedAI Drive
Everything. Pre-loaded. Shipped to you.
High-speed USB-C drive with models downloaded and ready on arrival.

  • Everything in Pro
  • High-speed USB-C drive included
  • Models pre-downloaded, ready on arrival
  • Full HuggingFace model library browser
  • Text-to-image generation, offline
  • iPhone, iPad & Mac
  • Trade-in program — upgrade anytime
  • Packs pre-installed
Reserve a drive
Living knowledge packs

Packs that remember
what they learn.

Knowledge packs are portable expertise modules — each with its own agent, curated knowledge base, and pre-computed embeddings stored on your drive as a .cap file.

Some packs can fetch live data when you're online — a stock pack pulls real-time quotes, a fantasy football pack gets this week's injury report, a weather pack fetches a current forecast.

But here's what makes CappedAI different: that live data doesn't disappear when the session ends. It gets embedded and written permanently into the pack's local knowledge base on your drive. The pack remembers what it learned.

On your next conversation — online or offline — the agent already knows what it fetched last time. The longer you use a pack, the smarter and more current it becomes.

Every other AI tool fetches live data and forgets it.
CappedAI packs remember.

How living knowledge works
💬
Your query
main chat · pack chat · cross-pack
RAG retrieval
🗄️
Local knowledge base
SQLite · bge-small embeddings · on drive
Grows over time
agent decides if live data needed
🌐
CappedAI data services
quotes · scores · forecasts · custom
Online only
embed + write back to drive
💾
Permanent local storage
live data becomes pack knowledge · persists offline
Stays forever
synthesized response
🤖
LLM inference
llama.cpp · Metal GPU · fully local
On device
write-back loop
🌺
Hawaiian Vacation
Agent: Kai
Offline
🏕️
Survivalist
Agent: Ranger
Offline
🤠
Austin TX Summer
Agent: Tex
Offline
📈
Stock Research
Agent: coming soon
Offline Living
🏈
Fantasy Football
Agent: coming soon
Offline Living
More coming
legal · weather · travel & more
Model library

The models are free.
Stop paying to access them.

Qwen 3, Llama 4, Phi-4, Mistral — open-source models now match GPT on most real-world tasks. The weights are free. You've been paying for someone else's server.

Drive owners browse and download from HuggingFace's full model catalog directly inside the app — chat models, image generators, embedding models. Downloaded to the drive, owned permanently. No API key. No billing relationship. No usage cap.

A 128 GB drive holds 15–20 models comfortably. When a better model ships, download it. The old one stays until you choose to remove it.

CappedAI — Model Library
Llama-3.2-3B-Instruct.Q4_K_M
Meta · Chat
2.0 GB
Qwen3-8B-Instruct.Q4_K_M
Alibaba · Chat
5.2 GB
FLUX.1-schnell.Q8
Black Forest Labs · Image
11.9 GB
Phi-4-mini-instruct.Q4_K_M
Microsoft · Chat
2.5 GB
nomic-embed-text-v1.5.Q4_0
Nomic · Embeddings
270 MB
Platform support

Built for Apple Silicon.

CappedAI uses Apple's Metal GPU framework for hardware-accelerated inference — the same GPU that powers your device handles the AI. No cloud. No external GPU. Just the silicon already in your hand.

Platform Chip GPU / Inference Framework Status
iPhone
iOS 16+
A16, A17 Pro, A18
Apple GPU · 5–6 core
Unified memory shared with CPU
Metal ✓ Supported
iPad
iPadOS 16+
M1, M2, M4 · A14+
Apple GPU · up to 10 core
M-series iPads match MacBook performance
Metal ✓ Supported
MacBook Air
macOS 13+
M1, M2, M3, M4
Apple GPU · 7–10 core
Fast enough for 7B–13B models comfortably
Metal ✓ Supported
MacBook Pro
macOS 13+
M1 Pro/Max, M2 Pro/Max, M3 Pro/Max, M4 Pro/Max
Apple GPU · up to 40 core
70B+ models viable on Max chips
Metal ✓ Supported
Mac mini / Mac Studio / iMac
macOS 13+
M1–M4 · M2 Ultra / M3 Ultra
Apple GPU · up to 76 core
Studio Ultra handles 70B+ at full quality
Metal ✓ Supported
Windows / Android
Various No unified GPU framework Not planned

Apple's unified memory architecture means the GPU and CPU share the same memory pool — no data transfer bottleneck. This is why Apple Silicon runs large models faster than discrete GPU setups with the same theoretical FLOPS.

Privacy

Not a policy.
A physical reality.

Every other AI tool processes your queries on a server somewhere. Policies change. Servers get breached. Companies get acquired.

CappedAI runs entirely on your own hardware. The models and your data live on your phone or drive — never on a server. Inference runs on your device's CPU or GPU. There is nothing in the cloud to breach, subpoena, or shut down.

When a pack fetches live data, that's an explicit network call made by the pack — and the result is written to your drive, not to our servers. Your conversation context never leaves your device.

  • Conversations never leave your device
  • No account means no data profile
  • License keys validated locally — no phone-home
  • Live data fetched by packs is stored on your drive, not ours
  • Works in air-gapped environments
  • Generated images stay on the drive
Where your data goes
💬
Conversations
Stored on device / drive
Local only
🤖
AI inference
llama.cpp · Metal · Apple platforms only
Local only
📄
Documents you upload
Indexed locally, never uploaded
Local only
💾
Live data fetched by packs
Written to your drive, not our servers
Yours permanently
🔑
License key validation
RSA signature · baked-in public key
No server call
☁️
Your queries / context
Never sent to any server
Never

Be first to get
CappedAI.

Join the list for launch updates. iOS app and drives coming soon.

🧪 iOS App Coming Soon
💾 Get notified when drives ship

Leave your email and we'll reach out when the iOS app and drives are ready.
Interested in a drive? Just mention it in a reply — we'll be in touch.

✓ You're on the list. Talk soon.

Built in Austin, TX  ·  CappedAI LLC