The easiest way to run local AI on your Mac.

Install in one click. Chat, code and make images with open models that run on your Mac's own GPU. Every dial is within reach, and nothing sits between the model and your hardware to slow it down.

Free · Version 2.4.0 · macOS 13+ · M1 or later
SHA-256 …

New chat

Switch between and in the sidebar, just like the real app. Code mode shows Sabine running tools on a project.

The idea

Easy by default. Powerful when you reach for it.

Most local-AI tools make you pick one: a friendly app with the knobs hidden, or a pile of flags and servers you assemble yourself. Sabine is built to refuse that trade.

1 click

Installs itself

Open the disk image and run the installer. It sets up Python, the libraries and the app. Pick a model, or let Sabine pick one that fits your Mac.

0 layers

Nothing in the way

The model runs in-process on Apple's MLX, straight on your GPU. No local server to proxy through, no wrapper re-reading your whole prompt every turn.

All dials

Yours to turn

Model, voice, content policy, thinking budget, permissions, spend limits, memory. The simple path is the default; every control is one click deeper.

Speed

Fast because it does less work, not because it cuts corners.

On a local model the slow part is usually not generating, it's re-reading. Sabine is engineered around not doing that.

Where one reply's time goes

A 200-token reply on an M-series, 16 GB Mac, from Sabine's own measurements.

Re-reading the conversation every turn 32 s

Sabine: attention cache kept between turns 16 s

re-computing old textreading only the new messagegenerating the reply

Bars are drawn to scale of the figures in Sabine's README. Your numbers depend on your chip, model and prompt; the stats line under every reply shows the real ones.

Measured, not promised

Sabine ships with an eval harness that runs real tasks through the app's real chat loop and grades the result.

0tokens/sec, a 3B model on an M3 (faster than you read)
0smaller system prompt in Code mode (6,000 → 660 tokens)
0less wait for the first token in Code mode (40.6 s → 7.8 s)
0score out of 100 on 14 agentic coding tasks, up from 61
  • Native MLX. Apple's own array framework, 4-bit weights, unified memory.
  • Thinking has a budget. A hard cap on reasoning tokens, and none at all for a quick “thanks”.
  • Tuned per mode. Code gets a lean prompt and low temperature; chat keeps its personality.

Control

Two dials. Every voice at every level.

Personality is how Sabine talks. Policy is what it will answer. They're independent, so you can have a blunt answer that still flags real stakes. Drag them and see.

bluntwarmtutorprofessional
1 loosest2345 strictest
off5121.5k4k8k tok

Levels 1–4 share one hard floor. The settings are real; the replies on the right are an illustration of tone.

ILLUSTRATION · “Can I get away with skipping my tests before this release?”
Can I skip tests before this release?

Off until you say so

File access, web search, a real browser, image generation: every capability is a switch that starts off, per mode. While it's off, the model isn't even told it exists.

Approve or deny

When the model wants to do something sensitive, it stages a request. Approve or deny it right in the chat, on the record.

A model per mode

Chat and Code each keep their own model, voice and switches. Use a small fast model for chat and a bigger one for code.

Incognito chats

Chats that keep nothing: not the text, not the pictures. Local files from them are cleared when the server starts.

Features

Everything in one real Mac app.

A native window and Dock icon, the same interface in a browser, and a terminal chat if that's how you live.

A model that fits your Mac

Tell Sabine your chip and memory and it recommends the best local model for chat and for code, and warns you when one fits but would crawl. 43 models to browse.

—CHAT
—CODE

Local by default, online when you choose

One conversation, two engines. Flip to a frontier model through OpenRouter with your own key, with the same prompt, personality and commands.

Weights on your disk. Nothing leaves the machine to run the model.

A Code workspace

Reads, writes and edits files, runs your tests and fixes what fails, in a project folder you choose. Asks before it touches anything.

Skills

Expert playbooks the model can load, in the same SKILL.md format Claude Code uses. A UI/UX one ships built in; add or import your own.

Image studio

Prompt, aspect, how many, and a searchable gallery. Run FLUX or Z-Image locally on MLX, or any image-capable model on OpenRouter.

Memory that's one paragraph you can read

Sabine learns lasting facts as you talk and brings them into every chat. It's a single paragraph: paste in a ChatGPT or Claude memory, edit it, or wipe it. Duplicates are dropped automatically.

“Prefers concise answers. Works mostly in Python and Swift. Building a menu-bar app. Lives in Lisbon.”

⌘K for everything

Every action, setting, model and conversation in one fuzzy search.

⌘Kswitch mod…

Spend tracking and a hard credit limit

Using online models? Every message is priced, counted in a ledger, and stopped at a daily, weekly or monthly cap you set. Local is always free.

This week · OpenRouter$1.84 of $5.00

The internet is opt-in

/online adds search and page fetching. /browser gives it a real Chromium that keeps your logins. Both start off, and you can narrow them to the sites you allow.

Signed updates

Updates download in the background and install on quit. Each is cryptographically signed; Sabine refuses anything that doesn't verify.

Attach, edit, export

Drag in files, edit a message to rewind the conversation, pin chats, export as Markdown. Syntax-highlighted code, offline.

App, browser or terminal

~/Sabine/run for the CLI, run web for the browser. One engine, three doors.

Models

Open-weights models, ready to run.

Pre-quantized to 4-bit from mlx-community, so a 3-billion-parameter model fits in about 1.7 GB instead of 6. Search them, filter by family, show only what fits.

Privacy

What you type stays on your Mac.

The weights sit on your disk and your GPU does the work. No API key, no account, and the server binds to your machine only.

  • Works with the Wi-Fi off. The interface has no build step and loads nothing from a CDN, no web fonts, no remote icons.
  • The network is a choice. Only installing and downloading models use it, plus whatever online features you switch on.
  • Accessible by design. Text meets WCAG AA contrast, and every looping animation respects reduced motion.

Install

From download to chatting in a few minutes.

No Homebrew, no virtual environments, no config files.

Open the disk image

Double-click Sabine.dmg in your Downloads folder.

Run “Install Sabine”

Terminal shows progress while it fetches Python and the libraries (about 600 MB, a few minutes).

Pick a model

Sabine opens. Choose one in Settings; its weights download the first time you do.

macOS may say it can't verify the installer. It isn't yet signed with an Apple Developer ID. Right-click Install Sabine, choose Open, then Open again, or press Open Anyway in System Settings → Privacy & Security. You can check the download against the SHA-256 above with shasum -a 256 Sabine.dmg.
MacApple Silicon: M1, M2, M3, M4 or later. Intel isn't supported.
macOS13 Ventura or later
Memory8 GB runs small models; 16 GB or more is comfortable.
DiskAbout 2 GB for the app, plus 2–40 GB per model.
InternetFor installing and downloading models. Chatting locally needs none.

FAQ

Questions

Is it as smart as ChatGPT or Claude?

No, and it doesn't pretend to be. Local models are a fast, private helper for drafting, summarizing, explaining code and quick questions. They can lose the plot on long multi-step reasoning. When you need more, switch the same conversation to a frontier model with one key.

Where does it install?

Sabine's code, Python environment, settings and chats live in ~/Sabine. The app is ~/Applications/Sabine.app. To remove it, delete both.

Does my data leave my Mac?

Not in local mode. Only the installer and model downloads use the network. If you turn on online mode, search or browsing, those specific requests go out.

How do I update?

Sabine checks for updates itself, downloads them in the background and installs them the next time you quit. Each update is cryptographically signed, and Sabine refuses anything that doesn't verify. You can also check by hand in Settings → Updates. Your settings, memory and chats are never touched.

Prefer the terminal?

After installing, ~/Sabine/run starts the command-line chat, and ~/Sabine/run web serves the same interface in a browser.

Is the content policy a lock?

No. It's a prompt, and Sabine says so plainly. Levels 1–4 share a hard floor (no routes to chemical or biological agents, mass-casualty weapons, working malware, or sexual content involving minors). “Off” sends no prompt at all, so it has no floor. What each level is worth is measured in the repository's policy-eval.

Put your Mac's GPU to work.

Free to download. Install in minutes. Nothing leaves your machine.

Download for Mac

Version 2.4.0 · Apple Silicon · macOS 13+