A lot of small businesses hold information they would rather not paste into a cloud AI: client files, payroll, medical notes, draft contracts, a customer list. The usual advice is "don't use AI for that". There is a better option. Run the model on your own computer, and the text never goes to an AI provider at all.
This guide covers how to do that with Ollama or LM Studio and Murage, what a local model is good at, where it struggles, and how to split your work between local and hosted bots.
What "local" means here
When a bot works, it sends what it needs for that turn to its engine: your message, its instructions, relevant memories and the files it is working on. With Claude Code, Codex or a provider key, that goes to the company that runs the model.
With a local model, the model runs on your own machine, or on another machine you own. The turn never leaves your network. Murage itself keeps your bots, memory, conversations and files on your computer either way.
What you need
- A model server. Ollama and LM Studio are the easiest. Murage also works with llama.cpp, vLLM and EXO.
- A model that can use tools. An agent does more than chat. It reads files, writes files and hands work to other bots, so it needs a model that handles tool calls.
- Enough computer. Bigger models need more memory. A recent Mac with plenty of unified memory, or a PC with a good graphics card, runs useful models. Small models run on almost anything, just with less skill.
Set it up
- Install Ollama or LM Studio and download a model with it.
- Start the model server.
- In Murage, open Settings → Models → Local models. Murage checks this computer for running model servers and lists what it finds. A server on another machine on your network or your tailnet can be added with Add a server.
- Choose the Test button next to a model. Murage runs seven checks on your computer, such as whether the model calls a tool when it should and whether it uses the result. No cloud provider is called and nothing is billed.
- When the test says Tools work, choose Use with a bot.
Murage sends a model server nothing until you pick one of its models for a bot.
A tested model can be used by a bot on the bundled Fuigo engine, and on several agent engines, including Hermes, OpenCode, Qwen and pi. The full list is in Connect an AI engine.
One Ollama tip: agents need a large context window, and some Ollama models ship with a small one. If the test flags it, Murage offers to create a copy of the model with a bigger context, then asks you to test that copy.
What local models do well
- Sorting and tagging. Reading a pile of emails, tickets or files and putting each one in a bucket.
- Summarizing private documents. Contracts, client notes, HR files.
- First drafts from your own notes. Letters, checklists, internal updates.
- Extracting details. Names, dates, amounts and promises from a messy document.
- Always-on, high-volume work. Local models cost nothing per call, so a bot can churn through a backlog all day.
Where they struggle
Be honest with yourself about the trade-off.
- Hard reasoning. Long plans, tricky analysis and careful writing are still better on the largest hosted models.
- Coordination. Your Chief of Staff plans, briefs and judges. It benefits most from a strong engine.
- Speed. Big models on ordinary hardware are slower than a hosted service.
- Tool use. Some models answer but can't use tools. The test tells you before you rely on one.
Mix local and hosted bots
You don't have to choose one for the whole business. In Murage, every bot has its own engine. A common pattern for a privacy-minded small business:
- Chief of Staff on Claude Code or Codex, for planning and judgment. It sees your requests, never the private files.
- Client-file reader on a local model. It reads and summarizes the sensitive documents.
- Writer on a hosted model. It polishes non-sensitive drafts.
- Tagger on a small local model. It sorts the inbox all day for nothing.
Give each bot only what it needs. The bot on the local model works in its own folder on your computer, and that is where the sensitive files go. See pick an engine per bot for more on matching engines to jobs.
What still leaves your computer
A local model keeps the AI turn private. It is worth knowing what else can go out, so you can decide:
- Web search. If a bot searches the web, the search words go to the search service. Tell a private bot not to search, or leave search off for that job.
- Connected apps. If a bot works in Gmail or Slack, those services see what it does there.
- Other bots. If a local bot hands its findings to a bot on a hosted engine, that bot's engine sees them. Keep sensitive work inside the local bots, and ask them to hand over only summaries.
- Update checks. Murage checks GitHub for new versions.
Today's releases send no usage analytics. Anonymous, opt-out analytics are planned. The full list is in the privacy and usage analytics guide.
A note on memory
Murage's memory lives on your computer in any case. Semantic memory search uses a small model that Murage downloads and runs locally, so recall doesn't need a cloud service either. On Intel Macs it uses keyword search instead.
Get started
Download Murage, start Ollama or LM Studio, and open Settings → Models → Local models. Test one model, put it on one bot, and give that bot a folder of documents you would never paste into a chatbot. The local models page has more detail, and why Murage runs on your computer covers the rest of the privacy picture.
