Skip to content

g and Homebase: A Local AI System I Own

A local AI system on hardware in my house: a terminal agent, a personal memory, an offline knowledge corpus, and a Slack bot, built to keep working with no subscription and no internet.

Overview

I wanted two things, in that order. First, coding help and real answers to real questions that cost nothing per token and kept working with the Wi-Fi off. Second, a hedge: if frontier models get expensive or go away, I wanted to already own the weights, the harness, and the data, on hardware in my house, as plain Markdown and git history.

The readme describes the whole system in one line: four things exist, a laptop that thinks, a laptop that never sleeps, a drive that remembers, and a phone that asks.

Fully autonomous, agentic app-building on local models at this size is not reliable yet. Real coding work routes to Claude Code while I'm subscribed to it, and only falls back to the local models when I'm not, so the tool design and safety rails below are built for that fallback to actually work.

The four machines

The brain is an M4 MacBook Pro with 16 GB of unified memory, running every real inference through Ollama and a CLI I built called g, plus a job runner that lets other machines hand it work. It sleeps, the same as any laptop, and everything else is designed around that.

Homebase, the machine the project is named after, is a 2017 13-inch MacBook Pro, a dual-core i5 with 8 GB of RAM, wiped and reinstalled hands-off from an Ubuntu Server 24.04 autoinstall stick. It runs with the lid closed, twenty-four hours a day, on wired ethernet, holding the canonical copy of my personal data, the Slack bot, the jobs API, and one small model on CPU. Its battery doubles as a free UPS, good for four to six hours idle, and it replaced a retired PC tower that idled at around 100 watts and lost to it on power, noise, and CPU. Its first-boot script taught me something too: the first version ran on every boot and kept undoing my own configuration changes, so the version that shipped joins Wi-Fi and the network once, confirms it, and deletes itself.

The drive that remembers is a 2 TB SSD on the M4 holding the model archive (42 files, 46 GB), the offline corpus, and nightly backups, never the only copy of anything that matters.

The phone that asks is Slack. Between the machines themselves, a private mesh VPN is the only path, and every service binds only to that network, so the bind itself is the access control, with a firewall behind it as a second line. Slack only connects outbound, through Socket Mode, with no public URL, and nothing load-bearing ever travels through it.

The models, and the seam that lets them change

The M4 runs Qwen3.5 9B at Q8 for chat and retrieval, Qwen3 14B dense at Q6_K (12.1 GB, 16K context) as the default coding model, Qwen3-Coder 30B-A3B as a slower second option, Qwen2.5-VL 7B for vision, and nomic-embed-text for embeddings. Homebase runs llama.cpp on CPU: Qwen3.5 2B as a fallback chat model at around 12 tokens a second, plus a second embedding server so retrieval keeps working while the M4 sleeps. The two machines' embeddings agree at a cosine similarity of 1.00000, checked rather than assumed, so either index is interchangeable with the other.

Every model is reached through config, never a hardcoded address. A future Mac with 64 or 128 GB of memory slots in by changing one endpoint, and the harness, the context engine, and the bots, the parts meant to last, never notice. g states its own design goal plainly: local to homebase to frontier is an edit, not a patch. Configured endpoints fail over in order, each given a 2.5 second window before the next is tried.

The bake-off

I assumed a mixture-of-experts coder in the 30B range would be the default, since MoE plus memory-mapping should let a model that size fit under real memory pressure. I measured it instead of trusting the assumption, the house rule this project runs on: measured, not assumed. Qwen3-Coder 30B-A3B decoded at 0.58 tokens a second, only 61 percent of it landing on GPU. Qwen3 14B dense decoded at 7.87 tokens a second with 89 percent on GPU, thirteen times faster. A later fix raised the dense model's context to 16K and cost it some speed, down to about 4.7 tokens a second, still roughly eight times faster than the model I had planned to default to.

I kept both and inverted which one is the default. g bench reruns that comparison through the real agent loop any time a model, quantization, or llama.cpp version changes, and a nightly job runs it automatically, flagging a regression only when a pass turns into a fail, because a cron job that pages on a slow run gets muted within a week.

g, the terminal agent

g is for Greg. One command fronts every model I run, from the chat and code models on the M4 to the small models on homebase to a frontier API on the rare occasion nothing local will do, plus the personal context engine all of them read from, effectively a Claude-Code-style terminal agent pointed at models I own. It's TypeScript on Node 22 with eight runtime dependencies and no LLM SDK, just a hand-written streaming client that speaks OpenAI-style server-sent events and Ollama's native API.

g chat is the agentic REPL, an Ink and React terminal interface where the model reads, writes, edits, and runs bash behind an approval gate, with about 25 slash commands like /context, /model, /compact, /undo, and /resume. g ask (also what runs when you just type g followed by a question) answers in one shot with retrieval and cited sources, with an --offline flag for the offline corpus, and g do runs one headless agentic turn for scripts and cron, with real exit codes. g note, g todo, and g inbox append into the context repo with a git commit in about a tenth of a second, and g review's weekly digest keeps its model summary opt-in, since a digest that always waits on a 9B is a digest you stop running.

Twenty-three agent personas sit on top of this: an orchestrator, PMs, a designer, engineers, QA, a copywriter, a researcher, a chief of staff, addressed with @name and a task for a bounded sub-turn, or pinned with a bare @name for the session. Sub-agents cannot spawn further sub-agents. A command is a prompt you fire, a persona is who you are talking to.

The tool surface is deliberately small, read, write, edit, and bash, since a small local model chooses better from a short menu and bash already covers ls, grep, find, and git. Describing those tools in prose in the system prompt once pushed the coder model toward hand-writing XML tool calls the server couldn't parse, so the prompt says nothing about them now, and a salvage parser catches what slips through. The edit tool matches loosely, but only so far: whitespace-tolerant matching took one benchmark task from two tool errors and six rounds down to zero and two, and I stopped there on purpose, since a looser match can span far more of a file than intended, silently.

The rest is craft that only shows up in daily use: a warm-up request hiding a 9.1 second cold start, a context gauge /context draws as a stacked bar, and a completion chime keyed to 30 seconds of keyboard idle instead of turn length, because on a local box every turn is slow. The KV cache mattered more than expected: a cold prompt took 23.4 seconds, the same prompt again took 0.15, and the same prompt with one token changed at the front took 24.3, since that early change invalidates everything cached after it, so stale tool results now get trimmed in steps of four turns to protect the cached prefix.

93 commits over eight days, ten releases from 0.1.0 to 0.4.1, 479 tests across 66 files, about 14,000 lines of source against roughly 7,500 lines of tests.

Keeping the agent honest

An agent that can write files and run bash on your own machine needs real limits. g cycles through four modes with shift-tab, from asking before every change to accepting edits automatically, to a read-only plan mode, to no tools at all. An "always allow" grant is scoped to a program only when its name states what it runs; wrappers like env, xargs, and ssh are scoped to the exact command string instead, and any pipe is always matched exactly. The approval preview shows a real diff, /undo checkpoints every write, and headless g do denies every change unless you pass --allow or --yolo. --root confines file access to a subtree and resolves real paths first, closing off symlinks pointing outside it.

A self-audit on September 4, at 400 of 400 tests passing, found one harmless merge bug and two real trust boundaries: read itself was ungated, which combined with Slack posting meant a file could leave the machine just by being read and summarized, and "always allow" grants were keyed on a command's first word, so allowing env once meant allowing it to run anything afterward. I closed both the same day, one with a program allowlist, the other with --root confinement. The audit document states plainly what was checked and what wasn't.

The context engine and the offline corpus

Personal memory is a git repository of plain Markdown, an inbox, todos, notes, a knowledge base, a people file, plus a derived vector index built with sqlite-vec. The canonical copy lives on homebase, the M4 works from a clone, and every push reindexes it, so my memory survives any single vendor, tool, or decade. Chunking depends on file type: an append-only log gets one chunk per line, prose splits at headings, and every chunk keeps its heading trail. The relevance floor of 0.62 cosine came from checking real queries, on-topic ones scoring 0.71 to 0.83, off-topic ones 0.44 to 0.57, so the floor sits in the measured gap between them. Backups run nightly with restore drills against them, because a bundle nobody has restored is a file, not a backup, and the nightly fetch fails loudly on purpose, since a stale bundle that looks fresh is worse than no bundle at all.

Underneath that sits an offline corpus for the day both frontier APIs and my local models are gone: 33 Kiwix archives, about 90 GB, full-text searchable with no embedding model needed. English Wikipedia's text at 52.7 GB, Wiktionary, Project Gutenberg, developer docs for JavaScript, TypeScript, Node, React, and Python, a WCAG and ARIA accessibility corpus, and 1,687 pages of Apple's Swift and SwiftUI documentation converted into an archive. It's slow, about 200 seconds for the first search across all of it cold and 10 to 17 seconds after that, and that's fine for the job it has. The refresh policy says it best: this corpus is the floor, not the ceiling.

The phone that asks

One Slack app in Socket Mode on the free tier, with one persona per channel. #assistant writes inbox, todo, and note entries into the context repo. #ask answers with retrieval and citations, using the M4 when awake at around 17 seconds, or the small model on homebase at around 4 when it's not, tagged as a quick answer. #ops runs and controls background jobs with commands like run <repo>: <task>, plus status and cancel. #do takes a task and is read-only by design, since a door you reach from your phone, where you can't see a diff before approving it, should not be able to write anything. #status announces boots, unclean shutdowns, and network recovery, and a user allowlist gates every handler. Model Markdown renders as raw asterisks on a phone screen, so I built a Markdown-to-Slack converter for it, and three of its bugs never showed up until I actually posted to a real channel and read the result on my own phone.

The job system

A small queue service runs on homebase, Fastify in front of SQLite, and a runner on the M4 polls it every 15 seconds and runs one job at a time, always through g do. A stranded job gets failed by a reaper rather than re-queued, since a spurious failure just costs a re-run while a spurious retry could corrupt a repository. Ask, summarize, and private jobs get 21 read-only programs, things like ls, grep, and diff, with pipes, redirects, and substitution denied even inside that list, confined to an empty scratch directory, since a prompt that arrived over Slack must not be able to read this machine's real files just by asking. I also found and closed a case where the model API was bound to the home network rather than the private mesh, so its pull and delete endpoints could technically have reached the model weights from any device on the LAN. One risk I accepted openly: homebase has no disk encryption, since that needs a passphrase typed at every boot, and a machine that boots with the lid closed has no way to type one.

Parent and child jobs, and scheduling on homebase instead of the M4, are built and unit-tested but still waiting on live verification, and automatic power-on after an outage and automated corpus refreshes are both still open.

What broke

g never sent a context size on its requests, and sending one wouldn't have helped, because Ollama's OpenAI-compatible endpoint silently ignores that field. Every model had been running at Ollama's 4,096-token default the whole time while the status line's gauge assumed 32K, with long prompts truncated from the front and no error anywhere, the chat model included. Switching to Ollama's native API fixed it.

One coder model I benchmarked went further than a normal failure: after rejecting a set of tool calls, it fabricated a tool call and a success payload for a file it had never written. I disqualified it on the spot. A model that fails honestly is fine, since I can build around a known failure mode; one that convincingly claims it acted is worse, because it removes the signal I need to catch it.

Fake test doubles cost me twice, both times by hiding a real bug behind something that only looked like the real system. Claude's headless mode prints nothing until the process exits, so a mock that emitted progress as it ran never caught that no progress was reaching Slack in production, and separately, a diff summary once credited my own uncommitted edits to a job that hadn't touched those files. Both surfaced only once I stopped trusting the mock and read the live queue and Slack output myself.

The default coding model sometimes ends a turn mid-thought with a bare "terminated." g sends one nudge asking it to finish, and the nightly bench run counts how often that nudge fires, so the defect stays a number instead of an anecdote. Running a laptop as a server surfaced its own quirks too: a lid-closed reboot powers the machine off instead of restarting it, the ethernet adapter only appears if plugged in at boot, and the machine doesn't power itself back on when AC returns, so silence in #status after a remote reboot means go open the lid.

One finding is still open. The job system's read-only confinement and the context engine's retrieval can combine into a job that gets handed citations it has no way to open and reports success without having answered the question, the worst version of the failure because it looks identical to one that worked.

What it taught me

The homelab side of this ran 99 commits in seven days against 49 tracked issues, 42 of them closed, and 91 of the 93 commits on g carry a Claude Code co-author line. I lead design for a living, not backend infrastructure, and I built almost none of this by typing it myself: the code came out of Claude Code while I made the calls that mattered, where the trust boundaries go, what fails loudly and what's allowed to fail quietly, what gets measured before I believe it.

What stuck with me most is how often the right fix was a boundary: binding the model API to the private mesh, confining a Slack-triggered job to an empty scratch directory, making #do read-only because a phone can't show a diff, and deciding which failures post to Slack, because on a machine nobody is watching, silence and "it worked" look identical. Those calls are why I trust homebase to run unattended with the lid closed, doing work I'd otherwise pay a subscription for.