← Writing

Running small open models on hardware you already own

Open-weight models run on a machine I administer, using software I can inspect.

I put open weights on hardware I run. The files sit on my disk, and the inference software is something I can read and pin. I have watched hosted tools change behavior without warning. I do not give them drafts I care about.

I do this on a laptop or desktop I already own. A small model that fits in RAM is enough to keep the work local.

Start with constraints

Before I download weights I write down how much system RAM the machine actually has after the desktop and the browser take their share. VRAM on a discrete card is useful when this box has one. Most of the laptops I still carry do not.

I also write down how much disk I can spare for GGUF files, and whether that disk has to answer with the radio off. A 7B quant is a few gigabytes. Extra quants I might try later are how a slice I meant to keep honest fills up. If the machine has to work offline, the model has to live on that disk before I leave the house.

Those answers narrow the field faster than any leaderboard. A 7B-class model quantized to 4-bit drafts and rubber-duck debugs on 16 GB of RAM if I leave headroom for the desktop. It has clear limits next to a cloud frontier model. Local inference is about sovereignty and iteration.

Sovereignty means the prompt and the draft stay on a disk I can name. Iteration means I can run the same file next week and know which binary answered. I want both before I care how a scoreboard moved.

Last spring I sat at this desk with free -h open. The box reports 16 GB. After Omarchy and a browser sitting on a PHP project, I had about 9 GB I could spare for a model. The SSD had a slice I was willing to fill, about 40 GB. I needed the laptop to keep answering on a train. That page of notes picked a 7B Q4 before I opened a comparison chart.

I did try a 13B Q5 the same afternoon. The first prompt after a cold load swapped the machine into the disk and made the editor hitch. I deleted that file. The 7B stayed. Extra quants of the same 7B make the restore slower. I keep the one I hashed.

Open weights are the point

When the weights and the inference stack are open, I can inspect what changed between versions. I can run the same prompt on Tuesday and on Thursday and know whether I moved a file. Sensitive drafts stay on disk I control.

The failure modes are inspectable. For personal writing and for code experiments on a box I administer, that trade is often worth the slower tokens.

Open weights are a file I can hash and a runtime I can read. A hosted chat window can look the same on Wednesday and still be a different model than the one that answered on Monday. I have watched that happen to a client outline I should have kept here. The window did not warn me. The next morning the tone had moved and I could not name the build.

I keep the GGUF in a directory I created. The filename has a quant tag I chose. After the download I run sha256sum and keep the digest next to the file. If a later page offers the same 7B with a friendlier name, I hash it before I point the runtime at it. A different digest is a different model.

A vendor that serves weights from a URL I do not pin can swap the blob between visits. A directory I administer cannot do that unless I copy a file into it. That is the control I wanted when I stopped pasting drafts I care about into a window I do not run.

I learned the same lesson on catalogs I used to administer. If I could not name the binary that opened the table, I did not know which answer I had. A model file is the same kind of object.

Runtime on this desk

The desktop under this habit is Omarchy, on a Linux box I built and still administer. The ISO does not ship a model. Install > AI will offer Ollama or LM Studio. I still pick the path and I still bring the file.

Ollama is the runtime I leave running when I want a prompt from a terminal. LM Studio is the one I open when I want to see the loaded file and the context length in a window I can read. Both point at the same directory when I tell them to. I pin the package versions in the tree I keep for this box.

When the work is code, I type c and OpenCode talks to that runtime. The stub is a line I can open. I have left the hosted-agent stubs unwired on this install. A coding harness that can edit the tree still needs a review step I can see.

I opened c on a PHP include I already intended to change. The review listed the edit before it applied. I accepted the hunk that renamed a function I had already picked. I discarded the rest, and the file stayed mine.

The model is a file on a disk I can name. The runtime is a package I can pin. If both sit on hardware I administer, I can run last month’s prompt and know which binary answered.

I did not wait for a new chassis to start this. The laptop I already owned ran the same 7B. The desktop I built later made the same files faster. The habit is the directory and the pin.

On a week with a ship date I want that pin more than I want a new score. I open the same terminal and load the same GGUF. Predictability is a file time I can read.

A practical first week

Pick one runtime you are willing to read the docs for, and download one small model for one task you already repeat. I used Ollama and a 7B Q4 on commit messages for a week of ordinary client work.

Log where it helps, and log the moment you reach for something bigger. After a week you will know more than another afternoon of forum arguments.

I kept a plain text file next to the weights. Monday’s commits were boilerplate I could have written by hand, and the 7B drafted them in a shape I still edited. I changed two lines on most of them. I kept the file.

Later that week I asked the same model to reason about a schema change that spanned more files than the context would hold. It started inventing columns I do not have. I stopped, used a hosted model for that throwaway, and came back here for the draft I still had to stand behind.

Two notes from that week still sit in the log. The 7B is enough when the task fits in the window I can see. I reach past it when the task does not.

I did not add a second model that week, and I did not switch runtimes. The point of the week was one path I could describe without opening a forum. Ollama and one GGUF on commit messages were enough to know where the local file earned its keep.

Checksums and pins

A checksum after a download is how I know the file I kept is the file I meant to keep. A pin on the runtime is how I know which binary loaded it. When a repeated prompt drifts, I look at this disk before I look at a status page.

After a weekly update the answers on a commit prompt I reuse started to wander. I hashed the GGUF. The digest matched the note I had written in April. The Ollama package did not match the pin. I rolled the package back. The old prompt matched again. I would have blamed the model if I had not kept both numbers.

I keep the digest in a text file beside the GGUF, in the same directory the runtime is allowed to read. If I have to rebuild the box I want that pair in the tree I clone. A filename is a hint. A hash is the record.

When I still use a hosted model

A frontier model on someone else’s hardware is fine for a throwaway I would paste into a public search box. A regex I cannot remember. A public API shape I have not looked up this year. That work can leave the building.

A draft I have to stand behind stays here. Client notes stay on this disk. So does a PHP file that still has a password in a comment I have not cleaned up.

DHH has said he reaches for hosted models when he wants the strongest ones. I believe him. I still bring the weights onto this disk for the work I refuse to send out. The strongest window is still a window I do not pin.

Last month I used a hosted model to remind me how a public pagination header is supposed to look. I would have searched for that page anyway. I wrote the client handler on this box against our own catalog. The hosted pass was a throwaway. The file I committed was local.

What I am not claiming

Local models will not replace every API call in a workflow tomorrow. They work when I scope tasks they can handle on this hardware.

A 7B Q4 on 16 GB will lose to a hosted frontier model on a long refactor. I already know that. I still run the 7B because I can name the file that answered, and I can run it again in October.

That is the sovereignty I wanted when I put the first GGUF on this disk. Iteration is the other half. I can open the same directory next month and get the same binary to load the same weights. A scoreboard does not give me that.