Post · #ai

Running open-source LLMs on a homelab

I wanted to know a simple thing: how far can you get running open models on hardware you already own, before renting a GPU stops being optional. The honest answer is "further than I expected, right up until it isn't." Here's the map.

What fits, and what doesn't

A 7B model quantized to 4-bit runs on the Pi 5 with 8GB, slowly. It's fine for a chat you're not in a hurry for, and useless for anything streaming. The moment you want a 13B or real context, you're swapping, and swap on an SD card is its own kind of punishment.

The setup that actually held up was boring: ollama on a mini-PC node, the Pis doing everything else.

# pull a small instruct model and talk to it
ollama pull llama3.2:3b
ollama run llama3.2:3b "summarize why my cluster is on fire"

The homelab isn't about saving money. It's about having somewhere to be wrong cheaply.

The rule I landed on

  • Prototyping and learning: local, quantized, patient.
  • Anything with a user waiting: rent the GPU, no guilt.
  • Keep the model layer swappable so the decision stays reversible.

Next I'm wiring this into a small agent that can read my cluster's events. That's the next note, assuming the Pi survives it.

← all writingkoreissi.com