minipc-01 · fedora · 25 containers · 28 ansible roles

One box.
Everything on it.

Media, files, a local language model and the automations around it, all on one mini PC in a closet. The whole machine rebuilds from one Ansible playbook.

The board

Every part of the machine and what runs on it. Point at a part to see its services, or at a service to see where it lives.

MINIPC-01192.168.0.5
ram
32 GB DDR4 · dual channel the model holds two thirds of it, permanently
  • language model, 35B parameters, ~3B read per word
  • embedding model for document search, 1.4 GB
  • the other 23 containers and Fedora together

Services

Sorted by how far each one is from the internet. The colour is the distance.

publicon the open internet, behind TLS

  • Jellyfin

    Films, shows and music, streamed anywhere, with hardware transcoding on the iGPU.

  • Nextcloud

    Files, calendars and contacts, synced from anywhere. MariaDB underneath, Redis in front.

home networkreachable from the LAN and the VPN

  • Grafana + Prometheus

    Dashboards for CPU and GPU load, both drives' temperature and fill, and the model's live reading and writing speed.

  • Forgejo

    Git hosting for every project on the box, with CI. Each push is built, tested and scanned in a sandboxed Docker, and an app that passes deploys itself, and rolls itself back if it comes up broken.

  • n8n

    The automation engine, wired to the local model. Telegram reaches it through a single public webhook path.

  • Lift document AI

    A 9B vision model. Send a PDF or a photo, get the fields back as a clean PDF.

  • Doc search

    Search over the box's own notes and code. Ask why a decision was made and get the answer with the file and line it came from.

  • Plex

    The second media server, for the TVs in the house.

  • NFS

    The media drive shared with every machine on the network.

  • Pi-hole

    Ad and tracker blocking for every device, at the DNS level.

internalprivate Docker networks, never published

  • llama.cpp

    Ornith 1.5 35B on the Vega 8 through Vulkan, 128k context. A mixture of experts: all 35 billion parameters stay in memory, about 3 billion are read per word. Prompts never leave the box.

  • Embedding model

    A small second model that turns text into vectors for search, kept apart so tuning it never disturbs the main one.

  • Qdrant

    Meaning and exact keywords indexed side by side, so a plain question and a variable name both find the right paragraph.

  • Transmission

    Downloads, driven entirely from Telegram. No web interface at all.

  • MariaDB

    The database behind Nextcloud, on a network nothing else can see.

  • SearXNG

    Web search for the automations and the coding agent, returning short extracts instead of whole pages.

Automations

n8n workflows call the local model. Everything is worked out on the box; only the finished message leaves it, to Telegram.

  1. daily 07:30

    Morning briefing

    Weather, today's calendar, headlines and how the server slept, written up by the model as one message.

  2. mondays 08:00

    Image update watcher

    Checks every pinned container image for a newer release; the model writes the change notice.

  3. mondays 08:15

    Security report

    The week's SSH attempts and fail2ban bans, summarised.

  4. nightly 04:00

    Doc index canary

    Re-indexes the documentation and checks it is still there. Silent when healthy.

  5. on request

    Document extraction

    Send the bot a receipt or invoice photo; the vision model reads it and replies with a filled-in PDF.

  6. on request

    Media sorter

    Send a magnet link. When the download finishes the model names it, splits season packs by episode and files it in the library.

Build log

Everything that changed, oldest first.

  1. 2026-06-04

    git init

    Repo born. Base roles: users, SSH hardening, firewalld, Docker, Pi-hole, fail2ban, automatic updates.

  2. 2026-07-04

    Media era

    Jellyfin with AMD VAAPI hardware transcoding, Plex, NFS media share — plus nginx reverse proxy and Let's Encrypt DNS-01 certificates.

  3. 2026-07-05

    Eyes and cloud

    This website. Uptime Kuma, hardware sensors, Prometheus + Grafana dashboards. Nextcloud on MariaDB with Redis.

  4. 2026-07-06

    AI era

    Ollama serving Qwen3 locally, n8n automation engine, and three LLM-powered workflows reporting to Telegram.

  5. 2026-07-07

    Document AI

    Datalab lift 9B vision model self-hosted on CPU. A gateway rasterizes PDFs, extracts schema-constrained JSON and renders the result back to PDF — driven from Telegram.

  6. 2026-07-07

    Media pipeline

    Transmission wired to Telegram: send a magnet, the LLM classifies the release and files it into the Plex/Jellyfin library. Downloads, sorting and naming, all on-device.

  7. 2026-07-09

    Morning briefing

    The daily digest grows up: weather, Google Calendar events and news headlines join the server metrics, composed by the local LLM into one 07:30 Telegram message.

  8. 2026-07-11

    Sharper brain, faster pipeline

    All LLM workflows switch to Ornith 9B after a CPU benchmark bake-off. The media sorter goes event-driven — Transmission fires a webhook the second a download completes — and learns to sort whole season packs, episode by episode.

  9. 2026-07-18

    Memory doubled, models tripled

    A second RAM stick unlocks dual-channel bandwidth — inference speed nearly doubles overnight. A three-way bake-off crowns Qwen3.6 35B as the new workflow brain, and the lift vision model steps up to a near-lossless quant with double the context.

  10. 2026-07-18

    Chat with the box

    Open WebUI puts a face on the local LLM — chat from any browser, anywhere, behind Cloudflare Access. Tuned for CPU inference: a "hello" that first took five minutes now answers in three seconds.

  11. 2026-07-22

    New brain

    Ornith 35B takes over from Qwen3.6 as the workflow and chat model. Same tokens per second, but it loads in three quarters of the time and answers straight away instead of thinking out loud first — so replies arrive sooner without the box working any harder.

  12. 2026-08-13

    The iGPU wakes up

    Ollama gives way to llama.cpp, and inference moves off the CPU onto the Vega 8 through Vulkan — prompt processing gets five and a half times faster. One model now serves everything, held in memory permanently so it remembers the start of a long conversation instead of re-reading it every turn. Chat, the document reader and every automation share it.

  13. 2026-08-14

    Ornith comes back

    The 9B model that ran the workflows back in July returns as the one brain for everything — chat, the coding agent and all five automations. Same size and the same speed as the model it replaces, because underneath they are the same architecture differently tuned; the swap is a bet on how it writes, not on how fast it runs. Roughly eight gigabytes resident, which leaves the long-conversation cache plenty of room. The previous model stays on disk, one line away, in case the new one disappoints.

  14. 2026-08-20

    The box reads its own notes

    Everything written about this machine — every Ansible role, every design note, every benchmark — becomes searchable by meaning as well as by name. A second small model, permanently in memory, turns the text into vectors; a store keeps those beside a keyword index and merges both rankings, so "why is the cache unquantized" and an exact variable name each land on the right paragraph. Answers come back with the file and line they came from. The coding agent on the laptop now asks the box about the box.

  15. 2026-08-20

    Four times the model, faster answers

    A 35-billion-parameter model takes over from the 9B — and it is quicker on every measure, because it wakes under a tenth of itself for any one word. Prompt reading gains a third, writing speed rises by four fifths, and it scores measurably better on the standard quality test rather than trading accuracy for speed. The cost is space: three times as much memory held permanently, which meant lifting the ceiling the graphics chip is allowed to borrow from system RAM and taking a reboot to do it. The 9B stays on disk as the way back, unused unless something asks for it by name.

  16. 2026-08-21

    Making room

    An accounting of where the memory actually goes: the model is 82% of it, and everything else on the box together is a fraction of one model. The browser chat window turned out to be the largest thing that was not the model — partly because it was quietly loading a second embedding model of its own, duplicating a job the retrieval stack already does better. That got repointed at the real embedder, and the chat front end retired. The same model still answers the coding agent, the automations and the document search; only the browser window is gone. Nothing is deleted — accounts and history sit on disk behind one flag.

  17. 2026-08-23

    One model, newer

    The brain moves to the next version of the same model — same shape and same settings underneath, only retrained weights, so nothing else on the box had to change. Everything else is cleared off the disk with it: the older versions kept as a way back are gone, and what remains is this model and the document reader, nothing more. Two honest caveats, because the notes say so too. No side-by-side speed or quality test was run first, so "newer is better" is a reasonable expectation rather than something measured here. And with no spare model on disk, going back means downloading one again instead of changing a single line.

  18. 2026-09-27

    A second drive

    An NVMe drive goes into the empty M.2 slot and the whole media library moves onto it — films, shows, music and the download folder together, because the sorter files a finished download by linking it rather than copying it, and that only works when both live on the same drive. The system drive gets back more than half of its space, and the models, databases and everything else stop competing with the library for it. Both drives now report their own temperature and fill level on the dashboards and in the morning briefing, with separate limits, because an NVMe drive runs warmer by nature.

  19. 2026-09-28

    Push to deploy

    Every project moves into its own repository on a Git server on the box, and a push now does the rest. The code is built and tested in a sandboxed Docker that cannot touch the rest of the machine, checked for known vulnerabilities, and started once against an empty database to prove it boots. Only then does the box swap it in, and if the new version does not answer like itself within a couple of minutes, the old one is put straight back. On the first run the vulnerability check caught a web framework 62 security fixes behind; it was updated the same hour. Dependency updates now arrive as a weekly pull request that has to pass the same checks.

Stack

os
Fedora Server · SELinux enforcing
containers
Docker + Compose · 25 running
automation
Ansible · 28 idempotent roles · secrets in Ansible Vault
ci/cd
Forgejo + Actions in Docker-in-Docker · tests, audit, image scan · Renovate · pull-based deploys with health-gated rollback
inference
llama.cpp + llama-swap · Ornith 1.5 35B mixture of experts, 128k context · lift 9B vision model
retrieval
Qwen3 embeddings · Qdrant hybrid dense + sparse · key-gated FastAPI gateway
ingress
nginx · Let's Encrypt DNS-01 via Cloudflare
data
MariaDB 11 · Redis · isolated Docker networks
security
SSH hardening · firewalld · fail2ban · automatic updates
monitoring
Prometheus + Grafana · smartd on both drives · Telegram alerts