Brad Westness
  • About
  • Archive
  • Categories
  • Feed

A Coding Agent on Every Desk, Part 3: A Fast Model in a Sidecar

Software AI Home Lab DIY

Last time in this series, I talked about configuring my home lab to run on two GPUs, which itself followed part one of configuring the home lab. This time I didn’t add a second GPU,b ut a second local model: a small, fast “classifier” model next to the main server, doing the small, fast tasks the large model is the wrong fit for.

A street-racing motorcycle with an integrated sidecar, in the manner of the Sidehackers from Mystery Science Theater 3000.
  • 15 Sep 2026

A Coding Agent on Every Desk, Part 2: Two GPUs and a Bigger Context Window

Software AI Home Lab DIY

The second graphics card arrived, and so did a new model. This is a follow-up to A Coding Agent on Every Desk. The Podman, Tailscale, and Open WebUI pieces from that post are unchanged. What changed is the card in the second slot, the model behind the qwen-coder alias, and most of the llama-server command line.

The inside of an open desktop computer case with an MSI GeForce RTX 3090 installed above a Gigabyte GeForce RTX 4060 Ti, a yellow-lit motherboard, and Arctic case fans.
  • 07 Sep 2026

Rethinking Build vs. Rent in the Era of Coding Agents

Software AI Infrastructure Libraries Testing

Back in 2014, I wrote a little thought experiment called Not Invented Here Mechanic.

Two industrial robots work on a partially assembled silver car body inside an automotive factory.
  • 28 Aug 2026
  • « Older Posts

© 2026 Brad Westness

Brad Westness

Wisconsinite, Software architect, Brewers baseball fan, he/him

  • Bluesky
  • GitHub
  • Mastodon
  • Email
This website uses cookies to ensure you get the best experience. Learn more