Day to day we code with flagship models from Anthropic and OpenAI. Full dependence on someone else’s infrastructure is, however, a real risk, so we spent a week measuring the alternative – Qwen3.6‑27B running locally on a Lenovo ThinkStation PGX (NVIDIA GB10, DGX Spark platform), borrowed courtesy of firit.pl (thank you!).
We got the hardware running without issues, and the model too, more or less – but it isn’t yet ready for our main workflow. What held us back wasn’t quality, but rather single‑session throughput (around 10 tok/s, a spec merge in 14 minutes vs. 1-2 minutes with Claude, complex analyses occasionally timing out) and the operational maturity of the stack stabilizing the backend took a while.
The economics close out the picture: at our current usage (we’re not tokenmaxing 😉) the hardware would pay for itself only over a horizon of years. The purchase would become rational only if subscription‑plan costs rose sharply.

For now, we’re sticking with subscriptions from the big US providers plus cheaper API access as a backup channel, and we’re keeping the local hardware as a deliberate option with a clear condition for revisiting it, and for tasks where what matters isn’t price but the fact that data never leaves the client’s infrastructure. This isn’t a safeguard “for worse times,” but knowledge of our own architecture’s limits knowledge that keeps choosing a provider a genuine choice, not a compulsion.