← All postsOpen models15 July 2026 · 5 min read

The real price of a free model: self hosting maths for a small business

In short

Free weights are the cheapest part of running an open model. The GPU bill, the engineer who babysits it, and the redundancy you discover you need are the real price. Here is the 2026 maths, and the third option that beats both extremes for most small teams.

The break even point moved, but not where you think

In 2026, the crossover where self hosting becomes cheaper than paying for APIs sits at roughly 5 to 10 million tokens per month depending on model tier and your team's infrastructure comfort, per SitePoint's self hosted LLM cost guide. Below that, APIs win decisively. For context, that threshold represents serious sustained usage: a typical early product doing a few thousand Ai interactions a day does not get near it.

The infrastructure line is only the visible half. Production self hosting realistically consumes 20 to 30 percent of a senior engineer's time, which is $3,000 to $6,000 a month in staffing cost before a single GPU hour, per AI Superior's 2026 deployment cost breakdown, with dedicated infrastructure typically adding $1,500 to $5,000 a month on top. A free model with a five figure monthly operating floor is not free. It is an enterprise purchase wearing a hoodie.

The third option: open models over someone else's GPUs

The 2026 twist is that you no longer choose between self hosting an open model and paying frontier API prices. Open models are served by hosted providers at prices the closed labs cannot match: GLM 5.2 costs $1.40 per million input tokens with MIT licensed weights, per devFlokers' June 2026 roundup, and comparison shopping across hosts is a solved problem, per SiliconFlow's 2026 hosting comparison. You get open model economics, zero infrastructure burden, and the freedom to switch hosts because the weights are public. Matt Wolfe's verdict after testing GLM 5.2 lands the point: quality similar to the frontier closed models "at like 1/5 of the cost", per his full GLM 5.2 guide on YouTube.

This mirrors the advice we give on Ai app build costs generally: the cost that kills projects is rarely the headline number, it is the recurring one nobody modelled. Price the engineer, not the download.

Frequently asked questions

Is self hosting an LLM worth it for a small business?

Usually not in 2026. The crossover versus APIs sits around 5 to 10 million tokens a month, and hidden staffing costs run $3,000 to $6,000 monthly. Hosted open models give most of the cost benefit with none of the operations burden.

What is the cheapest way to use open source models?

Hosted inference providers serving open weights. GLM 5.2 at $1.40 per million input tokens is frontier adjacent quality at a fraction of closed model pricing, and public weights mean you can switch providers freely.

When does self hosting genuinely make sense?

High predictable volume (hundreds of millions of tokens monthly), hard data residency requirements, or when inference cost is the core of your unit economics. Then the fixed costs amortise and control pays for itself.

© 2026 Dinimiciuil Labs. All rights reserved. Written on the build floor in Dublin. You are welcome to quote a short excerpt with a link back; please do not republish the full article without permission.

Keep reading
Open modelsShould your product run on an open model? A founder's checklistOpen models hit frontier quality this summer, but that answers the wrong question. Whether YOUR product should run on one comes down to six checks: quality fit, license, cost curve, data rules, supplier risk, and switching cost.Read →Open modelsOpen weight is not open source: the license trap in the model gold rushKimi K3 tops benchmarks under a Modified MIT license, Inkling ships Apache 2.0, and founders keep calling both 'open source'. The differences decide what you can legally build. A plain guide.Read →Open modelsThe month open models caught up: Kimi K3, Inkling and GLM 5.2In four weeks, open weight models went from a few months behind to benchmark parity: Kimi K3 topping arenas, Mira Murati's Inkling under Apache 2.0, GLM 5.2 at a fifth of frontier prices. What actually happened and what it changes.Read →

Build somethingAi nativeagenticroboticreal

We design, build and deploy Ai products end to end, from Dublin.

Start a projectSee how we ship
© 2026 Dinimiciuil Labs. All rights reserved.Dublin · Ireland