Sign in

Best Open Source LLM Serving Stacks in 2026: Hugging Face Archived TGI and Named Its Replacements, and One Identical Model Bills $0.17 or $0.95 per Million Output Tokens Depending on Who Runs It

Open source LLM serving stacks: one model, 24 sellers, bash.17 to bash.95 per million output tokens

The short answer

If you are choosing what to run behind an inference endpoint, two facts decide most of it and neither is on the results page Google shows you. The first is that Hugging Face has archived text-generation-inference, the stack Google’s own AI Mode still ranks fourth for this query, and its README now points at vLLM and SGLang instead. The second is that the price of running an open-weight model has almost nothing to do with the model. We priced gpt-oss-120b, which both OpenRouter and Fireworks describe as fitting on a single H100, across every provider that publishes a rate for it on 17 September 2026.

The same weights, behind the same API, bill $0.03 to $0.35 per million input tokens and $0.17 to $0.95 per million output tokens. That is 11.7 times on input and 5.6 times on output. And the spread is not a quality ladder: at the single most common price on that list, $0.15 in and $0.60 out, OpenRouter’s own measured throughput runs from 15 tokens a second to 259.

Three providers, one model, one price Per 1M in / out Throughput, P50
Groq $0.15 / $0.60 259 tok/s
Together AI $0.15 / $0.60 24 tok/s
SiliconFlow $0.15 / $0.60 15 tok/s
Rates and P50 throughput for gpt-oss-120b, read from OpenRouter’s provider table on 17 September 2026. Three providers charging the same rate to the cent, 17 times apart on the speed OpenRouter itself measures. We did not run this benchmark and we do not publish one.

Four things on this page that no other page ranking for this query carries: that TGI’s repository is archived and read-only as of 17 September 2026; that the same H100 hour is sold at $2.89 and at $8.00; that one of those vendors publishes a multiplier in its documentation that its pricing page does not mention; and that Groq has taken its pricing page down entirely and now answers “Contact Sales” for both Llama models.

Which of these projects is still shipping code

This is the check to run before any other one, and it takes a minute. We read all six repositories through GitHub’s REST API on 17 September 2026, not from the repository page, because the API returns the archive flag and the page mostly whispers it.

Project Licence badge Stars Last push State
vLLM Apache-2.0 92,012 17 September 2026 Active
Ollama MIT 181,243 17 September 2026 Active
SGLang Apache-2.0 36,100 17 September 2026 Active
NVIDIA TensorRT-LLM Other 14,637 17 September 2026 Active
BentoML OpenLLM Apache-2.0 12,535 14 September 2026 Active
Hugging Face TGI Apache-2.0 10,885 21 March 2026 Archived, read-only
All six rows read from the GitHub REST API on 17 September 2026. Four of the six were pushed the same day we read them.

Five of the six were pushed inside four days of each other. The sixth has not been pushed since 21 March 2026 and is now archived, which means nobody can open a pull request against it. Its last tagged release, v3.3.7, went out on 19 December 2025. It carries 324 open issues and 1,286 forks that are now frozen where they stand.

What Hugging Face says about its own serving stack

The README opens with a caution banner. The wording is text-generation-inference is now in maintenance mode, and the paragraph under it names the engines Hugging Face contributes to and recommends going forward: vLLM, SGLang, and local engines such as llama.cpp and MLX. That is the maintainer of the fourth-ranked answer on this results page telling you to use the first and second instead.

Google’s AI Mode, captured on 14 September 2026, ranks TGI fourth of five and describes it as Best for Native Hugging Face Ecosystem Integration. That was a fair description of TGI for about two years. It is not a description of an archived repository, and the gap between the two is the reason this page exists.

One licence badge on this list is not what it looks like

NVIDIA TensorRT-LLM, which Google’s AI Mode ranks third for this query, shows Other on GitHub and NOASSERTION through the API. On the RAG side of this catalogue that badge meant a genuinely restricted licence. Here it does not. The LICENSE file opens This project is licensed under the Apache 2.0 license and then runs 791 lines listing the licences of code it vendors in, from Bitcoin Core to CUTLASS to Diffusers. GitHub cannot classify a file that long and that mixed, so it gives up and prints Other.

This corrects our own earlier work, and it is recorded here rather than quietly fixed. Our 14 September 2026 capture of this roster concluded that there was no unclassifiable licence badge anywhere in it. That was true of the five projects we had listed and false of the roster Google actually returns, because TensorRT-LLM was not in our five. The rule this sets is the same one the badge itself teaches: read the file, not the badge, and read the whole roster, not the part you already had. The method is set out in our open source pillar, and the opposite case, a permissive badge over a restricted licence, is in the RAG frameworks article.

One model, twenty-four sellers, and an eleven-fold price spread

Here is the part that decides a budget. gpt-oss-120b is 116B mixture of experts, 5.1B active per forward pass, and both OpenRouter’s model page and Fireworks’ own model page state it fits on a single H100. It is the same weights everywhere. Nobody is selling you a different model. OpenRouter lists 24 provider rows serving it, and this is what they charge on 17 September 2026.

Provider Input / 1M Output / 1M Throughput P50 Latency P50 Uptime
CoreWeave $0.03 $0.17 45 tok/s 0.40 s 99.82%
AkashML $0.03 $0.17 20 tok/s 0.93 s 99.59%
DekaLLM $0.03 $0.18 53 tok/s 0.64 s 99.86%
DeepInfra (bf16) $0.037 $0.17 23 tok/s 4.73 s 99.35%
Crusoe $0.05 $0.25 196 tok/s 0.16 s 99.25%
Baseten $0.10 $0.50 164 tok/s 0.31 s 99.99%
SambaNova $0.14 $0.95 256 tok/s 0.51 s 97.15%
Groq $0.15 $0.60 259 tok/s 0.44 s 99.99%
Together AI $0.15 $0.60 24 tok/s 0.40 s 99.25%
SiliconFlow $0.15 $0.60 15 tok/s 1.44 s 97.22%
Amazon Bedrock $0.15 $0.60 not measured not measured 100.00%
DeepInfra (fp8) $0.20 $0.95 106 tok/s 1.85 s 73.06%
Cerebras $0.35 $0.75 794 tok/s 0.22 s 100.00%
Thirteen of the 24 provider rows OpenRouter lists for gpt-oss-120b, read on 17 September 2026. Throughput, latency and uptime are OpenRouter’s own measurements of its own traffic, not ours.

Read the first column and the fourth together and the usual story falls apart. The cheapest input rate on the list, $0.03, belongs to CoreWeave at 45 tokens a second. Together AI charges 5 times that and OpenRouter measures it at 24. The one place price and speed line up is the top of the list: Cerebras charges the most per input token and is measured at 794 tokens a second, which is a real product with a real reason to cost more.

The most common price on the list is not a market rate

8 of the 24 rows print exactly $0.15 in and $0.60 out. Amazon Bedrock, Nebius, DeepInfra Turbo, Phala, Together AI, Groq and SiliconFlow all land on the same two numbers to the cent, and Fireworks AI, which OpenRouter does not list for this model, publishes the same pair on its own page. That is a price point, not a market. What it buys differs enormously: on OpenRouter’s own measurements those identical rates return 259 tokens a second at the top and 15 at the bottom.

The headline rate is the cheapest seller, not the going rate

OpenRouter prints $0.03 / $0.17 per 1M at the top of the model page. Further down the same page it publishes what its customers actually paid: $0.09729 input and $0.3943 output, weighted across providers. The real average is 3.2 times the headline on input and 2.3 times on output. To OpenRouter’s credit, it is the only page in this entire research pass that publishes the gap between its own advertised number and its own billed number. Every other vendor here leaves you to find it.

If you serve it yourself: the same H100, at 2.77 times the price

vLLM, SGLang and TGI are free. The GPU is not, and the GPU is the bill. Every vendor below rents the same class of card. We opened all six pricing pages in a rendered browser on 17 September 2026 and operated every billing control on them, because on three of the six the number you see first is not the number you pay.

Where the H100 comes from Per GPU hour What the page does not lead with
Runpod, H100 PCIe 80 GB, Pods $2.89 Community Cloud and Secure Cloud tabs print the same rate
Runpod, H100 SXM 80 GB, Pods $3.49 same rate on both cloud tabs
Modal, H100 SXM5 $3.95 page defaults to $0.001097 per second; every GPU Function is preemptible
Together AI, HGX H100, GPU Clusters on demand $3.99 $1.99 preemptible on the same table
Runpod, H100 80 GB, Serverless $4.79 flex worker rate
Together AI, HGX H100, Dedicated Inference $5.49 promotion valid until 30 September 2026
Baseten, H100 80 GiB, Dedicated Deployments $6.50 page defaults to $0.10833 per MINUTE
Fireworks AI, H100 80 GB, On Demand $8.00 region-restricted deployments priced at 1.5x
Read from each vendor’s own pricing page in a rendered browser at a 1440 pixel viewport on 17 September 2026, after operating every toggle on it.

Three pricing pages that show you a different number first

Baseten defaults to a per-minute view. Its H100 row prints $0.10833, which reads like eleven cents. Click Hour, which is the control next to it, and the same row prints $6.50. The arithmetic closes exactly, so nothing is hidden, but the first number a buyer sees is 1/60th of the one they will be charged.

Together AI prints two prices for the same card on the same page. Its Dedicated Inference table lists the HGX H100 at $5.49 with a promotional $3.99 marked valid until 30 September 2026. Its GPU Clusters table, further down the same page, lists the same HGX H100 at $3.99 on demand with no promotion and $1.99 preemptible. So the promotion brings the dedicated price down to what the other table charges anyway, and on 1 October the dedicated line goes back to $5.49, 27.3 percent higher.

Runpod ships a toggle that changes nothing. Its pricing page offers Community Cloud and Secure Cloud as two tabs. We read all 21 GPU rows in both states on 17 September 2026. Every row printed the same rate in both. The control visibly switches and moves no number.

The multiplier that is in the docs and not on the pricing page

Modal prices its H100 SXM5 at $0.001097 a second, which is $3.95 an hour and the second cheapest figure in the table above. Its plan comparison table, near the bottom of the same page, carries a row reading Non-preemptible execution at three times base prices, listed identically under all three plans.

Modal’s own documentation says something different. The 3x multiplier applies to the list price for CPU and Memory usage, and the next line reads The nonpreemptible parameter is not supported for GPU Functions. You cannot buy non-preemption for a GPU function on Modal at any price. Every GPU function there is preemptible and will be interrupted and restarted. That is a reasonable design for batch work and it is a decision you want to make deliberately when the thing being interrupted is a live inference endpoint.

Two vendors here also charge extra to run in a region you pick, which is not optional for anyone with a data residency obligation. Modal multiplies the whole bill by 1.15x for a broad region such as us and 1.75x for a narrow one such as us-west, so the same H100 hour becomes $6.91. Fireworks prices a region-restricted deployment at 1.5x, so its $8.00 becomes $12.00.

The only break-even we will do on published numbers

We will not tell you how many tokens a second your stack will produce, because we have not measured it and neither has anyone else on this results page in a way you could reproduce. What can be settled with arithmetic is the other half: what a month of that H100 costs against what the same month of tokens costs from someone else.

At $2.89 an hour, the cheapest H100 in the table above, a month of 730 hours is $2,109.70. That money buys, at published rates for the same model:

  • 12.41 billion output tokens at $0.17 per million, the cheapest published rate
  • 3.52 billion at $0.60, the rate 8 of the 24 rows charge
  • 2.22 billion at $0.95, the dearest

Rent the dearest H100 instead, $8.00 an hour or $5,840.00 a month, and the same three rates put break-even at 34.35 billion, 9.73 billion and 6.15 billion. The range across those six combinations is 15.5 times, and every one of the twelve inputs is a published number. Which end of it you land on is a procurement decision, not an engineering one. The compute side of the same question, including what it costs to keep a card busy, is worked through in our self-hosting break-even article.

Who writes the answer Google shows you

On 14 September 2026, this query returned no AI Overview at all. The capture is sound rather than broken: organic results, People Also Ask, related searches and AI Mode all returned normally in the same request, and the AI Overview expansion came back Google hasn’t returned any results for this query. The discussions and forums module returned nothing either.

AI Mode did answer, and cited two sources: Yotta Labs and Spheron. Both sell GPU compute. On the organic side, 3 of the 9 results are published by 2 companies that sell a serving product. That is not an accusation of dishonesty. It is a note about what a page is for: a company selling inference capacity has a reason to publish a benchmark and no reason to publish the archive date of a competitor’s repository.

Two vendors moved the price off the pricing page

Groq has removed its pricing page. On 17 September 2026, groq.com/pricing returns HTTP 200 and serves the homepage, console.groq.com/pricing returns HTTP 404, and groq.com/sitemap.xml lists no pricing URL at all. The rates do exist, one subdomain away, in the supported models table in the developer documentation. We found them because a finding from the day before said to look: a vendor publishing no price on its pricing page is a claim about a page, and it has to be checked against every surface that page links to.

Fireworks AI does the same thing more politely. The Serverless Inference block on its pricing page prints embedding rates and then points to documentation for everything else, so the gpt-oss-120b rate lives on that model’s own page rather than on the page called Pricing.

What Groq stopped publishing

Model Model ID Speed Groq publishes Price Groq publishes
Llama 3.1 8B llama-3.1-8b-instant 560 tok/s Contact Sales
Llama 3.3 70B llama-3.3-70b-versatile 280 tok/s Contact Sales
GPT OSS 120B 500 tok/s $0.15 in, $0.60 out per 1M
GPT OSS 20B 1000 tok/s $0.075 in, $0.30 out per 1M
Read from Groq’s supported models table on 17 September 2026. The two Llama rows carry an Enterprise tag and return Contact Sales in both the price column and the rate limit column. The two gpt-oss rows keep published rates.

Groq still publishes a throughput figure for both Llama models, 560 and 280 tokens a second, so the models are running. It is only the price that has gone. If your plan was to serve Llama 3.3 70B through Groq at a known rate, that rate is now a sales conversation.

What each of these stacks is actually for

This grouping is Google’s AI Mode answer from 14 September 2026, reported because it is what a reader searching this query is shown, and attributed for the same reason. It is not our recommendation and it does not survive the archive finding above intact.

Rank in AI Mode Stack What AI Mode says it is for What we found
1 vLLM general production APIs, multi-model hosting, mixed hardware fleets Apache-2.0, 92,012 stars, pushed 17 September 2026
2 SGLang agents, multi-turn chat, structured and constrained generation Apache-2.0, 36,100 stars, pushed 17 September 2026, and one of the two engines Hugging Face now points at
3 NVIDIA TensorRT-LLM maximum throughput on static all-NVIDIA clusters badge reads Other, file reads Apache 2.0, 14,637 stars, pushed 17 September 2026
4 Hugging Face TGI native Hugging Face ecosystem integration archived and read-only, last push 21 March 2026
5 Ollama local development, workstations and edge, not web-scale multi-tenant MIT, 181,243 stars, pushed 17 September 2026
Left two columns quoted from Google AI Mode, captured 14 September 2026. Right column read from the GitHub REST API on 17 September 2026.

What this page does not tell you

  • Tokens per second, time to first token, and every throughput or latency comparison between any two of these stacks. We have not benchmarked vLLM against SGLang against TensorRT-LLM on identical hardware with an identical model, and we will not borrow a benchmark from a company that rents GPUs. The per-provider throughput figures above are OpenRouter measuring the traffic it routes, which is a different claim, and it is labelled as theirs.
  • Any total cost of ownership for a self-hosted stack. Engineer hours, the cost of an on-call rotation and the utilisation you will actually hit are real and no vendor publishes an authoritative page for any of them.
  • Groq rates for Llama 3.1 8B and Llama 3.3 70B. They are Contact Sales, so there is nothing to report.
  • Reserved and committed-spend rates. Together, Runpod and Fireworks all quote lower numbers against a term commitment and most of those tiers read Contact us, so none of them can be placed beside an on-demand rate on equal footing.
  • Whether the price spread reflects quantization. Several providers on that list serve different precisions of the same weights, OpenRouter labels only some of them, and we did not test outputs. It is a plausible part of the explanation and we cannot size it.
  • One published rate we read and will not use. Runpod’s public endpoints table, on a page stamped September 13, 2026, prints $0.00001 per 1m tokens for deep-cogito / Deep Cogito v2 Llama 70B while pricing qwen / Qwen3 32B AWQ at $10.00 per 1m tokens three rows away. One of those two is wrong by several orders of magnitude and we cannot tell which, so neither is in any table here.

How we checked this

Every price on this page was read on 17 September 2026 from the vendor’s own page in a rendered browser at a 1440 pixel viewport, US locale, from a settled frame after every billing control on the page had been operated. That is not caution for its own sake: on this pass alone, one page defaulted to a per-minute view, one carried a promotional price with an expiry date next to a permanent price for the same card, and one shipped a toggle that moved no number.

Repository rows come from GitHub’s REST API rather than the repository page, because the API returns the archive flag as a field. Licence text was read from the file on the named branch, not from the badge. Where a vendor’s documentation and its pricing page disagreed, as they did on Modal, the documentation is quoted and the disagreement is printed rather than resolved silently. Anything that could not be verified was cut, and the cuts are listed above. The same method and the licence-badge check it introduced are set out in our open source pillar.

Questions people actually ask about LLM serving frameworks

Which is the best open source LLM serving framework in 2026?

Google’s AI Mode, captured 14 September 2026, ranks vLLM first for general production APIs and SGLang second for agents and structured output. What we can verify is narrower and more useful: on 17 September 2026, vLLM carries 92,012 stars under Apache-2.0 and SGLang 36,100, both were pushed the day we read them, and Hugging Face’s own README names those two as the engines it recommends now that it has archived its own. We have not benchmarked them against each other and do not publish a speed ranking.

Is Hugging Face TGI dead?

The repository is archived and read-only as of 17 September 2026, which means no new pull requests and no new issues. Its last push was 21 March 2026 and its last tagged release, v3.3.7, was 19 December 2025. The README says text-generation-inference is now in maintenance mode and points at vLLM and SGLang. Existing containers keep running; the project is not accepting work.

What does it cost to run gpt-oss-120b?

Between $0.03 and $0.35 per million input tokens and between $0.17 and $0.95 per million output tokens, across the 24 provider rows OpenRouter lists on 17 September 2026. The most common single price is $0.15 in and $0.60 out, charged by 8 of them. OpenRouter’s own weighted average of what its customers actually paid is $0.09729 in and $0.3943 out.

Is it cheaper to self-host or to use an inference API?

On published numbers, a month of the cheapest H100 in this article, $2.89 an hour for 730 hours, is $2,109.70. That equals 12.41 billion output tokens at the lowest published rate for the same model, or 2.22 billion at the highest. If your product will not produce billions of output tokens a month, the arithmetic favours the API before any engineer time is counted. What this page cannot tell you is how close to that ceiling your own stack would get, because we have not measured it.

Why does the same open source model cost different amounts?

Because you are not buying the model, you are buying someone’s GPUs, their utilisation and their margin. The weights are identical and free. On 17 September 2026 the spread across OpenRouter’s provider list for one model was 11.7 times on input tokens. The spread does not track speed either: at the single rate of $0.15 in and $0.60 out, OpenRouter measures 15 tokens a second at the bottom and 259 at the top.

How much does an H100 cost per hour?

$2.89 to $8.00 an hour on demand across the six vendors in this article, all read on 17 September 2026. Runpod is lowest at $2.89 for an H100 PCIe, Modal is $3.95, Together AI $3.99 on its GPU cluster table, Baseten $6.50 and Fireworks $8.00. Two of those rise further if you pin a region: Modal by 1.15x or 1.75x, Fireworks by 1.5x.

Does Groq publish its prices?

Not on a pricing page. As of 17 September 2026, groq.com/pricing returns HTTP 200 and serves the homepage, console.groq.com/pricing is a 404, and the sitemap lists neither. The per-token rates are published in the supported models table in the developer documentation. Two models have no public rate at all: Llama 3.1 8B and Llama 3.3 70B both read Contact Sales in the price column.

Can I run a serving stack on a preemptible GPU?

Yes, and on one of these vendors you have no choice. Modal’s documentation states that its non-preemptible parameter is not supported for GPU Functions, so every GPU function there can be interrupted and restarted. Together AI sells preemptible H100 capacity at $1.99 an hour against $3.99 on demand, which is the clearest published price for what interruptibility is worth: about half.

Sources

  1. Together AI pricing, serverless rates and GPU cluster rates, read 17 September 2026.
  2. Groq supported models and per-token rates, read 17 September 2026.
  3. Groq sitemap, which lists no pricing URL, read 17 September 2026.
  4. Baseten pricing, model APIs and dedicated deployments, read 17 September 2026.
  5. Fireworks AI pricing, on-demand GPU rates, read 17 September 2026.
  6. Fireworks AI gpt-oss-120b model page, read 17 September 2026.
  7. Runpod GPU cloud pricing, updated 13 September 2026, read 17 September 2026.
  8. Modal pricing, per-second GPU rates and plan table, read 17 September 2026.
  9. Modal docs, region selection multipliers, read 17 September 2026.
  10. Modal docs, preemption and the GPU exclusion, read 17 September 2026.
  11. OpenRouter gpt-oss-120b, provider rates, measured throughput and weighted average paid, read 17 September 2026.
  12. Hugging Face text-generation-inference, archived repository and maintenance-mode README, read 17 September 2026.
  13. NVIDIA TensorRT-LLM LICENSE, main branch, read 17 September 2026.
  14. GitHub REST API, repository records for all six projects, read 17 September 2026.

Comments

Sign in to join the discussion.

Login to comment