Sign in

Fine-Tuning an LLM in 2026: The Same 30-Million-Token Run Bills $12 or $90, and One Platform Google Recommends Has Redirected Its Whole Domain to Another Company

Table comparing the cost of one 30 million token LoRA fine-tuning run on Together AI and Fireworks AI against rented H100 hours

The short answer

If you are a small team weighing whether to fine-tune a model, the decision is usually framed as which library to use. That is the wrong first question, because the libraries are free and the compute is not. The question that moves money is what one run costs, and on 18 September 2026 the answer depends far more on who you buy it from than on what you are training.

Take one concrete job: a LoRA supervised fine-tune on a 10 million token dataset for three epochs. Both managed vendors state on their own pricing pages that they bill dataset size multiplied by epochs, so that is 30 million training tokens on either bill. Here is what it costs.

30M training tokens, LoRA supervised Together AI Fireworks AI Apart by
gpt-oss-20B $12.00 $90.00 7.5x
Llama 3.3 70B $60.90 $90.00 1.48x
Llama 3.3 70B, full parameter $67.20 $180.00 2.68x
Both vendors publish a rate per million training tokens. Read from their own pricing pages on 18 September 2026 and multiplied out. Nothing here is estimated.

The 7.5x on the first row is not a discount and not a quality difference. Together AI prices each model individually. Fireworks prices by parameter band, and gpt-oss-20B lands in the band that starts at 16.1 billion parameters. One tenth of a billion parameters either side of that line is the difference between $0.50 and $3.00 per million training tokens, which is six times.

Three more things on this page that no other result for this query carries. OpenAI has closed its fine-tuning platform to new users and says so in its API documentation, two redirects away from the URL that used to hold it. Predibase, which Google’s own AI Mode recommends by name, no longer resolves: its root and its pricing page both return a permanent redirect to a different company. And the same H100 hour is sold at $1.99 and at $8.00, a factor of 4, which is the number that decides whether renting beats buying the managed job at all.

Who publishes the page Google shows you

Start with the results page itself, captured on 14 September 2026. For a commercial query with a real budget attached, it is unusually thin: no AI Overview and zero People Also Ask questions, while organic, related searches, a discussions block and AI Mode all returned normally in the same request. The capture is sound. Google simply has no consensus answer to summarise here.

What it does return leans commercial. Google’s AI Mode cites three sources, and two of them sell into the decision: one sells an evaluation platform, one sells fine-tuning compute. Of the eight organic results, three are published by companies selling in the category, including one whose own blog ranks above its own product. The discussions block is two posts with the same headline shape, both promising four libraries.

The single genuinely independent source on the page is organic 4, an arXiv survey. It is good on method and it is dated 23 August 2024, with its last posted version on 30 October 2024. That is more than two years before this article. It predates every price below, it predates Unsloth shipping a desktop application, and it predates both of the platform closures in the next section. Read it for the concepts. Do not read it for a vendor.

Two of the managed options Google offers you cannot be bought

Google’s AI Mode closes with a section on managed platforms for teams that do not want to run infrastructure. We opened each one on 18 September 2026.

Predibase redirects to another company

AI Mode describes Predibase as a managed infrastructure platform for efficient LoRA fine-tuning. On 18 September 2026, predibase.com returns HTTP 301 and its pricing page returns HTTP 301, both to a product page on rubrik.com. The documentation subdomain redirects to the same domain’s home page. We tested with two independent clients, a rendered browser and a plain HTTP request, because a single failed fetch is a property of one client and not of a page. Both returned the same server-issued redirect header. There is no Predibase pricing page to read.

OpenAI is winding down fine-tuning

This one takes two redirects to find, which is why it is not in any listicle on this results page. openai.com/api/pricing/ redirects to openai.com/business/pricing/#api, a business pricing page that carries inference rates and no fine-tuning at all. platform.openai.com/docs/pricing redirects to developers.openai.com/api/docs/pricing, and that is where the notice sits. The wording is OpenAI is winding down the fine-tuning platform, followed by a line saying the platform is no longer accessible to new users and that existing users can keep creating jobs for the coming months.

One model is left with a published training price: o4-mini-2025-04-16 at $100.00 an hour, billed by the hour rather than by the token, which is the opposite of how every other managed vendor here bills. There is a second number worth seeing. Inference on a fine-tuned model costs $4.00 per million input tokens and $16.00 per million output tokens, and the same model with data sharing enabled costs $2.00 and $8.00. Sharing your training traffic is worth exactly half the bill.

SiliconFlow sells fine-tuning and publishes no fine-tuning price

SiliconFlow is one of the three sources Google’s AI Mode cites for this query, and it sells fine-tuning compute. Its pricing page carries per-token inference rates only. Its products page has a section headed Pricing with two cards: one links back to that same token page, the other says Contact Sales. Both of the obvious direct URLs, /pricing/finetune and /pricing/gpu, return HTTP 404. There is no fine-tuning rate and no GPU hourly rate anywhere on the site we could reach.

That leaves Together AI as the only one of the managed platforms Google names that you can price today, which is why it carries half the arithmetic below.

What one run costs, converted into the only unit that compares

A managed service quotes tokens. A rented GPU quotes hours. Those do not compare directly and anyone who tells you otherwise is guessing at a training time. So invert it: take the managed price and ask how many GPU hours it buys. That is arithmetic on two published numbers, it needs no benchmark, and it hands the decision back to the only person who knows how long the job runs, which is you.

Managed price for the run Cost Hours it buys at $2.89 an H100
gpt-oss-20B on Together AI $12.00 4.2 hours
Llama 3.3 70B on Together AI $60.90 21.1 hours
Llama 3.3 70B on Fireworks AI $90.00 31.1 hours
Llama 3.3 70B on Fireworks AI, against a Fireworks GPU at $8.00 $90.00 11.25 hours
Managed rates and GPU hourly rates both read on 18 September 2026. The right-hand column is the break-even: if your run finishes faster than that, renting is cheaper, and if it takes longer, the managed job is. We do not publish a training time and we did not estimate one.

The last row is the one to sit with. Fireworks charges $90.00 to run that job for you, and rents you the GPU to do it yourself at $8.00 an hour. Its own managed price is therefore worth 11.25 hours on its own hardware. If a 70 billion parameter LoRA run on 30 million tokens takes you longer than that on one card, and it may well, the managed job is the cheaper of the two things Fireworks sells you.

The minimum charge that nobody on this results page computes

Together AI publishes a minimum charge column beside every rate, and it is doing more work than it looks. Llama 3.1 8B trains at $0.34 per million training tokens with a minimum of $4.00. Divide one by the other: every job under 11.76 million training tokens costs the same $4.00, whether it is one million tokens or eleven. For a small team testing whether fine-tuning helps at all, that is the real price of the first experiment, and it is also cheap enough that the answer is worth buying.

The method multipliers are constants, not costs

Every guide on this query will tell you that LoRA is cheaper than a full fine-tune and that preference optimisation costs more than supervised. True. What none of them says is that the size of those gaps is a number each vendor picked, and the two vendors picked differently.

Choice Together AI charges Fireworks AI charges
Preference optimisation instead of supervised 2.47x to 2.50x, on all 39 rows 2.00x exactly, on all 4 bands
Full parameter instead of LoRA 1.10x to 1.12x 2.00x exactly, on all 4 bands
Derived by dividing one published column by another on each vendor’s own pricing page, read 18 September 2026. Neither multiplier moves with the model, which is what tells you it is a policy rather than a measured cost.

Read the second row again. Going from a LoRA adapter to a full parameter fine-tune doubles the bill on Fireworks and adds about a tenth to it on Together AI. That is the same technical choice priced 1.8 times apart, and for a team that has decided it needs a full fine-tune it is a larger factor than anything in the library comparison they have been reading.

There is a catch on the cheaper side, and it is the kind that only shows up when you operate the page rather than read it. Together AI lists 29 models on its LoRA tab and 10 on its full fine-tuning tab. Switching tabs silently removes 19 models. gpt-oss-120B, DeepSeek, Kimi and GLM all have a LoRA price and no full fine-tuning price at all. Fireworks prints one for every band. So the cheaper multiplier is real and it is only available on the ten models it covers. For the other 19 the comparison does not exist.

Together AI, per 1M training tokens LoRA supervised LoRA preference Full supervised Minimum charge
Llama 3.1 8B $0.34 $0.84 $0.38 $4.00
Qwen3.5 9B $0.34 $0.84 $0.38 $4.00
gpt-oss-20B $0.40 $1.00 not offered $4.00
Qwen3.5 27B $1.05 $2.62 $1.16 $4.00
Llama 3.3 70B $2.03 $5.08 $2.24 $4.00
gpt-oss-120B $2.50 $6.25 not offered $6.00
DeepSeek-V3.1 $7.00 $17.50 not offered $20.00
Kimi K2.6 $15.00 $37.50 not offered $60.00
GLM-5.2 $40.00 $100.00 not offered $60.00
Nine of the 29 models Together AI prices for LoRA training, read on 18 September 2026. Billing basis, in its own words: training dataset size × number of epochs, plus evaluation tokens.
Fireworks AI, per 1M training tokens LoRA supervised LoRA preference Full supervised Full preference
Up to 16B parameters $0.50 $1.00 $1.00 $2.00
16.1B to 80B $3.00 $6.00 $6.00 $12.00
80B to 300B $6.00 $12.00 $12.00 $24.00
Over 300B $10.00 $20.00 $20.00 $40.00
Four parameter bands, not a model list, read on 18 September 2026. The step from the first band to the second is 6 times across one tenth of a billion parameters.

One honest note on direction, because this is not a table where one vendor wins every row. Above 300 billion parameters Fireworks charges $10.00 per million training tokens and names Kimi K2 as an example of that band, while Together AI prices Kimi K2.6 at $15.00. At the very top of the catalogue the cheaper vendor swaps over. The lesson is not that one of these is the cheap one. It is that a per-model rate and a per-band rate cross, and where they cross depends entirely on the model you picked.

The same H100 hour, $1.99 to $8.00

If the break-even table above sent you toward renting, this is the number that decides how good that deal is. Every one of these is an 80 gigabyte class H100 hour, read from the vendor’s own page on 18 September 2026.

Vendor Product Card Per hour As printed
Together AI GPU Clusters, preemptible HGX H100 $1.99 $1.99 per GPU per hour
Runpod Pods H100 PCIe 80GB $2.89 $2.89/hr
Runpod Pods H100 SXM 80GB $3.49 $3.49/hr
Modal GPU Tasks H100 SXM5 $3.95 $0.001097 / sec
Together AI GPU Clusters, on-demand HGX H100 $3.99 $3.99 per GPU per hour
Runpod Serverless H100 80GB $4.79 $4.79/hr
Fireworks AI On-demand deployments H100 80GB $8.00 $0.134 per minute
Runpod Clusters H100 SXM no price Contact sales
SiliconFlow Reserved GPUs not stated no price Contact Sales
Seven priced rows and two that refuse, read 18 September 2026. The Runpod page carries its own stamp: Updated September 13, 2026.

2.77 times between the cheapest on-demand hour and the dearest, and 4 times once you allow a preemptible one. That is the same spread we measured in the serving lane one day earlier, in the LLM serving article, on a different set of vendors. It is not a quirk of one pricing page.

Three things the hourly number does not tell you

One vendor prints the second, not the hour. Modal lists its H100 at $0.001097 per second. Nothing is hidden, the arithmetic closes, and the hour is $3.95. Modal confirms it on the same page, in its own savings comparison, where it prices itself at $3.95 per GPU hour. That comparison is worth a second look for this use case: it beats a $3.00 baseline only by assuming you would have provisioned 75 GPUs to Modal’s average of 50. A training run has no idle capacity to reclaim. Take the idle saving away and what is left is 1.37 times the cheapest hour in the table.

One vendor advertises an option it does not sell. Modal’s plan comparison lists Non-preemptible execution, 3x base prices under all three plans. Modal’s own documentation says, in as many words: The nonpreemptible parameter is not supported for GPU Functions. This matters more for training than for inference. A preempted request is retried; a preempted training run loses work, which is why Modal’s preemption guide tells you to checkpoint. Region pinning is a separate multiplier, 1.15x to 1.75x, which takes that $3.95 hour to $6.91.

One vendor charges the same for two different cards. Fireworks prices an H100 with 80 gigabytes and an H200 with 141 at $8.00 an hour each. Its region-restricted premium of 1.5x takes that to $12.00. And its two columns do not quite reconcile: $0.134 a minute is $8.04 an hour beside a printed $8.00. That is rounding rather than concealment, and it is worth knowing only because a buyer who budgets from the per-minute column overstates the bill.

Two rows in that table have no price. Runpod sells an H100 SXM by the hour under Pods and answers Contact sales for the same card under Clusters, which is where a multi-GPU training job would actually run. SiliconFlow answers Contact Sales for reserved GPUs. Both are the withheld-numerator shape this desk keeps finding: the specification is public and the rate is not.

The speed claim Google repeats is not the one Unsloth prints

Google’s AI Mode ranks Unsloth first for this query and summarises it with a specific pair of numbers: up to 5x faster with 80% less memory. That is a vendor claim, so we went to the vendor. On 18 September 2026, Unsloth’s own README says something different: 2× faster with 70% less VRAM.

Unsloth notebook Speed, as Unsloth states it Memory saved, as Unsloth states it
Gemma 4 (E2B) 1.5x faster 50% less
Qwen3.5 (4B) 1.5x faster 60% less
gpt-oss (20B) 2x faster 70% less
gpt-oss (20B): GRPO 2x faster 80% less
Llama 3.1 (8B) Alpaca 2x faster 70% less
embeddinggemma (300M) 2x faster 20% less
Six rows from the table in Unsloth’s own README, read on 18 September 2026. These are the project’s own figures for its own notebooks and we did not measure them.

Not one row in Unsloth’s own table says five times. The speed column runs 1.5x to 2x and the memory column runs 20 percent to 80 percent depending on the model, with an embedding model at the bottom of that range. Google took the top of one range and the top of another, dropped the words that qualify them, and handed a small team a planning number that the vendor does not claim for the model they are probably training. We are not calling this a false claim by Unsloth. Unsloth wrote 2x and 70 percent. We are saying the number reaching the buyer is not that one.

There is a second change worth recording, because it affects what you are choosing. Unsloth’s README headline on 18 September 2026 reads Unsloth is the first desktop app to run and train models, and its repository description now leads with a local application for running and training models. Google still describes it purely as a library for consumer GPUs and free notebooks, which it still is. But the project’s own front door has moved, and a team arriving from that summary should expect a different landing page than the one being described.

The libraries themselves are the cleanest part of this market

After all of the above, the open source side is reassuring. We read all five repositories through GitHub’s REST API on 18 September 2026, not from the repository page, because the API returns the archive flag.

Project Licence Stars Last push Open issues
Unsloth Apache-2.0 76,362 18 September 2026 1,241
LLaMA-Factory Apache-2.0 74,854 14 September 2026 1,156
DeepSpeed Apache-2.0 43,132 18 September 2026 1,436
Hugging Face PEFT Apache-2.0 21,694 18 September 2026 80
Axolotl Apache-2.0 12,486 18 September 2026 239
All five rows from the GitHub REST API on 18 September 2026. Four of the five were pushed the same day we read them.

All five are plain Apache-2.0 and none is archived. That is worth stating because it is not what we found one lane over. In the RAG frameworks article the most starred project on the results page carried a modified licence requiring a commercial agreement, and in the serving article the fourth-ranked project turned out to be archived and read-only. The training layer is the clean floor of this stack and the application layer above it is not.

One housekeeping item, since a rename breaks links. LLaMA-Factory now lives at hiyouga/LlamaFactory and the old path redirects. It is also the only one of the five that had not been pushed on the day we read it, by four days, which is well inside normal for a project of that size and is recorded here only so the table is not misread.

A unit switch worth catching before you sign for tracking

Experiment tracking is the one paid tool almost every fine-tuning guide recommends and almost none of them prices. Weights & Biases costs nothing on its free tier, which carries 5 model seats and 5 gigabytes of storage a month. Pro starts at $60.00 a month for up to 10 seats and 100 gigabytes.

Two rows of the same comparison table carry the overage rates, and they are in different units. Additional storage is $0.03 per gigabyte. Additional data ingestion is $0.10 per megabyte. At a thousand megabytes to the gigabyte, that second rate is $100.00 a gigabyte sitting two rows below $0.03 a gigabyte. They meter different things, so this is not a contradiction. It is a unit switch inside one table, and a team that budgets tracking overage off the storage row will be out by more than three thousand times.

The Pro plan also carries a condition rather than a price: For early-stage teams fewer than 50 employees. Above that headcount the published number stops applying and the plan becomes a conversation.

What we cut, and why

  • How long a fine-tuning run takes. Every managed price above could have been turned into a dollar comparison by assuming a training time. We do not have a measured one for these models on these cards, so the break-even is published in hours and the hours are yours to fill in.
  • Any quality comparison between libraries. Whether a model fine-tuned with one of these is better than the same model tuned with another is a benchmark we have not run, and no page on this results page has run it either.
  • The street price of a GPU you buy outright. Cut for the reason recorded in the self-hosting break-even article: no single authoritative page carries one.
  • Predibase pricing. Not withheld, not estimated. There is no page to read.
  • SiliconFlow fine-tuning and GPU rates. Same reason. We looked at the pricing page, the products page and two direct URLs.

How we tested

Every price in this article was opened in a rendered browser on 18 September 2026 and read from the painted page, never from raw HTML, because a hidden billing state is often still in the document. Tabs and toggles were operated rather than assumed, which is how the full fine-tuning catalogue turned out to be 10 models rather than 29. Repository rows come from GitHub’s REST API on the same date. Where a redirect mattered, it was confirmed from two independent clients. Anything we could not verify was cut rather than hedged. More on the method is in our open source pillar.

Questions this results page left unanswered

What does one LLM fine-tuning run actually cost?

On published rates read 18 September 2026, a LoRA supervised fine-tune of Llama 3.3 70B on 30 million training tokens costs $60.90 on Together AI and $90.00 on Fireworks AI. The same job on gpt-oss-20B costs $12.00 and $90.00. Both vendors bill dataset size multiplied by the number of epochs, so a three epoch run on ten million tokens is billed as thirty million.

Is it cheaper to rent a GPU or use a managed fine-tuning service?

It depends on how long your run takes, and you can answer it without a benchmark. Divide the managed price by the GPU hourly rate to get a break-even in hours. Together AI’s $60.90 for the Llama 3.3 70B run buys 21.1 hours of an H100 at $2.89. Finish faster than that and renting wins. Take longer and the managed job does.

Do I need an expensive GPU to fine-tune a model?

Not to start. Rented consumer cards on Runpod run from $0.27 an hour for an RTX A5000 to $0.74 for an RTX 4090, read 18 September 2026, and LoRA and QLoRA exist precisely so a small model fits on one. The libraries themselves cost nothing: all five in the table above are Apache-2.0.

Can I still fine-tune a model with OpenAI?

Not as a new customer. OpenAI’s API documentation, read 18 September 2026, states that it is winding the fine-tuning platform down and that it is no longer accessible to new users, with existing users able to create jobs for the coming months. One model keeps a published training price, o4-mini-2025-04-16 at $100.00 an hour. The notice is not on the public pricing page, which carries no fine-tuning section at all.

Unsloth or Axolotl?

Both are Apache-2.0 and both were pushed on 18 September 2026. Unsloth is the consumer GPU and free notebook route and its own README claims 2× faster with 70% less VRAM rather than the larger figure Google repeats. Axolotl is the YAML configured, reproducible, multi-GPU route. The choice is about how you want to define a run, not about cost: neither charges anything, and the compute bill is identical either way.

Why does the same model cost 7.5 times more to fine-tune on one platform?

Because the two platforms meter differently. Together AI publishes a rate per model. Fireworks AI publishes four parameter bands, so gpt-oss-20B falls into its 16.1B to 80B band at $3.00 per million training tokens while Together prices it at $0.40. The band edge itself is a 6 times step across one tenth of a billion parameters.

What does full fine-tuning cost compared with LoRA?

Two different answers from two vendors, read 18 September 2026. Fireworks AI charges exactly twice the LoRA rate for a full parameter fine-tune on all four of its bands. Together AI charges 1.10x to 1.12x, but only offers full fine-tuning on 10 of the 29 models it prices for LoRA. Neither multiplier varies with the model.

How much should I budget for experiment tracking?

Weights & Biases is free for 5 model seats with 5 gigabytes of storage a month, and Pro starts at $60.00 a month, read 18 September 2026. Watch the overage units: extra storage is $0.03 a gigabyte and extra trace ingestion is $0.10 a megabyte, which is $100.00 a gigabyte. Pro is also limited to teams below fifty employees.

Sources

Every figure above traces to one of these, each opened on the date shown.

  1. Runpod GPU cloud pricing, www.runpod.io, read 18 September 2026.
  2. Modal pricing, modal.com, read 18 September 2026.
  3. Modal preemption documentation, modal.com, read 18 September 2026.
  4. Together AI pricing, fine-tuning and GPU clusters, www.together.ai, read 18 September 2026.
  5. Fireworks AI pricing, fireworks.ai, read 18 September 2026.
  6. OpenAI API pricing documentation, developers.openai.com, read 18 September 2026.
  7. SiliconFlow pricing, www.siliconflow.com, read 18 September 2026.
  8. SiliconFlow products, www.siliconflow.com, read 18 September 2026.
  9. Weights & Biases pricing, wandb.ai, read 18 September 2026.
  10. Unsloth README on GitHub, github.com, read 18 September 2026.
  11. GitHub REST API, five repositories, docs.github.com, read 18 September 2026.
  12. Fine-tuning survey, arXiv 2408.13296, arxiv.org, read 18 September 2026.

Comments

Sign in to join the discussion.

Login to comment