Hugging Face’s guide says you can fine-tune a 33-billion-parameter model on one 24-gigabyte graphics card [1]. That fact changes a choice many teams get wrong: whether to build a local model or keep paying for tokens. The true test is not the hardware specs, but the price you compare it against.
The big picture:
A local model and a rented one use different pricing units. Those units decide the answer.
Rented models sell by the token. On one October 2026 afternoon, a cheap rented model cost 30 cents per million input tokens [3]. A top-tier model cost $10 per million [2].
That is a 33-times gap for the exact same work. What strikes me is how rarely this huge gap shows up in team budgets.
On the local side, you buy the card once and pay for power. A card drawing 575 watts needs a 1,000-watt power supply [5]. The US average retail price for power was 12.68 cents per kilowatthour in the latest yearly data [4].
Run that card flat out for a year, and the power costs about $639, or near $53 a month. That is basic math on published specs and public prices, not a vendor guess.
Teams often skip the memory limits. Hugging Face uses a memory-saving method called QLoRA, which keeps a compressed copy of the model while it trains. That compressed path fits a 33-billion-parameter model on 24 gigabytes, and a 65-billion-parameter model on 46 gigabytes [1].
By the numbers
- 33 billion — Fine-tuning floor: The compressed QLoRA method fine-tunes a 33-billion-parameter model on one 24-gigabyte card, and a 65-billion-parameter model on a 46-gigabyte card [1].
- 12.68 cents — Power price: The US average retail power price per kilowatthour [4]. A 575-watt card run all year uses about 5,037 kilowatthours, costing near $639 [5][4].
- $0.30 vs. $10 — Rented prices: A MiniMax M3 model costs 30 cents per million input tokens [3], while a top-tier model costs $10 per million [2].
- 532 million vs. 13 million — Token yield: A $639 power bill buys about 532 million output tokens at $1.20 per million [3], or about 13 million at $50 per million [2] — a 42-fold swing.
What I’d watch:
Smart teams price their builds against a cheap rented model, not the top-tier one. The part I keep circling is how the rig’s power bill becomes the smallest number in the room.
- The math test: A payback sheet using a $10-per-million rate will always say “build local.” That same sheet at 30 cents says “keep renting” [2][3].
- The token mix: Cheap rented output tokens cost four times their input tokens [3]. A job that writes more than it reads changes the math faster than any hardware choice.
- The memory floor: A 33-billion-parameter model needs 24 gigabytes to fine-tune. The 65-billion-parameter floor needs about 46 gigabytes, which most single consumer cards lack [1].
- The power line: A 575-watt card needs a 1,000-watt power supply [5]. This means “free” local compute hits a hard power limit before it ever prints a bill.
The biggest lever is the rented price you chose not to beat.
The catch
Power is just the running cost, not the full bill. The graphics card, the server box, and the engineer’s time sit outside that math.
My read: 24 gigabytes is a hard floor for fine-tuning. It only works for a compressed path that keeps most of the model’s quality, and it cannot fix bad data [1]. A team with no labeled examples has nothing to train, no matter how cheap the power gets.
The pricing side is just as tricky. The two rented rates above are vendor prices from October 2026, and vendors change rates often [2][3]. The power figure is a national average, so a factory rate or a European rate tells a different story [4].
A field note shows the real shape of a win. One legal-tech team cut $14,200 a month in rented costs down to a $180 power bill by fine-tuning a small local model [6]. That win came from a narrow task and a huge pile of labeled text, not a promise that every job works the same way.
At a glance
- The Big Shift: Teams often treat local model training as a hardware choice, but the real deciding factor is the rented token price you compare it against. That price showed a 33-times gap on the exact same day [2][3].
- Why It Matters: A local rig’s running cost is just power, which runs about $639 a year for a 575-watt card at the national average rate [4][5]. Your payback math swings 42-fold depending on which rented model sits in your comparison.
- What I’d Watch: Whether teams price their local builds against a cheap rented model instead of a top-tier one.
- The math test: The rented rate used in the math, where a high rate proves “build local” and a cheap rate proves “keep renting” [2][3].
- The token mix: How much of the job is output rather than input, since cheap rented output tokens cost four times more [3].
- The memory floor: The video memory a fine-tune needs. This takes about 24 gigabytes for a 33-billion-parameter model and 46 gigabytes for a 65-billion-parameter one [1].
- The Catch: Power costs ignore the card’s upfront price and the labeled data a fine-tune needs. Also, both rented rates are vendor prices that can change without warning [1][2][3].
Related reading
- Self-Hosted AI: When to Buy vs Rent GPUs — more on Major Purchases & Assets
- Ebike Costs: Per-Use Calculator for Commuters — more on Major Purchases & Assets
- Is a $3,000 Espresso Maker Machine Worth It? — more on Major Purchases & Assets
Sources
[1] Hugging Face, “Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA” — https://huggingface.co/blog/4bit-transformers-bitsandbytes [2] Anthropic, “Claude API pricing” (top-tier model rate card, retrieved 2026-10-06) — https://www.anthropic.com/pricing [3] Together AI, “Pricing” (MiniMax M3 rate card, retrieved 2026-10-06) — https://www.together.ai/pricing [4] US Energy Information Administration, “Electricity Profile” (state data, 2024) — https://www.eia.gov/electricity/state/ [5] Nvidia, “GeForce 5090” product page (specifications) — https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/ [6] Editorial-factory field notes, workstation_compute_economics Anecdote 3 (internal) — context/growth_os/customer-truth.md