· 18 min read

AI Model Development Cost in 2026: Fine-Tuning, Training and GPUs

A fine-tuned model built by a vendor costs about $25,000–$80,000 in 2026 when you tune a commercial model through its provider's API, $60,000–$200,000 when you fine-tune and host an open-weight model yourself, and $120,000–$500,000+ for a custom predictive or computer vision model trained on your own data. The surprise for most buyers: GPU training time is usually under 3% of that budget. The money goes into labeled data, engineering experiments, evaluation and the infrastructure that keeps the model running. Below are the ranges, current cloud GPU prices, labeling costs, a calculator for build and running cost, and a quiz that tells you whether you should be training a model at all.
A GPU chip connected to a labeled dataset grid and a stack of coins
Thinking about training or fine-tuning a model?Send us the task and a sample of your data. We will tell you whether a custom model is worth it, or what to try first.
Explore my options

The short answer

$25–80kFine-tune a commercial model through the provider's API, with dataset and evaluation
$60–200kFine-tune and self-host an open-weight model (Llama, Mistral, Qwen class)
$120–500k+Custom predictive ML or computer vision model trained on your data
$1M+Pretraining a language model from scratch. Rarely the right call

These are 2026 build budgets for US companies working with a development vendor: ML engineers, a data engineer, MLOps, QA and a part-time lead, plus the cost of data labeling and GPU time. Running cost comes on top and we deal with it separately, because for self-hosted models it is a fixed monthly bill that starts on day one.

This article is about the model itself: when you need your own, what data and compute it takes, and what it costs to keep it working. For AI projects in general (LLM features, retrieval systems, agents) start with our AI development cost guide. For agents that take actions across tools, see AI agent development cost.

First decision: prompt, retrieve, fine-tune or train

"We need our own model" is one of the most common sentences in our first calls about AI, and one of the most expensive if taken literally. There are four levels of owning a model, and each one up the ladder costs more, takes longer and creates more to maintain.

How much model do you need to own? 1. Prompting 2. Retrieval (RAG) 3. Fine-tuning 4. Custom model $5–25k $40–150k $25–200k $120–500k+ Model reads your documents at query time Model learns a behavior from your examples Needs 500–50k good examples and an eval set Model learns the task from your data only Labeled data, GPUs, MLOps, retraining Pretraining from scratch: millions more data work, more infrastructure, more upkeep Vendor-built, US buyer, 2026. Fine-tuning and RAG are not strictly ordered by price; they solve different problems and are often combined.
Move up a step only when the step below has been measured on your real data and falls short.

The two middle steps get confused the most. Retrieval gives a model knowledge: it looks up your documents at query time, so the answers stay current when the documents change. Fine-tuning changes behavior: the model learns a format, a tone, a narrow classification or a domain's way of phrasing things from examples. Fine-tuning is a poor way to teach facts, because the facts go stale the day after training and the model still invents details around them.

Data needs are smaller than people expect for behavior and larger than they expect for knowledge. Meta's 2023 LIMA paper showed that a 65B model fine-tuned on just 1,000 carefully chosen examples could produce answers people often rated close to much larger commercial assistants. Quality of examples beat quantity. A custom vision model detecting rare defects, on the other hand, may need tens of thousands of labeled images before it beats a human inspector.

  • What it is: you upload training examples to a model provider and get a private tuned version of their model, served through the same API.
  • Build: $25k–80k, mostly dataset building and evaluation. Training fees are usually a few hundred dollars. Timeline: 6–12 weeks.
  • Running: per-token pricing, typically somewhat higher than the base model. No GPUs to manage.
  • Good fit: strict output formats, classification and extraction at volume, moving a task from a large model to a cheaper small one.
  • Watch out: you don't own the weights. When the provider retires the base model, you redo the fine-tune on its successor.

Where the money goes

Buyers who read about nine-figure training runs expect GPUs to dominate the bill. For a business model they almost never do. The Stanford AI Index 2025, working with Epoch AI, estimated that Google's Gemini 1.0 Ultra cost about $192 million to train, and Epoch AI found frontier training costs growing 2–3× a year. That is the frontier lab problem. Fine-tuning an existing model on a few thousand examples uses the same kind of GPU for a few hours.

A typical $100k fine-tuning project, and its first year Build $100k Year one running, $40k 50% 24% 12% 12% GPU training: ~2% 55% 25% 20% same scale as the build bar Engineering and experiments Data collection and labeling (build) / retraining and new labels (run) Evaluation and monitoring MLOps and deployment GPU compute (training / inference)
Illustrative split from fine-tuning estimates we have prepared and reviewed: an open-weight model, self-hosted on one dedicated GPU. Compute is tiny during the build and becomes the biggest line once the model is serving traffic.

During the build, GPU time is a rounding error: 50 experiment runs on rented H100s cost less than a week of an engineer's time. After launch it flips. A self-hosted model needs a GPU awake around the clock, whether ten people use it or ten thousand.

Data collection and labeling

Every custom model needs examples of the right answer. Either you already have them (historical tables with outcomes, past support tickets with resolutions, documents with extracted values typed in by staff) or someone has to create them. Creating them is where budgets grow.

Labeling taskWho does itTypical cost per item10,000 items
Simple classification (spam / not spam, topic, sentiment)Crowd or labeling vendor, model-assisted$0.02–0.10$200–1,000
Bounding boxes on imagesTrained annotators$0.04–1.00 per object$400–10,000 (one object each)
Text spans, entities, document fieldsTrained annotators$0.20–2.00$2,000–20,000
Pixel segmentation, video trackingSpecialized annotators$2–10 per image or clip$20,000–100,000
Ideal answers for LLM fine-tuning, preference rankingYour experts or skilled writers$3–20$30,000–200,000
Medical, legal, engineering judgmentLicensed domain experts$10–50+$100,000–500,000
2026 market ranges. Add 20–30% for quality review, disagreements and relabeling. AWS publishes SageMaker Ground Truth pricing at about $0.036 per bounding box for its crowd workforce; enterprise vendors with managed quality cost several times more.

To keep this line under control, label a pilot batch of 500 items first and measure how often two annotators disagree; if they agree less than 85–90% of the time, your label definitions are the problem, not the model. Then pre-label with a pretrained model and use active learning (people label only the items the model is least sure about). Labeling vendors commonly report this cuts volume by half or more.

Built by Gilzor

Results we’ve shipped

70+products launched
98%delivered on time
85%clients come back
Art Scherbakov, Co-FounderAndrew Laminsky, CTOYuri Rudenya, Head of Mobile Development at GilzorAlena Timofeeva, Product Marketing Lead

Talk to the people who build it. Tell us about your project and get a free estimate of scope, timeline and cost.

See how we’d approach yours

GPU training and inference costs

Cloud GPU prices fell sharply between 2024 and 2026 as supply caught up. Here are typical on-demand rates for the GPUs used most in business model work.

On-demand cloud GPU price, USD per GPU-hour (2026, typical midpoints)

H100, hyperscaler 8-GPU instance ($5–7.50)$6.25
H100, specialist GPU cloud ($2–3.50)$2.70
A100 80GB ($1.50–3.50)$2.50
L40S ($0.80–1.80)$1.15
L4 24GB ($0.50–1.00)$0.75
Bar length relative to $6.25. Hyperscalers such as AWS sell H100s in 8-GPU instances, so the per-GPU price is derived. One-year commitments typically cut 30–40%; spot capacity is cheaper but can be taken away mid-run. Egress, storage and idle time add to the bill.

What that means for training, in GPU-hours per run:

  • LoRA fine-tune of a 7–8B model on 10,000–50,000 examples: roughly 2–6 hours on one A100 or H100. $5–20 per run.
  • QLoRA fine-tune of a 70B model: roughly 8–36 hours on one 80GB GPU. $20–100 per run.
  • Full fine-tune of a mid-size model or training a vision model from a pretrained backbone: tens to hundreds of GPU-hours on a multi-GPU node. $200–2,000 per run.
  • Provider fine-tuning APIs charge per training token instead, typically a few dollars to about $25 per million tokens depending on the model. A 5,000-example dataset trained for three epochs is usually a few hundred dollars or less.

A real project runs 20 to 50 experiments with different data mixes and settings. Even then, compute rarely passes a few thousand dollars for a fine-tune.

Inference is the opposite. A dedicated H100 for 730 hours a month costs about $2,000 on a specialist cloud and $4,500 on a hyperscaler. A production service wants two for redundancy. A quantized 8B model often runs fine on an L4 or L40S for $500–900 a month, which is one of the strongest reasons to fine-tune a small model instead of serving a large one. Serverless GPU platforms that bill per second help at low or spiky traffic, at the cost of cold starts.

The break-even question

Self-hosting a fine-tuned model is a fixed cost; API calls are a variable one. If your API bill for the same task is under $2–3k a month, self-hosting rarely pays for the extra engineering and on-call time. Above $10k a month for a stable, narrow task, a small fine-tuned model on your own GPUs usually wins.

Calculate your model cost

Pick the approach, describe the data you need labeled and the compute, and choose where the team sits. The calculator splits the build into engineering, labeling and GPU time, then shows the running cost. Engineering covers ML engineers, a data engineer, MLOps, QA and a part-time lead.

Custom AI model: build and running cost

Build budget (expect ±25%)
Engineering, evaluation and MLOps
Data labeling, including 25% for review and relabeling
GPU training compute, including storage and failed runs
Timeline with 3–4 people
Inference and monitoring per month (with 0 GPUs: a base estimate for API fees)
First-year total: build, 12 months of serving, 22% upkeep

Engineering hours are calibrated on typical 2026 scopes and multiplied by the compliance factor. Labeling cost is items × price per item × 1.25. Compute is runs × GPU-hours × hourly price × 1.3. Serving assumes the same GPU type as training, 730 hours a month, plus $300 for monitoring and logging. Upkeep covers retraining, new labels and evaluation at 22% of the build per year.

Moving from LoRA to a full fine-tune multiplies compute by fifty, and at the default settings it still stays well below the engineering line. Moving the labelers from skilled writers to licensed experts can double it. The serving slider decides whether the model is affordable to keep: each hyperscaler H100 adds about $55k a year.

MLOps and evaluation: the part demos skip

A model in a notebook is an experiment. A model in production needs a pipeline around it, and that pipeline is 10–25% of the build and a recurring cost after launch.

  • Data and model versioning. Which dataset produced which model, so a bad release can be rolled back and explained. Open-source tools like MLflow or DVC cost nothing in licenses; managed platforms typically run from a few hundred to a few thousand dollars a month.
  • Training pipeline. Retraining should be a button, not a week of an engineer remembering what they did in March.
  • Evaluation harness. A held-out test set of real cases, scored automatically on every new model. For LLM fine-tunes, add a comparison against the base model with a good prompt. If the fine-tune doesn't beat it clearly, ship the prompt.
  • Serving. An inference server, autoscaling, request logging, rate limits and a fallback when the model is down or unsure.
  • Monitoring. Accuracy on fresh labeled samples, input drift, latency and cost per request. Models degrade quietly; the dashboards are how you find out before customers do.

Google researchers made the same point in their 2015 paper "Hidden Technical Debt in Machine Learning Systems": the model code is a small box inside much larger systems for data, serving and monitoring. Our QA team treats evaluation sets the way it treats regression suites: versioned, run on every change, and owned by someone.

Team rates by region

ML engineering bills above general development. For context, the US Bureau of Labor Statistics put the median data scientist salary at $120,230 in May 2025, and senior ML engineers at US product companies earn well above that before benefits and recruiting fees. Vendor rates to US clients in 2026 typically look like this:

RoleUS onshoreLatin AmericaCentral & Eastern EuropeSouth & Southeast Asia
Senior ML / AI engineer$130–200/h$60–95/h$60–95/h$35–60/h
Data engineer$110–170/h$50–80/h$50–80/h$30–50/h
MLOps / platform engineer$120–180/h$55–85/h$55–85/h$30–55/h
Annotator (general)$20–40/h$8–15/h$10–18/h$4–10/h
Typical 2026 vendor rates to US clients. Domain-expert annotators (clinicians, lawyers) bill at professional rates in any region.

Latin America and Central Europe price about the same; the difference is time zone. Gilzor works from Poland and Cyprus, which for US clients is offshore with two to four shared hours with the East Coast on a shifted schedule. Model work suits that better than most: training runs, labeling and evaluation happen asynchronously, and a short daily overlap covers decisions. For country-level detail see nearshore software development rates, and for role-by-role numbers on app work, app development cost per hour.

Should you train a model at all?

Six questions about your task, data and volume. The result points to the cheapest approach likely to work.

Prompt, RAG, fine-tune or custom model?

Hidden costs to budget for

CostTypical sizeWhy it exists
Upkeep and retraining20–30% of build per yearThe usual software rule is 15–20% a year. Models sit higher: data drifts, labels need refreshing and retraining needs an engineer.
Base model retirement2–6 weeks, once or twice a yearAPI fine-tunes are tied to a base model version. When it is retired, you re-tune and re-evaluate on the successor.
Idle GPUsOften 30–60% of the serving billDedicated GPUs bill while nobody uses them. Nights and weekends are paid for unless you scale down.
Your experts' timeHours per week for monthsGold examples, label disputes and error reviews need people who know the domain.
Data access and legal$3k–30kData processing agreements, anonymization, rights to use customer data for training, and checks on open-weight model licenses.
Security and compliance$5k–50kPenetration tests, HIPAA safeguards and BAAs for health data, SOC 2 evidence, model access controls.
Storage and egress$100–2,000 a monthDatasets, checkpoints (a 70B checkpoint is well over 100 GB) and moving data between clouds.

Three budgets, worked through

Typical requests from US companies, priced with the calculator logic above at Central European or Latin American rates. Illustrative, not Gilzor quotes.

1. Ticket classifier moved from a large model to a fine-tuned small one

A SaaS company classifies 1.5 million support tickets a month into 40 categories with a frontier model, and the API bill has passed $12k a month. They have two years of tickets already categorized by agents.

  • No new labeling, only 6,000 tickets cleaned and checked by the team. Fine-tune a small model via the provider's API. Build: roughly $40k, eight weeks.
  • Running: the tuned small model costs a fraction of the frontier model per ticket, so the bill drops to low four figures a month. The project pays back within the first year.
  • What made it work: the labels already existed and the task never changes shape.

2. Private document model for a healthcare company

A health-tech company needs extraction and summaries of clinical notes. Data must stay in its own cloud account. Clinicians must write and review 3,000 gold examples.

  • LoRA fine-tune of a 70B-class open model, served on two GPUs for redundancy. Engineering about $95k with the HIPAA factor; expert labeling about $75k; compute under $5k.
  • Running: two A100s around the clock, about $3,600 a month, plus around $40k a year of upkeep.
  • Where it gets expensive: clinician time, not GPUs. The schedule follows how many hours the clinicians can give each week.

3. Visual defect detection on a production line

A manufacturer wants to catch surface defects that human inspectors miss on night shifts. Defects are rare, so collecting enough examples is slow.

  • About 1,700 engineering hours, 40,000 images with segmentation masks, a vision model trained from a pretrained backbone. Build: roughly $135k engineering plus $100k+ labeling.
  • The honest check: if defects are fewer than a few hundred a year, a model may never see enough of them. Synthetic images and anomaly detection can help; sometimes better lighting and a camera solve most of it.

What we see in estimates and first calls

  • "Our own model" usually means "a model that knows our stuff". That is a retrieval problem in most cases, and it costs less and stays fresher than a fine-tune.
  • Nobody has compared against a good prompt. We ask for this baseline before any fine-tuning estimate. In a meaningful share of cases the prompt plus a few examples is close enough, and the project shrinks to a few weeks.
  • The labels exist, but nobody trusts them. Years of tickets or records labeled inconsistently by different people are common. A one-week audit of 300 samples tells you whether the history is usable or needs relabeling.
  • Serving cost was never estimated. Teams budget the training and forget that a self-hosted model bills every hour. We put the monthly serving number next to accuracy in every proposal.
  • The evaluation set gets built last. Without it, nobody can say whether model version two is better than version one, and model upgrades turn into arguments. Our business analysis work often starts by defining the test cases and the accuracy the business case actually needs.

How to reduce the cost without breaking the model

  1. Measure the cheap options firstA strong prompt and retrieval, scored on 200 real cases, before anyone rents a GPU. It is the cheapest experiment and the baseline every later option must beat.
  2. Fine-tune small models, not big onesA 7–8B model tuned for one task often matches a frontier model on that task and serves on an L4 instead of an H100.
  3. Use LoRA before full fine-tuningParameter-efficient methods cut compute and storage by an order of magnitude and are easier to version and roll back.
  4. Label smarterPre-label with a model, have people correct, and use active learning to label only uncertain cases. Fix label definitions on a 500-item pilot before scaling.
  5. Right-size servingQuantize, batch requests, scale to zero overnight where latency allows, and commit to reserved capacity only once traffic is stable.
  6. Automate retraining earlyA pipeline that retrains and re-evaluates with one command turns yearly upkeep from a project into a routine.
  7. Don't cut evaluation or monitoringThey are the first lines removed to win a bid and the reason a model can be trusted, compared and replaced later.
The most expensive mistake

Starting with labeling 50,000 items before a pilot proves the labels are consistent and the approach works. Label 500, train, measure, then decide. The pilot costs a few thousand dollars; relabeling a full dataset costs the budget.

A custom model is usually one part of a product. For the full picture of AI budgets see AI development cost; for apps with AI inside, AI app development cost; for support assistants, chatbot development cost. If the model needs data pulled together from several systems first, read cost of data integration. To keep the GPU bill in check after launch, see cloud cost optimization.

FAQ

How much does it cost to develop a custom AI model?
For a US company working with a vendor in 2026: fine-tuning a commercial model through the provider's API typically costs $25k–80k including data preparation and evaluation. Fine-tuning and self-hosting an open-weight model costs $60k–200k. A custom predictive or computer vision model trained on your own data usually costs $120k–500k+, with data collection and labeling often the largest single line. Training a large language model from scratch is a different category that runs into millions.
How much does it cost to fine-tune an LLM?
The compute is cheap: a LoRA fine-tune of a 7–8B open model on a few thousand examples takes a few GPU hours, so tens of dollars per run on rented A100 or H100 GPUs. Provider fine-tuning APIs charge per training token, usually a few dollars to about $25 per million tokens depending on the model. The real cost is the project around it: building a dataset of good examples, running 20–50 experiments, evaluating against a baseline and deploying the result. That is why fine-tuning projects land at $25k–200k, not $500.
How much do cloud GPUs cost per hour in 2026?
Typical on-demand prices in 2026: an NVIDIA L4 around $0.50–1.00 per hour, an A100 80GB around $1.50–3.50, and an H100 around $2–3.50 on specialist GPU clouds and $5–7.50 per GPU on the big hyperscalers, where H100s are sold in 8-GPU instances. One-year commitments typically cut 30–40% off on-demand rates.
Should we fine-tune a model or use RAG?
Use retrieval (RAG) when the model needs to know facts that change: documents, prices, policies, tickets. Use fine-tuning when the model needs to behave differently: a fixed output format, a narrow classification task, a house style, or when a small fine-tuned model can replace an expensive large one at high volume. Many production systems combine both. Start with prompting and an evaluation set, add retrieval, and fine-tune only when the measurements show a gap.
How much does data labeling cost?
It depends on who has to do the labeling. Simple classification by a crowd or a labeling vendor can cost a few cents per item. Bounding boxes typically cost $0.04–1 per object, pixel-level segmentation several dollars per image, and examples written or reviewed by domain experts such as clinicians, lawyers or engineers $5–50 per item. Add 20–30% for quality review and relabeling.
What does it cost to run a custom model after launch?
Two parts. Inference: a self-hosted model on a dedicated GPU is a fixed bill of roughly $500–4,500 a month per GPU depending on the GPU and cloud, doubled if you need redundancy. Upkeep: monitoring, evaluation, new labeled data and periodic retraining typically cost 20–30% of the original build per year, above the 15–20% rule of thumb for ordinary software, because data drifts and models get stale.

Where Gilzor fits

We build AI and ML solutions for startups and product companies from our teams in Poland and Cyprus: fine-tuned and custom models, the data pipelines that feed them, and the web and mobile products around them. For US clients we are offshore, with an agreed two-to-four-hour overlap with the East Coast. Before we estimate a custom model, we ask for the baseline it has to beat, a sample of the data and the monthly volume, so the proposal shows build cost, labeling, serving cost and upkeep as separate lines you can check against any other quote.

No sales pitch

Get a straight answer for your project

Tell us what you’re building. We’ll reply with options, a rough cost and timeline. If we’re not the right fit, we’ll say so.

Next, a few optional questions so the first call is useful. We use your details only to reply to your request. Privacy Policy

Andrew Laminsky
Written byAndrew Laminsky

CTO of Gilzor. Responsible for architecture and the engineering standards our teams work by.

LinkedIn →

Gilzor · AI/ML Solutions partner

Need a team for your AI product?

95%referred by business partners
70+successful launches
85%repeat business
98%delivered on time

The team behind them

Art Scherbakov
Art ScherbakovCo-Founder
Andrew Laminsky
Andrew LaminskyCTOLinkedIn
Yuri Rudenya
Yuri RudenyaHead of Mobile Development at GilzorLinkedIn
Alena Timofeeva
Alena TimofeevaProduct Marketing LeadLinkedIn
Tell us what you’re buildingOptions, a rough cost and timeline for your project. No commitment.

More insights