AI Model Development Cost in 2026: Fine-Tuning, Training and GPUs

In this article
- In 2026 a vendor-built fine-tuned model costs about $25k–80k on a provider's fine-tuning API, $60k–200k as a self-hosted open-weight model, and $120k–500k+ as a custom ML or computer vision model trained on your own data.
- GPU training time is the smallest line. A LoRA run on an 8B model costs tens of dollars; the money goes into labeled data, engineering, evaluation and MLOps.
- Self-hosted inference is a fixed bill: one H100 running around the clock is roughly $1,500–4,500 a month depending on the cloud, before redundancy.
- Fine-tune only after prompting and retrieval have been measured and fall short. Most first calls we have about custom models end with a cheaper plan.
Jump to
- The short answer
- First decision: prompt, retrieve, fine-tune or train
- Where the money goes
- Data collection and labeling
- GPU training and inference costs
- Calculate your model cost
- MLOps and evaluation: the part demos skip
- Team rates by region
- Should you train a model at all?
- Hidden costs to budget for
- Three budgets, worked through
- What we see in estimates and first calls
- How to reduce the cost without breaking the model
- Related cost guides
- Where Gilzor fits
The short answer
These are 2026 build budgets for US companies working with a development vendor: ML engineers, a data engineer, MLOps, QA and a part-time lead, plus the cost of data labeling and GPU time. Running cost comes on top and we deal with it separately, because for self-hosted models it is a fixed monthly bill that starts on day one.
This article is about the model itself: when you need your own, what data and compute it takes, and what it costs to keep it working. For AI projects in general (LLM features, retrieval systems, agents) start with our AI development cost guide. For agents that take actions across tools, see AI agent development cost.
First decision: prompt, retrieve, fine-tune or train
"We need our own model" is one of the most common sentences in our first calls about AI, and one of the most expensive if taken literally. There are four levels of owning a model, and each one up the ladder costs more, takes longer and creates more to maintain.
The two middle steps get confused the most. Retrieval gives a model knowledge: it looks up your documents at query time, so the answers stay current when the documents change. Fine-tuning changes behavior: the model learns a format, a tone, a narrow classification or a domain's way of phrasing things from examples. Fine-tuning is a poor way to teach facts, because the facts go stale the day after training and the model still invents details around them.
Data needs are smaller than people expect for behavior and larger than they expect for knowledge. Meta's 2023 LIMA paper showed that a 65B model fine-tuned on just 1,000 carefully chosen examples could produce answers people often rated close to much larger commercial assistants. Quality of examples beat quantity. A custom vision model detecting rare defects, on the other hand, may need tens of thousands of labeled images before it beats a human inspector.
- What it is: you upload training examples to a model provider and get a private tuned version of their model, served through the same API.
- Build: $25k–80k, mostly dataset building and evaluation. Training fees are usually a few hundred dollars. Timeline: 6–12 weeks.
- Running: per-token pricing, typically somewhat higher than the base model. No GPUs to manage.
- Good fit: strict output formats, classification and extraction at volume, moving a task from a large model to a cheaper small one.
- Watch out: you don't own the weights. When the provider retires the base model, you redo the fine-tune on its successor.
- What it is: LoRA or full fine-tuning of an open-weight model on rented GPUs, then serving it on your own cloud account.
- Build: $60k–200k. Timeline: 3–6 months including serving infrastructure and evaluation.
- Running: a fixed GPU bill, roughly $500–4,500 a month per GPU, doubled for redundancy, plus MLOps time.
- Good fit: data that cannot leave your infrastructure, high and stable volume where per-token fees add up, latency-sensitive or offline use.
- Watch out: you now operate a GPU service: scaling, patching, uptime, security reviews.
- What it is: forecasting, churn, fraud, pricing or risk scoring trained on your historical tables.
- Build: $120k–400k, a third or more of it data engineering. Timeline: 4–9 months.
- Running: compute is cheap (often under $1k a month); retraining, monitoring and a data engineer's attention are not.
- Good fit: years of clean history, a decision made thousands of times, and a measurable payoff per percent of accuracy.
- Watch out: a simple statistical baseline sometimes gets within a few percent. Always build it first.
- What it is: detection, segmentation or classification of images and video: defects, documents, shelves, medical images.
- Build: $150k–600k+. Image collection and labeling often decide the budget. Timeline: 5–12 months.
- Running: GPU inference in the cloud or on edge devices, $1k–15k a month, plus hardware if cameras are part of the scope.
- Good fit: repetitive visual inspection with a clear definition of "bad", high volume, expensive misses.
- Watch out: lighting, camera and product changes break models and mean new images.
Where the money goes
Buyers who read about nine-figure training runs expect GPUs to dominate the bill. For a business model they almost never do. The Stanford AI Index 2025, working with Epoch AI, estimated that Google's Gemini 1.0 Ultra cost about $192 million to train, and Epoch AI found frontier training costs growing 2–3× a year. That is the frontier lab problem. Fine-tuning an existing model on a few thousand examples uses the same kind of GPU for a few hours.
During the build, GPU time is a rounding error: 50 experiment runs on rented H100s cost less than a week of an engineer's time. After launch it flips. A self-hosted model needs a GPU awake around the clock, whether ten people use it or ten thousand.
Data collection and labeling
Every custom model needs examples of the right answer. Either you already have them (historical tables with outcomes, past support tickets with resolutions, documents with extracted values typed in by staff) or someone has to create them. Creating them is where budgets grow.
| Labeling task | Who does it | Typical cost per item | 10,000 items |
|---|---|---|---|
| Simple classification (spam / not spam, topic, sentiment) | Crowd or labeling vendor, model-assisted | $0.02–0.10 | $200–1,000 |
| Bounding boxes on images | Trained annotators | $0.04–1.00 per object | $400–10,000 (one object each) |
| Text spans, entities, document fields | Trained annotators | $0.20–2.00 | $2,000–20,000 |
| Pixel segmentation, video tracking | Specialized annotators | $2–10 per image or clip | $20,000–100,000 |
| Ideal answers for LLM fine-tuning, preference ranking | Your experts or skilled writers | $3–20 | $30,000–200,000 |
| Medical, legal, engineering judgment | Licensed domain experts | $10–50+ | $100,000–500,000 |
To keep this line under control, label a pilot batch of 500 items first and measure how often two annotators disagree; if they agree less than 85–90% of the time, your label definitions are the problem, not the model. Then pre-label with a pretrained model and use active learning (people label only the items the model is least sure about). Labeling vendors commonly report this cuts volume by half or more.
Built by Gilzor
Results we’ve shipped




Talk to the people who build it. Tell us about your project and get a free estimate of scope, timeline and cost.
GPU training and inference costs
Cloud GPU prices fell sharply between 2024 and 2026 as supply caught up. Here are typical on-demand rates for the GPUs used most in business model work.
What that means for training, in GPU-hours per run:
- LoRA fine-tune of a 7–8B model on 10,000–50,000 examples: roughly 2–6 hours on one A100 or H100. $5–20 per run.
- QLoRA fine-tune of a 70B model: roughly 8–36 hours on one 80GB GPU. $20–100 per run.
- Full fine-tune of a mid-size model or training a vision model from a pretrained backbone: tens to hundreds of GPU-hours on a multi-GPU node. $200–2,000 per run.
- Provider fine-tuning APIs charge per training token instead, typically a few dollars to about $25 per million tokens depending on the model. A 5,000-example dataset trained for three epochs is usually a few hundred dollars or less.
A real project runs 20 to 50 experiments with different data mixes and settings. Even then, compute rarely passes a few thousand dollars for a fine-tune.
Inference is the opposite. A dedicated H100 for 730 hours a month costs about $2,000 on a specialist cloud and $4,500 on a hyperscaler. A production service wants two for redundancy. A quantized 8B model often runs fine on an L4 or L40S for $500–900 a month, which is one of the strongest reasons to fine-tune a small model instead of serving a large one. Serverless GPU platforms that bill per second help at low or spiky traffic, at the cost of cold starts.
Self-hosting a fine-tuned model is a fixed cost; API calls are a variable one. If your API bill for the same task is under $2–3k a month, self-hosting rarely pays for the extra engineering and on-call time. Above $10k a month for a stable, narrow task, a small fine-tuned model on your own GPUs usually wins.
Calculate your model cost
Pick the approach, describe the data you need labeled and the compute, and choose where the team sits. The calculator splits the build into engineering, labeling and GPU time, then shows the running cost. Engineering covers ML engineers, a data engineer, MLOps, QA and a part-time lead.
Custom AI model: build and running cost
Engineering hours are calibrated on typical 2026 scopes and multiplied by the compliance factor. Labeling cost is items × price per item × 1.25. Compute is runs × GPU-hours × hourly price × 1.3. Serving assumes the same GPU type as training, 730 hours a month, plus $300 for monitoring and logging. Upkeep covers retraining, new labels and evaluation at 22% of the build per year.
Moving from LoRA to a full fine-tune multiplies compute by fifty, and at the default settings it still stays well below the engineering line. Moving the labelers from skilled writers to licensed experts can double it. The serving slider decides whether the model is affordable to keep: each hyperscaler H100 adds about $55k a year.
MLOps and evaluation: the part demos skip
A model in a notebook is an experiment. A model in production needs a pipeline around it, and that pipeline is 10–25% of the build and a recurring cost after launch.
- Data and model versioning. Which dataset produced which model, so a bad release can be rolled back and explained. Open-source tools like MLflow or DVC cost nothing in licenses; managed platforms typically run from a few hundred to a few thousand dollars a month.
- Training pipeline. Retraining should be a button, not a week of an engineer remembering what they did in March.
- Evaluation harness. A held-out test set of real cases, scored automatically on every new model. For LLM fine-tunes, add a comparison against the base model with a good prompt. If the fine-tune doesn't beat it clearly, ship the prompt.
- Serving. An inference server, autoscaling, request logging, rate limits and a fallback when the model is down or unsure.
- Monitoring. Accuracy on fresh labeled samples, input drift, latency and cost per request. Models degrade quietly; the dashboards are how you find out before customers do.
Google researchers made the same point in their 2015 paper "Hidden Technical Debt in Machine Learning Systems": the model code is a small box inside much larger systems for data, serving and monitoring. Our QA team treats evaluation sets the way it treats regression suites: versioned, run on every change, and owned by someone.
Team rates by region
ML engineering bills above general development. For context, the US Bureau of Labor Statistics put the median data scientist salary at $120,230 in May 2025, and senior ML engineers at US product companies earn well above that before benefits and recruiting fees. Vendor rates to US clients in 2026 typically look like this:
| Role | US onshore | Latin America | Central & Eastern Europe | South & Southeast Asia |
|---|---|---|---|---|
| Senior ML / AI engineer | $130–200/h | $60–95/h | $60–95/h | $35–60/h |
| Data engineer | $110–170/h | $50–80/h | $50–80/h | $30–50/h |
| MLOps / platform engineer | $120–180/h | $55–85/h | $55–85/h | $30–55/h |
| Annotator (general) | $20–40/h | $8–15/h | $10–18/h | $4–10/h |
Latin America and Central Europe price about the same; the difference is time zone. Gilzor works from Poland and Cyprus, which for US clients is offshore with two to four shared hours with the East Coast on a shifted schedule. Model work suits that better than most: training runs, labeling and evaluation happen asynchronously, and a short daily overlap covers decisions. For country-level detail see nearshore software development rates, and for role-by-role numbers on app work, app development cost per hour.
Should you train a model at all?
Six questions about your task, data and volume. The result points to the cheapest approach likely to work.
Prompt, RAG, fine-tune or custom model?
Hidden costs to budget for
| Cost | Typical size | Why it exists |
|---|---|---|
| Upkeep and retraining | 20–30% of build per year | The usual software rule is 15–20% a year. Models sit higher: data drifts, labels need refreshing and retraining needs an engineer. |
| Base model retirement | 2–6 weeks, once or twice a year | API fine-tunes are tied to a base model version. When it is retired, you re-tune and re-evaluate on the successor. |
| Idle GPUs | Often 30–60% of the serving bill | Dedicated GPUs bill while nobody uses them. Nights and weekends are paid for unless you scale down. |
| Your experts' time | Hours per week for months | Gold examples, label disputes and error reviews need people who know the domain. |
| Data access and legal | $3k–30k | Data processing agreements, anonymization, rights to use customer data for training, and checks on open-weight model licenses. |
| Security and compliance | $5k–50k | Penetration tests, HIPAA safeguards and BAAs for health data, SOC 2 evidence, model access controls. |
| Storage and egress | $100–2,000 a month | Datasets, checkpoints (a 70B checkpoint is well over 100 GB) and moving data between clouds. |
Three budgets, worked through
Typical requests from US companies, priced with the calculator logic above at Central European or Latin American rates. Illustrative, not Gilzor quotes.
1. Ticket classifier moved from a large model to a fine-tuned small one
A SaaS company classifies 1.5 million support tickets a month into 40 categories with a frontier model, and the API bill has passed $12k a month. They have two years of tickets already categorized by agents.
- No new labeling, only 6,000 tickets cleaned and checked by the team. Fine-tune a small model via the provider's API. Build: roughly $40k, eight weeks.
- Running: the tuned small model costs a fraction of the frontier model per ticket, so the bill drops to low four figures a month. The project pays back within the first year.
- What made it work: the labels already existed and the task never changes shape.
2. Private document model for a healthcare company
A health-tech company needs extraction and summaries of clinical notes. Data must stay in its own cloud account. Clinicians must write and review 3,000 gold examples.
- LoRA fine-tune of a 70B-class open model, served on two GPUs for redundancy. Engineering about $95k with the HIPAA factor; expert labeling about $75k; compute under $5k.
- Running: two A100s around the clock, about $3,600 a month, plus around $40k a year of upkeep.
- Where it gets expensive: clinician time, not GPUs. The schedule follows how many hours the clinicians can give each week.
3. Visual defect detection on a production line
A manufacturer wants to catch surface defects that human inspectors miss on night shifts. Defects are rare, so collecting enough examples is slow.
- About 1,700 engineering hours, 40,000 images with segmentation masks, a vision model trained from a pretrained backbone. Build: roughly $135k engineering plus $100k+ labeling.
- The honest check: if defects are fewer than a few hundred a year, a model may never see enough of them. Synthetic images and anomaly detection can help; sometimes better lighting and a camera solve most of it.
What we see in estimates and first calls
- "Our own model" usually means "a model that knows our stuff". That is a retrieval problem in most cases, and it costs less and stays fresher than a fine-tune.
- Nobody has compared against a good prompt. We ask for this baseline before any fine-tuning estimate. In a meaningful share of cases the prompt plus a few examples is close enough, and the project shrinks to a few weeks.
- The labels exist, but nobody trusts them. Years of tickets or records labeled inconsistently by different people are common. A one-week audit of 300 samples tells you whether the history is usable or needs relabeling.
- Serving cost was never estimated. Teams budget the training and forget that a self-hosted model bills every hour. We put the monthly serving number next to accuracy in every proposal.
- The evaluation set gets built last. Without it, nobody can say whether model version two is better than version one, and model upgrades turn into arguments. Our business analysis work often starts by defining the test cases and the accuracy the business case actually needs.
How to reduce the cost without breaking the model
- Measure the cheap options firstA strong prompt and retrieval, scored on 200 real cases, before anyone rents a GPU. It is the cheapest experiment and the baseline every later option must beat.
- Fine-tune small models, not big onesA 7–8B model tuned for one task often matches a frontier model on that task and serves on an L4 instead of an H100.
- Use LoRA before full fine-tuningParameter-efficient methods cut compute and storage by an order of magnitude and are easier to version and roll back.
- Label smarterPre-label with a model, have people correct, and use active learning to label only uncertain cases. Fix label definitions on a 500-item pilot before scaling.
- Right-size servingQuantize, batch requests, scale to zero overnight where latency allows, and commit to reserved capacity only once traffic is stable.
- Automate retraining earlyA pipeline that retrains and re-evaluates with one command turns yearly upkeep from a project into a routine.
- Don't cut evaluation or monitoringThey are the first lines removed to win a bid and the reason a model can be trusted, compared and replaced later.
Starting with labeling 50,000 items before a pilot proves the labels are consistent and the approach works. Label 500, train, measure, then decide. The pilot costs a few thousand dollars; relabeling a full dataset costs the budget.
Related cost guides
A custom model is usually one part of a product. For the full picture of AI budgets see AI development cost; for apps with AI inside, AI app development cost; for support assistants, chatbot development cost. If the model needs data pulled together from several systems first, read cost of data integration. To keep the GPU bill in check after launch, see cloud cost optimization.
FAQ
How much does it cost to develop a custom AI model?
How much does it cost to fine-tune an LLM?
How much do cloud GPUs cost per hour in 2026?
Should we fine-tune a model or use RAG?
How much does data labeling cost?
What does it cost to run a custom model after launch?
Where Gilzor fits
We build AI and ML solutions for startups and product companies from our teams in Poland and Cyprus: fine-tuned and custom models, the data pipelines that feed them, and the web and mobile products around them. For US clients we are offshore, with an agreed two-to-four-hour overlap with the East Coast. Before we estimate a custom model, we ask for the baseline it has to beat, a sample of the data and the monthly volume, so the proposal shows build cost, labeling, serving cost and upkeep as separate lines you can check against any other quote.
No sales pitch
Get a straight answer for your project
Tell us what you’re building. We’ll reply with options, a rough cost and timeline. If we’re not the right fit, we’ll say so.

CTO of Gilzor. Responsible for architecture and the engineering standards our teams work by.
LinkedIn →Gilzor · AI/ML Solutions partner
Need a team for your AI product?
Services
AI/ML SolutionsComputer vision, NLP, predictive analytics, automation.→By company type
Selected projects






The team behind them





