· 11 min read

Staff Augmentation Metrics: What to Measure and When to Worry

Most companies measure staff augmentation with two numbers: the invoice and a feeling. The invoice is precise and the feeling is late. By the time a manager says out loud that an engineer or a vendor isn't working out, the data usually showed it two months earlier. Here are the metrics we'd track if we were on the client side, where each one comes from, the thresholds that should make you act, and a dashboard tool to score your own numbers.
A dashboard with gauges and red, amber and green indicators next to a team of engineers
Want an engagement you can measure?We agree the metrics and thresholds with you before the first engineer starts.
Explore my options

Why augmentation needs its own metrics

Delivery metrics such as the DORA set (deployment frequency, lead time, change failure rate, time to restore) tell you how well a team ships. They're the right tool for a self-organizing vendor team, and we cover them in how to manage a dedicated development team. With staff augmentation they're necessary but not enough, for three reasons.

First, augmented engineers are mixed into your team, so team-level numbers can't tell you whether the extra people are paying off. Second, a large part of the value is the vendor's promise: fast candidates, fast replacements, people who stay. That promise needs its own measures. Third, you're paying a monthly rate per person, so the question "what do we get per dollar?" is fair and answerable.

So we split the metrics into four layers. Each answers a different question, and each has a different owner on the vendor's side.

Business value Delivery and quality Engineer ramp and retention Vendor service Cost per delivered unit, trend per quarter Throughput, cycle time, rework, QA returns First PR, time to productivity, attrition Time to present, time to start, replacement, billing Is it worth it? Is the work good? Are the peopleworking out? Is the vendorkeeping promises? Quarterly Every sprint Monthly Per event
Problems usually start at the bottom and travel up. A slow replacement shows up as a ramp problem, then as a delivery dip, then as a worse cost per unit a quarter later.

The twelve metrics, with thresholds

These are the thresholds we use as defaults when a client asks us to agree on measurable terms. They assume mid or senior engineers and a reasonably healthy codebase. Adjust them after the first quarter, once you have your own baseline. A metric without a threshold is just a chart; a threshold tells everyone when to talk.

MetricHow to measureGreenAmberRed
Vendor service
Time to present candidatesBusiness days from your request to profiles you'd interview≤ 56–10> 10
Interview-to-offer ratioCandidates interviewed per accepted engineer≤ 34–5> 5
Time to startCalendar days from selection to first working day≤ 1415–28> 28
Time to replaceBusiness days from replacement request to a new engineer working≤ 1516–30> 30
Engineer ramp and retention
Time to first merged PRWorking days from start to first merged change (Git)≤ 56–10> 10
Time to productivityDays until the engineer closes routine tickets without help≤ 3031–50> 50
Attrition of placed engineersUnplanned departures in 12 months / average headcount≤ 10%11–20%> 20%
Delivery and quality (after day 90)
Relative throughputItems completed per sprint vs your team's median per engineer≥ 80%60–79%< 60%
PR cycle timeMedian hours from PR opened to merged≤ 4849–96> 96
Rework rateShare of PRs needing more than two review rounds≤ 15%16–30%> 30%
QA return rateShare of tickets sent back from QA to development≤ 10%11–20%> 20%
Business value
Cost per delivered unitMonthly cost of augmented engineers / items delivered, change vs last quarterFlat or downUp 10–20%Up > 20%

A few notes on the less obvious lines. Time to productivity is the only judgment call in the list; define "routine ticket" with your team once, then let the buddy or tech lead mark the date. PR cycle time mixes the engineer's work and your review speed, which is deliberate: if it goes red because your reviewers are slow, that's a finding about your side. For reference on the QA line, our own internal number is 5% of tickets returned from QA to developers, so 10% as the green line is a generous default.

Attrition deserves a closer look than most clients give it. Every unplanned departure costs you a new ramp of one to three months, which is why a vendor that keeps people is worth a higher rate. Ask for attrition across the vendor's whole company as well as on your account; the company-wide figure predicts your future.

Cost per delivered unit, worked through

This is the metric finance will ask for, and the one most teams skip because "delivered unit" sounds hard to define. It doesn't need to be perfect. It needs to be consistent.

Pick the unit your team already counts: story points if your estimates are stable, completed tickets of a normal size class if not. Then divide the monthly cost of the augmented engineers by the units they delivered. An example:

Example: three augmented engineers, quarter two vs quarter three

  • Q2: 3 engineers at $8,800 a month each = $26,400 a month. They complete on average 66 points a month. Cost per point: $400.
  • Q3: same rate, same people, now past the ramp. They complete 84 points a month. Cost per point: $314, down 21%. Green.
  • Q4 (warning case): one engineer leaves, a replacement starts mid-quarter. Output drops to 60 points. Cost per point: $440, up 40% against Q3. Red, and the cause is already visible in the attrition and replacement lines.

Never compare cost per point across teams. Points aren't a currency, and comparing them encourages inflation. Compare each team with its own past. For the rate side of the equation, see IT staff augmentation cost.

Built by Gilzor

Results we’ve shipped

70+products launched
98%delivered on time
85%clients come back
Art Scherbakov, Co-FounderAndrew Laminsky, CTOYuri Rudenya, Head of Mobile Development at GilzorAlena Timofeeva, Product Marketing Lead

Talk to the people who build it. Tell us about your project and get a free estimate of scope, timeline and cost.

See how we’d approach yours

Score your engagement

Enter your current numbers. The tool applies the thresholds above, marks each metric green, amber or red and tells you which layer needs attention first. Leave a field at its default if you don't track it yet.

Staff augmentation health dashboard

Thresholds as in the table above. The overall verdict weights the bottom layers more, because vendor and ramp problems turn into delivery problems a few weeks later.

How to collect the data without micromanaging

Every metric above comes from a system you already pay for. That's the point: nobody fills in a form, and nobody watches anyone's screen.

  • Git hosting (GitHub, GitLab, Bitbucket): first merged PR, cycle time, review rounds. Most platforms export this; a small script or an engineering analytics tool does the rest.
  • Your tracker (Jira, Linear, Azure Boards): throughput, QA returns via status transitions, and the date someone first closed a ticket alone if you tag it.
  • Vendor reports and your inbox: request and start dates for candidates and replacements. A shared spreadsheet with four columns is enough.
  • Finance: invoices and timesheets for cost per unit and billing accuracy.
Measure the same way for everyone

Apply the delivery metrics to your in-house engineers too, from the same queries. If augmented engineers are the only people measured, they'll notice, and the numbers will become a performance theater. Report delivery metrics at team level in shared meetings and look at individual numbers only in 1:1s, as a starting point for a conversation, not a verdict.

Be careful with tools that capture screenshots, count keystrokes or score "activity". They measure presence, not output, and good engineers leave clients who use them. Many CEE vendors' engineers are employed under EU privacy law, and their employer may not be allowed to run that kind of monitoring for you anyway. If hours need verifying for billing, a timesheet tied to tickets does the job.

A word on AI coding tools: they inflate commit and line counts, which is one more reason to ignore both. Flow and quality measures still work. We look at how AI changes the picture in AI in staff augmentation.

A dashboard that fits on one page

The monthly vendor review needs one page, not a BI project. The layout we recommend:

  1. Top row: four tiles, one per layer, each colored by its worst metric. Anyone can read the state of the engagement in five seconds.
  2. Middle: the twelve metrics with this month, last month and a three-month sparkline. Trends matter more than single values; one amber month after a holiday is noise.
  3. Bottom: the people table: each engineer, start date, ramp status, any open feedback item, planned leave. This is where attrition risk gets discussed.
  4. Footer: actions agreed last month and their status.

Keep it in whatever your team already uses: a spreadsheet, a Notion page, a Looker or Grafana board. The format matters far less than the habit of reading it together every month.

Review cadence and who owns what

RhythmWhat you look atWho is in the room
Every sprintCycle time, rework, QA returns, throughput (team level)Your tech lead and the team in the retro
MonthlyThe full dashboard, people table, vendor service eventsYour engineering manager and the vendor's account manager
QuarterlyCost per delivered unit, attrition, whether the team size and model still fit the roadmapBudget owner, engineering lead, vendor management
Per eventCandidate requests, starts, replacements, incidentsWhoever raised it, logged the same day

When a metric turns red

A red metric is a question, not an answer. Here's where each one usually points, from what we see in practice:

Red metricMost common causeFirst move
Time to present, time to replaceThe vendor has no bench in your stack, or your requirements are too narrowAsk how many candidates were screened; drop one "nice to have"
First merged PRAccess or local setup, almost never the engineerWalk through the setup yourself; fix the README
Time to productivityNo buddy, scattered starter tickets, missing docsAssign a buddy with blocked time and a single area to start in
AttritionRate pressure on the vendor, dull work, or no feedback from youAsk the vendor directly; share praise as well as issues
Throughput, cycle timeSlow reviews, unclear tickets, too little overlapCheck where PRs wait; set a review SLA on your side
Rework, QA returnsUnwritten standards or a skills gapWrite the review checklist; if it persists, the feedback route in how to manage staff augmentation
Cost per unitRamp after a replacement, scope churn, or the team is the wrong sizeTrace it down the layers; usually the cause is already red below

If the same layer stays red for two months after a clear conversation, that's a vendor problem rather than a people problem, and it may be time to rethink the vendor or the model. Our guide to vetting augmented developers covers how to raise the bar at the front of the funnel so fewer metrics go red later.

Put the vendor metrics in the contract

The vendor service layer only works if the numbers are written down. Sales decks promise "candidates in 48 hours" and "replacement at no cost"; contracts often say nothing measurable. Before signing, ask for the following in the agreement or a service schedule:

  • Candidate profiles within a set number of business days of a written request.
  • A trial period per engineer (two to four weeks is common) with no penalty for release.
  • A replacement target in business days, and whether the replacement's ramp-up is billed.
  • A monthly report with hours, changes in the team and known risks.
  • Notice periods in both directions and what happens to work in progress.

Treat these as service targets with a review clause, not as penalties. Penalty-heavy SLAs push vendors to send the first available person rather than the right one. The software outsourcing contract guide covers the legal side, and our staff augmentation RFP guide shows how to ask for these numbers before you shortlist.

FAQ

What are the most important staff augmentation metrics?
If you track only five: time to productivity (days until an engineer closes routine work alone), cycle time from pull request to merge, rework or QA return rate, retention of placed engineers, and cost per delivered unit. Add vendor service metrics such as time to present candidates and time to replace once you have more than two or three augmented engineers.
How do you measure the productivity of augmented developers fairly?
Compare them with your own team on the same metrics, from the same systems, after the ramp-up period (usually around day 90). Use flow and quality measures such as cycle time, throughput relative to the team median and rework rate, and look at trends over several sprints. Never use lines of code, commit counts or hours online as productivity measures.
Should I use time-tracking or monitoring software for augmented staff?
For billing, a simple timesheet is enough and is normal in time-and-materials contracts. Screenshot capture, keystroke logging and activity scores tell you little about output, push good engineers away and can create privacy issues in the European Union, where many vendors employ their people. Measure delivered work instead.
What SLA should I ask a staff augmentation vendor for?
Reasonable 2026 targets: candidate profiles within 5 business days of a request, a start within 2–3 weeks of selection, a replacement presented within 10–15 business days of a request, a trial period during which you can release an engineer without penalty, and a monthly report on hours, changes and risks. Put the numbers in the contract, not just in the sales deck.
How often should staff augmentation metrics be reviewed?
Look at delivery and quality signals every sprint inside the team, review the full dashboard with the vendor once a month, and do a quarterly review of cost per delivered unit, retention and whether the engagement still fits your plans. Monthly is the right rhythm for most vendor conversations; weekly reviews of individual numbers drift into micromanagement.

Where Gilzor fits

We like being measured, because the numbers usually help us. Before an engagement starts we agree the metrics and thresholds with you, and the monthly report covers hours, team changes, ramp status and risks. Internally we track the same quality signals we recommend here; 98% of our work is delivered on time and 85% of our customers come back. If you're comparing vendors, our team extension page explains how we start, and our projects show what the teams have shipped.

No sales pitch

Get a straight answer for your project

Tell us what you’re building. We’ll reply with options, a rough cost and timeline. If we’re not the right fit, we’ll say so.

Next, a few optional questions so the first call is useful. We use your details only to reply to your request. Privacy Policy

Art Scherbakov
Written byArt Scherbakov

Co-Founder of Gilzor. Works with founders and product companies on how to staff and run engineering: team extension, dedicated teams, and getting stalled projects moving again.

Gilzor · Development Support partner

Need a team for your product?

95%referred by business partners
70+successful launches
85%repeat business
98%delivered on time

The team behind them

Art Scherbakov
Art ScherbakovCo-Founder
Andrew Laminsky
Andrew LaminskyCTOLinkedIn
Yuri Rudenya
Yuri RudenyaHead of Mobile Development at GilzorLinkedIn
Alena Timofeeva
Alena TimofeevaProduct Marketing LeadLinkedIn
Tell us what you’re buildingOptions, a rough cost and timeline for your project. No commitment.

More insights