Staff Augmentation Metrics: What to Measure and When to Worry

In this article
- Measure the engagement, not just the code. Staff augmentation needs four layers of metrics: vendor service, engineer ramp and retention, delivery and quality, and cost per delivered unit.
- Twelve numbers are enough. Each one below has a definition, a data source you already own and green, amber and red thresholds you can adopt as a starting point.
- Collect from systems, not from people. Git, your tracker, CI and invoices give you everything. Screenshot and keystroke trackers measure activity and damage trust.
- One red metric is a conversation, two reds in the same layer are a plan. The dashboard tool below scores your current numbers and tells you where to look first.
Jump to
- Why augmentation needs its own metrics
- The twelve metrics, with thresholds
- Cost per delivered unit, worked through
- Score your engagement
- How to collect the data without micromanaging
- A dashboard that fits on one page
- Review cadence and who owns what
- When a metric turns red
- Put the vendor metrics in the contract
- Where Gilzor fits
Why augmentation needs its own metrics
Delivery metrics such as the DORA set (deployment frequency, lead time, change failure rate, time to restore) tell you how well a team ships. They're the right tool for a self-organizing vendor team, and we cover them in how to manage a dedicated development team. With staff augmentation they're necessary but not enough, for three reasons.
First, augmented engineers are mixed into your team, so team-level numbers can't tell you whether the extra people are paying off. Second, a large part of the value is the vendor's promise: fast candidates, fast replacements, people who stay. That promise needs its own measures. Third, you're paying a monthly rate per person, so the question "what do we get per dollar?" is fair and answerable.
So we split the metrics into four layers. Each answers a different question, and each has a different owner on the vendor's side.
The twelve metrics, with thresholds
These are the thresholds we use as defaults when a client asks us to agree on measurable terms. They assume mid or senior engineers and a reasonably healthy codebase. Adjust them after the first quarter, once you have your own baseline. A metric without a threshold is just a chart; a threshold tells everyone when to talk.
| Metric | How to measure | Green | Amber | Red |
|---|---|---|---|---|
| Vendor service | ||||
| Time to present candidates | Business days from your request to profiles you'd interview | ≤ 5 | 6–10 | > 10 |
| Interview-to-offer ratio | Candidates interviewed per accepted engineer | ≤ 3 | 4–5 | > 5 |
| Time to start | Calendar days from selection to first working day | ≤ 14 | 15–28 | > 28 |
| Time to replace | Business days from replacement request to a new engineer working | ≤ 15 | 16–30 | > 30 |
| Engineer ramp and retention | ||||
| Time to first merged PR | Working days from start to first merged change (Git) | ≤ 5 | 6–10 | > 10 |
| Time to productivity | Days until the engineer closes routine tickets without help | ≤ 30 | 31–50 | > 50 |
| Attrition of placed engineers | Unplanned departures in 12 months / average headcount | ≤ 10% | 11–20% | > 20% |
| Delivery and quality (after day 90) | ||||
| Relative throughput | Items completed per sprint vs your team's median per engineer | ≥ 80% | 60–79% | < 60% |
| PR cycle time | Median hours from PR opened to merged | ≤ 48 | 49–96 | > 96 |
| Rework rate | Share of PRs needing more than two review rounds | ≤ 15% | 16–30% | > 30% |
| QA return rate | Share of tickets sent back from QA to development | ≤ 10% | 11–20% | > 20% |
| Business value | ||||
| Cost per delivered unit | Monthly cost of augmented engineers / items delivered, change vs last quarter | Flat or down | Up 10–20% | Up > 20% |
A few notes on the less obvious lines. Time to productivity is the only judgment call in the list; define "routine ticket" with your team once, then let the buddy or tech lead mark the date. PR cycle time mixes the engineer's work and your review speed, which is deliberate: if it goes red because your reviewers are slow, that's a finding about your side. For reference on the QA line, our own internal number is 5% of tickets returned from QA to developers, so 10% as the green line is a generous default.
Attrition deserves a closer look than most clients give it. Every unplanned departure costs you a new ramp of one to three months, which is why a vendor that keeps people is worth a higher rate. Ask for attrition across the vendor's whole company as well as on your account; the company-wide figure predicts your future.
Cost per delivered unit, worked through
This is the metric finance will ask for, and the one most teams skip because "delivered unit" sounds hard to define. It doesn't need to be perfect. It needs to be consistent.
Pick the unit your team already counts: story points if your estimates are stable, completed tickets of a normal size class if not. Then divide the monthly cost of the augmented engineers by the units they delivered. An example:
Example: three augmented engineers, quarter two vs quarter three
- Q2: 3 engineers at $8,800 a month each = $26,400 a month. They complete on average 66 points a month. Cost per point: $400.
- Q3: same rate, same people, now past the ramp. They complete 84 points a month. Cost per point: $314, down 21%. Green.
- Q4 (warning case): one engineer leaves, a replacement starts mid-quarter. Output drops to 60 points. Cost per point: $440, up 40% against Q3. Red, and the cause is already visible in the attrition and replacement lines.
Never compare cost per point across teams. Points aren't a currency, and comparing them encourages inflation. Compare each team with its own past. For the rate side of the equation, see IT staff augmentation cost.
Built by Gilzor
Results we’ve shipped




Talk to the people who build it. Tell us about your project and get a free estimate of scope, timeline and cost.
Score your engagement
Enter your current numbers. The tool applies the thresholds above, marks each metric green, amber or red and tells you which layer needs attention first. Leave a field at its default if you don't track it yet.
Staff augmentation health dashboard
Thresholds as in the table above. The overall verdict weights the bottom layers more, because vendor and ramp problems turn into delivery problems a few weeks later.
How to collect the data without micromanaging
Every metric above comes from a system you already pay for. That's the point: nobody fills in a form, and nobody watches anyone's screen.
- Git hosting (GitHub, GitLab, Bitbucket): first merged PR, cycle time, review rounds. Most platforms export this; a small script or an engineering analytics tool does the rest.
- Your tracker (Jira, Linear, Azure Boards): throughput, QA returns via status transitions, and the date someone first closed a ticket alone if you tag it.
- Vendor reports and your inbox: request and start dates for candidates and replacements. A shared spreadsheet with four columns is enough.
- Finance: invoices and timesheets for cost per unit and billing accuracy.
Apply the delivery metrics to your in-house engineers too, from the same queries. If augmented engineers are the only people measured, they'll notice, and the numbers will become a performance theater. Report delivery metrics at team level in shared meetings and look at individual numbers only in 1:1s, as a starting point for a conversation, not a verdict.
Be careful with tools that capture screenshots, count keystrokes or score "activity". They measure presence, not output, and good engineers leave clients who use them. Many CEE vendors' engineers are employed under EU privacy law, and their employer may not be allowed to run that kind of monitoring for you anyway. If hours need verifying for billing, a timesheet tied to tickets does the job.
A word on AI coding tools: they inflate commit and line counts, which is one more reason to ignore both. Flow and quality measures still work. We look at how AI changes the picture in AI in staff augmentation.
A dashboard that fits on one page
The monthly vendor review needs one page, not a BI project. The layout we recommend:
- Top row: four tiles, one per layer, each colored by its worst metric. Anyone can read the state of the engagement in five seconds.
- Middle: the twelve metrics with this month, last month and a three-month sparkline. Trends matter more than single values; one amber month after a holiday is noise.
- Bottom: the people table: each engineer, start date, ramp status, any open feedback item, planned leave. This is where attrition risk gets discussed.
- Footer: actions agreed last month and their status.
Keep it in whatever your team already uses: a spreadsheet, a Notion page, a Looker or Grafana board. The format matters far less than the habit of reading it together every month.
Review cadence and who owns what
| Rhythm | What you look at | Who is in the room |
|---|---|---|
| Every sprint | Cycle time, rework, QA returns, throughput (team level) | Your tech lead and the team in the retro |
| Monthly | The full dashboard, people table, vendor service events | Your engineering manager and the vendor's account manager |
| Quarterly | Cost per delivered unit, attrition, whether the team size and model still fit the roadmap | Budget owner, engineering lead, vendor management |
| Per event | Candidate requests, starts, replacements, incidents | Whoever raised it, logged the same day |
When a metric turns red
A red metric is a question, not an answer. Here's where each one usually points, from what we see in practice:
| Red metric | Most common cause | First move |
|---|---|---|
| Time to present, time to replace | The vendor has no bench in your stack, or your requirements are too narrow | Ask how many candidates were screened; drop one "nice to have" |
| First merged PR | Access or local setup, almost never the engineer | Walk through the setup yourself; fix the README |
| Time to productivity | No buddy, scattered starter tickets, missing docs | Assign a buddy with blocked time and a single area to start in |
| Attrition | Rate pressure on the vendor, dull work, or no feedback from you | Ask the vendor directly; share praise as well as issues |
| Throughput, cycle time | Slow reviews, unclear tickets, too little overlap | Check where PRs wait; set a review SLA on your side |
| Rework, QA returns | Unwritten standards or a skills gap | Write the review checklist; if it persists, the feedback route in how to manage staff augmentation |
| Cost per unit | Ramp after a replacement, scope churn, or the team is the wrong size | Trace it down the layers; usually the cause is already red below |
If the same layer stays red for two months after a clear conversation, that's a vendor problem rather than a people problem, and it may be time to rethink the vendor or the model. Our guide to vetting augmented developers covers how to raise the bar at the front of the funnel so fewer metrics go red later.
Put the vendor metrics in the contract
The vendor service layer only works if the numbers are written down. Sales decks promise "candidates in 48 hours" and "replacement at no cost"; contracts often say nothing measurable. Before signing, ask for the following in the agreement or a service schedule:
- Candidate profiles within a set number of business days of a written request.
- A trial period per engineer (two to four weeks is common) with no penalty for release.
- A replacement target in business days, and whether the replacement's ramp-up is billed.
- A monthly report with hours, changes in the team and known risks.
- Notice periods in both directions and what happens to work in progress.
Treat these as service targets with a review clause, not as penalties. Penalty-heavy SLAs push vendors to send the first available person rather than the right one. The software outsourcing contract guide covers the legal side, and our staff augmentation RFP guide shows how to ask for these numbers before you shortlist.
FAQ
What are the most important staff augmentation metrics?
How do you measure the productivity of augmented developers fairly?
Should I use time-tracking or monitoring software for augmented staff?
What SLA should I ask a staff augmentation vendor for?
How often should staff augmentation metrics be reviewed?
Where Gilzor fits
We like being measured, because the numbers usually help us. Before an engagement starts we agree the metrics and thresholds with you, and the monthly report covers hours, team changes, ramp status and risks. Internally we track the same quality signals we recommend here; 98% of our work is delivered on time and 85% of our customers come back. If you're comparing vendors, our team extension page explains how we start, and our projects show what the teams have shipped.
No sales pitch
Get a straight answer for your project
Tell us what you’re building. We’ll reply with options, a rough cost and timeline. If we’re not the right fit, we’ll say so.

Co-Founder of Gilzor. Works with founders and product companies on how to staff and run engineering: team extension, dedicated teams, and getting stalled projects moving again.
Gilzor · Development Support partner
Need a team for your product?
Services
Development SupportTeam extension, maintenance, bug fixing, scaling.→By company type
Selected projects






The team behind them





