Skip to main content
When agencies underdeliver: a TA vendor scorecard, SLA thresholds, and remediation playbook

When agencies underdeliver: a TA vendor scorecard, SLA thresholds, and remediation playbook

Because "they're just not sending good people" isn't a metric you can act on

Most TA teams manage recruiting agencies on vibes. There's a Slack channel, a few email threads, a quarterly lunch where someone gets a steak, and a vague sense that Agency A is "solid" while Agency B "hasn't been great lately." Then renewal season hits, procurement asks for the data, and everyone scrambles to reconstruct performance from memory.

The problem with vibes is they favor agencies with the best account managers, not the best recruiters. A charming rep can smooth over months of slow submits and mediocre candidates. Meanwhile a quiet agency that actually delivers gets treated the same as everyone else because nobody's tracking the difference.

An agency performance scorecard for recruiting fixes this by turning fuzzy feelings into fields you can defend. But most scorecards you'll find floating around are useless — they measure vanity numbers like "number of resumes sent" and ignore the two things that actually predict whether an agency is worth the fee.

The two fields that predict everything: time-to-submit and quality conversion

If you only track two things per agency, track these.

Time-to-submit is the gap between when you release a req to an agency and when the first qualified submission lands. Not the first resume — the first candidate who clears your basic screen. Agencies love to submit fast and dirty to look responsive, so you have to measure time-to-qualified-submit or you'll reward spray-and-pray behavior.

Quality conversion is the more revealing number, and almost nobody tracks it properly. It's the ratio of submissions that advance past your first hiring-manager screen. An agency that sends 20 candidates and gets 2 into interviews has a 10% conversion rate. An agency that sends 6 and gets 3 through is running at 50% and is worth twice the money even if they "sent fewer people."

Agencies optimize for whatever you complain about most. If you nag them about volume, they flood you with bodies. If you nag them about speed, they submit half-baked profiles at midnight. Quality conversion is the one metric that's hard to game, because the hiring manager — not the agency — decides who advances.

FieldWhat it measuresHealthy rangeWatch threshold
Time-to-first-qualified-submitReq release → first screened candidate≤ 3 business days> 5 days
Submission-to-screen conversion% of submits that pass HM screen≥ 30%< 20%
Screen-to-onsite conversion% of screened who reach onsite≥ 50%< 35%
Submit-to-hire ratioTotal submits per hire≤ 8:1> 15:1
Fall-off rateHires who leave/renege inside 90 days< 8%> 15%
Resubmit/duplicate rateCandidates already in your ATS< 5%> 10%

The last row matters more than people think. Agencies resubmitting candidates you already sourced — or that another agency already submitted — is a common way ownership disputes and wasted screening time creep in. If your resubmit rate with a particular agency is climbing past 10%, they're probably working your req lazily, pulling from the same LinkedIn searches your internal team already ran.

Why the numbers drift even when the agency "seems fine"

A req gets assigned to three agencies at once because the role is hot and someone wants "maximum coverage." All three start submitting. Nobody's tracking per-agency conversion, just aggregate pipeline. Two of the three agencies figure out within a week that this is a low-effort contingency race, so they dump their existing bench into your inbox and move on to reqs where they have a better shot. Your pipeline looks full — 30 submissions in ten days — but most of them are recycled profiles nobody screened carefully.

By the time you notice conversion is trash, you've burned two weeks of hiring-manager screening time on candidates who were never a fit. The role's still open, the hiring manager's frustrated, and every agency blames "a tight market."

This is why aggregate pipeline health can look fine while individual agency performance quietly rots. You need the fields broken out per agency, per req, not rolled up. If you're already running a triage system to keep strong candidates from disappearing — and if you're not, this workflow for stopping top candidates from slipping through the cracks is a good starting point — the agency scorecard should feed into the same routing logic. A submission from a high-conversion agency should get faster human review than one from an agency running at 15%.

SLA thresholds that mean something

An SLA is worthless if it only exists in the MSA and nobody checks it. The trick is setting thresholds tight enough to matter but realistic enough that a good agency can hit them on a normal week.

  1. Response SLA — Agency acknowledges the req and confirms they're working it within 1 business day. Sounds trivial, but silence in the first 24 hours is the single strongest predictor that a req will get deprioritized on their end.
  2. First-submit SLA — First qualified candidate within 3–5 business days depending on role seniority. Senior and niche roles get the longer window; volume roles get the shorter one.
  3. Sustained-pipeline SLA — For roles open past two weeks, a minimum cadence of qualified submits per week (usually 2–3) so the req doesn't go dark after the initial burst.

The mistake most teams make is writing one SLA for every role type. A staff-level backend engineer and a mid-market SDR do not have the same realistic time-to-fill, and holding both to a 3-day submit window just teaches your agencies that your SLAs are arbitrary. Tier them by role family and difficulty.

One more thing worth building in: SLA credit for req quality on your side. If your intake was garbage, the job description was a copy-paste from three years ago, or the hiring manager took six days to give feedback, the agency's SLA clock should pause. Otherwise you'll get pushback — fairly — and the whole framework loses credibility. Good agencies will respect an SLA that cuts both ways far more than one that only points at them.

Remediation gates: what happens when an agency misses

This is the part almost every vendor program skips, and it's the part that actually changes behavior. Thresholds without consequences are just a spreadsheet.

A remediation gate is a defined trigger with a defined response. When an agency crosses a watch threshold, something specific happens — not "we'll keep an eye on it."

A workable escalation ladder:

  1. Gate 1 — Coaching flag. Two consecutive weeks below conversion threshold, or one missed first-submit SLA. Response: a short written note documenting the miss and a 15-minute call to realign on the role. No penalty yet. Most misses resolve here because they're just a bad intake or a misunderstanding about the profile.
  2. Gate 2 — Formal remediation. A full month below threshold across two or more reqs. Response: agency submits a written remediation plan (who's staffing your account, what changes), and you pause new req assignments until conversion recovers. Existing reqs continue.
  3. Gate 3 — Volume reduction. Two consecutive quarters of underperformance. Response: reduce their share of reqs, remove them from priority/first-look status, and put renewal explicitly at risk in writing.
  4. Gate 4 — Exit. Sustained failure past remediation. Response

    wind down active reqs, no new assignments, contract non-renewal.

Process diagram

A simple diagram like this helps keep the escalation ladder factual and predictable.

The reason the gates work is that they're predictable. An agency that knows exactly what triggers a volume cut behaves very differently from one that thinks they can charm their way through a quarterly review. And when you eventually do cut an underperformer, the documented gate history makes the conversation boring and factual instead of a fight.

The other benefit: gates protect your good agencies. When you reduce a weak vendor's share, that req volume has to go somewhere. Routing it to your high-conversion agencies — and having the capacity to absorb it internally — ties directly into how you forecast recruiter capacity and build in buffers so the shift doesn't overload your own team.

The quarterly review template

QBRs with agencies usually devolve into the account manager presenting a deck of flattering numbers they picked. Flip it. You bring the scorecard, you drive the agenda, they respond.

  1. Scorecard walkthrough — Their fields vs. thresholds vs. last quarter. No narrative from them yet; just the numbers on the table.
  2. Wins — Roles they filled well, standout candidates, anything genuinely good. Say it plainly; good agencies deserve to hear it.
  3. Misses and gates — Any thresholds crossed, any active remediation. This is where their explanation goes, and where you separate real market difficulty from account neglect.
  4. Req forecast — What's coming next quarter, what role families you'll lean on them for. Agencies staff your account based on how much future work they see.
  5. Two commitments each — Two things they'll change, two things you'll fix on your side (faster feedback, cleaner intake, whatever your own data shows).

The two-commitments-each structure is the part people underrate. When the review is only about their failures, agencies get defensive and the relationship sours. When you own your side — the slow hiring managers, the vague reqs — it becomes a working session instead of a performance trial, and the agency actually invests more of their good recruiters in your account.

Sample contractual language

You don't need a lawyer to draft these — you need your lawyer to approve language you've already shaped around how you actually operate. Rough starting points:

Performance metrics clause: > "Agency performance shall be evaluated quarterly against the following metrics: (a) time to first qualified submission, (b) submission-to-screen conversion rate, and (c) 90-day retention of placed candidates. Metric definitions and target thresholds are set forth in Exhibit [X] and may be updated by Client with 30 days' written notice."

SLA and clock-pause clause: > "Agency shall deliver a first qualified submission within the timeframe defined for the applicable role tier. SLA timeframes shall be suspended during any period in which Client has not provided required intake information or has not returned candidate feedback within [2] business days."

Remediation and reduction clause: > "Sustained performance below defined thresholds, as documented across two consecutive quarterly reviews, entitles Client to reduce requisition allocation, remove Agency from priority requisition status, or terminate for cause without penalty upon [30] days' written notice."

Candidate ownership / resubmit clause: > "Agency shall receive placement credit only for candidates not already present in Client's applicant tracking system at the time of submission. Candidates previously submitted by Client's internal team or another vendor shall not qualify for fees."

That last one prevents the ownership fights that eat up admin time when multiple agencies work the same req.

A real scenario

A mid-sized fintech running four contingency agencies had a submit-to-hire ratio hovering around 22:1 and no idea which agency was dragging it down. Everyone screened everything; nobody tracked conversion per vendor. Time-to-fill on engineering roles sat around 60 days and hiring managers were furious about "resume spam."

They stood up a per-agency scorecard — nothing fancy, a shared sheet updated weekly — tracking time-to-qualified-submit, screen conversion, and resubmit rate. Within about six weeks the picture was obvious: two agencies were converting around 30–35%, and one was sitting near 12% while accounting for almost half the total submissions. That same agency was responsible for most of the duplicate candidates.

They hit the low performer with a Gate 2 remediation, paused new reqs, and shifted that volume to the two strong agencies plus one new specialist vendor. A quarter later the blended submit-to-hire ratio was down around 11:1, hiring-manager screening time had roughly halved, and engineering time-to-fill had dropped by a couple of weeks. The underperforming agency eventually recovered its conversion after reassigning a senior recruiter to the account — which only happened because the gate made the consequence real.

Where scorecards backfire

This framework isn't free, and it's the wrong move in a couple of situations.

If you're running a single retained search for a rare executive role, per-week conversion thresholds are meaningless — the whole engagement might produce eight candidates over two months, and that's correct. Retained and highly specialized searches need a lighter-touch review, not a conversion funnel.

If you only send an agency two or three reqs a year, the sample size is too small to score fairly. One bad req can tank their numbers on pure noise. Below a certain volume, manage the relationship directly and skip the formal scorecard.

And if your own intake process is a mess — vague reqs, ghosting hiring managers, no feedback loop — building an agency scorecard first puts the blame in the wrong place. Fix your side of the SLA before you start grading vendors on theirs, or the whole exercise reads as bad faith and your best agencies will quietly deprioritize you.

The point

The agencies that make your program look good and the ones that quietly drain it often feel identical on a Tuesday afternoon. The difference only shows up when you measure time-to-qualified-submit and quality conversion per vendor, set thresholds that cut both ways, and attach real consequences to the misses.

None of this requires a huge system. A shared sheet, six honest fields, and the discipline to actually run the quarterly review with your numbers instead of their deck will tell you within a quarter which agencies to feed and which to phase out — and give you the paper trail to defend both decisions when renewal season arrives.

Built for Recruiters Optimized for recruitment workflows and team collaboration
Save Time Automate scheduling and streamline candidate management
Engage Candidates Faster communication and transparent hiring updates
Hire Better Data-driven insights to improve hiring decisions