On 26 March 2026, Zendesk closed its acquisition of Forethought. Two months earlier it had already split automated outcomes into three tiers — assisted escalation, contained resolution, and verified resolution, the only one that draws from your allowance — with the verdict produced by an LLM after a 72-hour window with no customer follow-up. Put those two facts together and one company now sells you the AI agent, defines what counts as a resolution, runs the model that decides whether the definition was met, sends the invoice, and sells the quality-assurance product that grades the whole arrangement.
That is not an accusation of bad faith. It is a description of a measurement problem. This stack is the audit layer: it scores every conversation, samples the ones most likely to be wrong, keeps a reconciliation ledger the agent vendor did not write, and turns the difference between the machine score and the human score into a calibration decision rather than an argument.
The shape
- Zendesk is the system of record and the meter. Suite Team is $55 and Suite Professional $115 per agent per month paid yearly. Automated resolutions are included at 5, 10, or 15 per agent per month depending on plan, capped at 10,000 per year, then billed at $1.50 committed or $2.00 pay-as-you-go. Every number in your dispute has to be keyed to a conversation ID that lives here.
- Decagon or Fin (Intercom) is the graded party. Whichever agent resolves, this stack treats its output as evidence rather than reporting. The head-to-head on the agents themselves is Decagon vs Fin; the deployment decisions sit in the AI support agent stack.
- Zendesk QA or Solidroad is the scorer, and which one you pick is the whole argument. Zendesk QA runs AutoQA across every interaction including AI agents and voice, flags churn risk and stuck loops through Spotlight, scores live conversations, and puts bot and human scores side by side in AI Agent QA. It is sold in the Workforce Engagement Bundle at $50 per agent per month paid yearly, which also covers workforce management. Solidroad sells two separate product lines — QA automation, coaching and training simulations for humans; QA and benchmarking, process improvements and custom AI scorecards for agents — and connects directly to Decagon, Sierra and Fin to evaluate them, importing policies and macros from Guru and Notion so the scorecard reflects your actual rules. It raised a $25M Series A on 16 April 2026 led by Hedosophia with First Round Capital, Y Combinator and Sony Innovation Fund, and names Ryanair, ŌURA, Crypto.com and ActiveCampaign as customers. It publishes no pricing; the only route to a number is a demo.
- Assembled is the capacity layer, and it publishes its prices. Workforce management runs $25 (Core), $45 (Pro) and $75 (Enterprise) per agent per month plus a platform fee, AI Copilot starts at $35 per agent per month, and its own AI agents start at $0.65 per conversation. It sells itself on managing in-house agents, BPO vendors and AI in one place, which matters here for a narrow reason: when the agent absorbs 60% of volume, the humans keep the residual, and the residual is harder per conversation than the average was.
- n8n draws the sample and keeps the ledger. It is the cheapest component and the only record in the stack that no vendor being audited can rewrite.
- Slack is where calibration happens. Not a dashboard — a weekly thread with the conversations where the machine score and the human score disagreed, and a decision at the end of it.
Named handoffs
- Conversation closes in Zendesk → webhook fires → n8n stratifies the sample. Risk buckets first: reopened within seven days, escalated after agent handling, CSAT of 1 or 2, accounts over a revenue threshold. Random draw fills whatever quota is left.
- Every conversation → the QA scorer, automatically → the sampled subset → a human reviewer. The machine covers 100%; the human covers the part where being wrong is expensive.
- The 72-hour window closes → Zendesk assigns a tier → n8n writes the row. Conversation ID, tier, QA score, reopened-within-seven-days flag, escalation flag, in your own store.
- Month end → reconcile the ledger against the invoice. A conversation billed as a verified resolution that also produced a reopen inside seven days is a line item to raise, not a rounding error to absorb.
- Machine score diverges from human score by more than the agreed threshold → the pair goes to the Slack calibration thread. The output is either a corrected human, a corrected scorecard, or a corrected prompt.
- QA flags a knowledge gap → coaching assignment for humans, scorecard revision for the agent. Same finding, two different remediation paths, because a human who lacks context and a bot that lacks context fail differently.
Cost baseline
A 12-agent support org handling 6,000 conversations a month, with the agent taking 3,600 of them and 2,400 landing as verified resolutions:
- Zendesk QA via the Workforce Engagement Bundle: $50 × 12 seats × 12 months = $7,200/year, and workforce management comes with it.
- Assembled Pro instead, if you buy the capacity layer separately: $45 × 12 × 12 = $6,480/year plus an unpublished platform fee. Ask for that fee as a line item before comparing it to the bundle.
- Solidroad: quote-only. Budget against the Zendesk bundle as your reference point and make the vendor beat it on something other than price — the reason to buy it is that it does not report to the company sending your resolution invoice.
- n8n Pro: €60/month for 10,000 executions, roughly €720/year.
- Human review time: sampling 2% of 6,000 conversations is 120 reviews a month; at eight reviews an hour that is 15 hours, about a tenth of an FTE. This is the line everyone forgets and the only one that makes the scores mean anything.
Now the uncomfortable arithmetic. Those 2,400 verified resolutions cost $3,600 a month at the committed rate, $43,200 a year. If 5% of them are mis-tiered in the vendor’s favour, recovering all of it returns about $2,160 a year against an audit layer that costs $8,000 or more. The billing dispute does not pay for this stack. What pays for it is the reopen: a conversation billed as resolved that comes back is charged twice, once by the meter and once by the human who then handles it. Price the stack against reopen volume and CSAT recovery, and treat the invoice accuracy as the byproduct it is.
Variations and when to swap
- Zendesk QA when you are already all-Zendesk and the agent is somebody else’s. If Decagon or Fin is doing the resolving, Zendesk QA is grading a competitor’s bot rather than its employer’s, and the independence problem largely disappears. The $50 bundle also absorbs workforce management, which removes the Assembled line entirely.
- Solidroad when the agent and the helpdesk are the same vendor. Running Zendesk’s own AI agents and Zendesk QA means the same company owns the outcome definition, the verdict, the invoice and the report card. That is the case where paying a second vendor for the scorecard is worth the second contract.
- Drop Assembled under about 15 agents. A forecast that fits in a spreadsheet does not need a platform. Add it back the moment BPO vendors enter the mix, because vendor-managed headcount is where forecast error becomes someone else’s SLA.
- Swap n8n for native helpdesk triggers only while the sampling rule reads from one system. The moment the sample depends on account revenue from the CRM or reopen data from a second tool, keep n8n.
What this stack does not replace
- It does not decide what the agent is allowed to resolve. Intent scoping is the ticket deflection agent template.
- It does not deploy the agent or own the escalation path. That is the AI support agent stack, which this one audits.
- It does not retain customers. A resolved ticket is not a renewed contract — see the customer retention stack.
- It does not write the scorecard. Every layer here assumes somebody has decided what good looks like and written it down.
- It does not replace a manager reading conversations. It replaces a manager reading a convenience sample and calling it a QA program.
Watch-outs, each with a guard
- The grader and the graded can be the same company. Zendesk has owned Forethought since 26 March 2026 and sells the QA product that scores AI agents. Guard: keep the reconciliation ledger in a store the agent vendor cannot write to, or buy the scorer from a different company than the agent.
- AutoQA scores are model judgments, and they move when the model or the prompt moves. A scorecard revision can shift your average without a single agent behaving differently. Guard: freeze a gold set of 50 hand-scored conversations and re-run it after every scorecard or model change; a shift on the frozen set is a measurement change, not a performance change.
- 100% coverage removes the sampling excuse but not the sampling bias. Human review still drifts toward short, legible conversations. Guard: make the risk strata mandatory quotas in the n8n sampler — reopens, escalations, low CSAT, high-value accounts — and let the random draw fill only the remainder.
- The billing tier and the QA score arrive on different clocks. The tier lands 72 hours after the last message; the score lands immediately. Guard: key every row on conversation ID and reconcile monthly. Weekly reconciliation reads as a discrepancy when it is just latency.
- QA priced per human seat gets more expensive per unit of work as the agent absorbs volume. At $50 per agent per month, 12 seats and a rising bot share, you pay the same for a shrinking human workload. Guard: check whether your QA contract meters seats or conversations before the AI share passes 50%, and renegotiate at that crossing rather than at renewal.
- Coaching against a rubric nobody has read is a performance review in disguise. Guard: publish the scorecard to the team before it scores anyone, and treat the first month of scores as calibration data rather than performance data. When a specific escalation goes badly, run the escalation RCA.
Match rules
Right pick when: you run 2,000 conversations a month or more, at least part of your resolution volume is billed per outcome, humans and AI both touch the queue, and someone — a CFO, a board, a renewal committee — will eventually ask how you know the deflection number is real. It fits best at 10 to 50 support seats, where the meter is large enough to matter and there is no analytics function to build a scoring pipeline in-house.
Wrong pick when: you handle a few hundred conversations a month and a manager can read all of the hard ones, or your support is entirely human on flat per-seat pricing, where the audit layer costs more than the accuracy it buys. It is also the wrong pick if nobody will own the scorecard — an unowned rubric produces scores that everyone quietly ignores, and you will have bought a second dashboard to ignore alongside the vendor’s.
If you can only do one thing: hand-score 50 conversations against a written rubric and keep them. That gold set costs a day, needs no contract, and is the only way to find out whether any scorer you buy later — vendor-owned or not — agrees with you about what a resolution is. Pair it with CSAT on agent-handled conversations and you have the two numbers that survive a vendor changing its definitions.