What's the Answer to "So, When Can This Go Live"?
Three hours into the go-live review meeting, the customer closes their laptop and asks the plainest …
Good fences make good neighbours. — Robert Frost, "Mending Wall" (1914)
Three hours into the go-live review meeting, the customer closes their laptop and asks the plainest question in the room: "So — when can this go live?" The room goes quiet for ten seconds. The vendor wants to say "the results are already good"; the business side wants to ask "who's on the hook if it breaks" — both sentences sit right at the tip of the tongue, and neither one gets said.
This scene keeps replaying across the industry because "going live" was never just a technical question — it's a two-way contract nobody has ever drafted: the vendor proves it hit the pre-agreed metrics, and the customer takes on operational responsibility against a pre-agreed checklist.
With nobody drafting the contract, both sentences stay stuck at the tip of the tongue, indefinitely.
This chapter is about how to draft that contract.
I. What Is Actually Being Asked, When Someone Asks About Going Live
Start by standing on the other side of the table. As the person responsible for bringing an AI system into the enterprise, the list of questions running through your head at the review meeting usually looks like this: what if the data connection fails? Who owns the permissions? Who's responsible if it gets an answer wrong? How do we trace accountability if something goes wrong? And when does this actually count as ready to go live?
Run down that list and one thing jumps out: every item is an acceptance question, and not one of them is a model question. The model's capability was already settled on demo day. The pressure point was never the model — it's acceptance.
"So when can this go live" sounds like it's asking for a date. What it's actually asking is: who decides, against what standard, that this thing has succeeded.
Without answers to those three things, any date offered is a guess pulled out of thin air. With answers to all three, a date is just the natural, final signature.
So the standard answer isn't a sentence — it's a contract made of three documents: an outcome acceptance sheet (metric + baseline + who decides), an operational readiness checklist (monitoring / degradation / rollback / on-call), and a responsibility handoff table (who's accountable for each incident scenario). One document handles "is it worth going live," one handles "do we dare go live," and one handles "who do we call when something breaks."
II. The Outcome Acceptance Sheet: Five Elements, and Missing Even One Wrecks It
"The results are good" can't be accepted against anything. A metric that can actually be accepted has to spell out five elements in full: baseline value (the level of human performance before going live — the commitment from Chapter 6's filter three, and the numbers measured by Chapter 7's Minimum Viable Deployment, land in the contract here), target value (the level to be reached after going live), measurement window (tracked weekly or monthly), denominator definition (who counts, who doesn't), and who decides (who has final say) (see Figure 8-1).

Figure 8-1: Going live means clearing three gates at once — outcomes, operations, and responsibility — and ramp-up validation still follows.
Of the five elements, the denominator definition is where the landmines hide most often — the same "90% resolution rate" can mean the denominator is every conversation, or only conversations routed to the AI channel, or only conversations the AI actually attempted to answer. Those three denominators produce numbers a full tier apart.
This is exactly where metric-definition discipline lives. The core value metrics for enterprise deployments — resolution rate, deflection rate — are precisely the metrics where every vendor's definition is the messiest. Take ServiceNow's official definition as an example: a "deflection" only counts if two conditions hold at once — no ticket filed within 24 hours of the interaction, and the user subsequently engaged positively (continued using the product, for instance). So before quoting any vendor's resolution rate or deflection rate, ask three things first: what's the denominator, how long is the judgment window, and who's the judge. One line nails this discipline exactly: "an 86% with a murky definition is worth less than a 51% with a clean one."1
What does a metric with all five elements spelled out actually look like? The insurance-industry company mentioned earlier gives a good template: the quality standard they agreed with the customer was 99.5% — translated into something checkable, that's "getting it right 995 times out of 1000 every week"2. What makes this number good comes down to three things: it's countable (995 out of 1000), it has a window (every week), and it can be double-checked (the customer can count it themselves). Countable, windowed, double-checkable — every metric on an acceptance sheet should survive all three questions.
The acceptance sheet has a negotiating function too: freeze the metric ahead of time, and go-live day loses the "this month's data is a special case" excuse. Every objection to the definition gets pushed back to the cheapest possible moment to raise it — before anything has even started running, when fixing it means changing one line of text.
III. The Operational Readiness Checklist: Who's Watching When the Alert Fires
Hitting the outcome target only answers "is it worth going live." The second document answers "do we dare go live." Four questions, one at a time:
Is anyone watching the alerts? When an alert fires, does it map to a name and an on-call schedule — not fall into a junk folder, not sink into a group chat. How the monitoring system is built is a later engineering topic; here it's only the acceptance-level question: when the alert fires, who sees it?
Is there a working degradation path, and has it been rehearsed? Including the plainest version: an explicit fallback to Excel and manual work. A lot of teams feel like writing "fall back to manual" into the plan is embarrassing, as if putting it on paper is admitting AI doesn't work. It's exactly the opposite: a plan willing to write down its own rollback is the plan that dares to go live — it admits the system will break, and it's already agreed in advance what happens when it does. This point gets repeated constantly in China: it's the item on the acceptance checklist most often skipped, and the one people are most grateful for when something actually breaks. A degradation path isn't a fig leaf for failure — it's a safety net, built on purpose.
Does a runbook exist, and has anyone on the customer side actually read it? The incident-response steps have to be written down, rehearsed once, and the customer's own point of contact has to hold a copy and have read it. A runbook that only lives on the vendor's laptop is the same as no runbook at all.
Who's on the on-call rotation? For the first few weeks after going live, who on the vendor's side is on call, and what's the committed response time — write it down in black and white. It's the least conspicuous item across all three documents, and the first thing anyone asks when something breaks.
The cautionary tale is short but it stings. An FDE at an AI startup noted in a weekly diary: an upstream provider (Cloudflare) went down, there was no contingency plan, and firefighting had to happen live — every one of their customers' systems felt it at once3. An upstream outage was never within your control. "No contingency plan" was. That's exactly what the operational readiness checklist is there to catch.
IV. The Responsibility Handoff Table: Every Way to Fail Gets a Name
The third document is the one most often missing, and the one that most tests both sides' good faith. The method is to list out every way the system could fail, one by one, and write down the first responder and the escalation path (how long, escalating to whom) for each. Here's an example:
| What Went Wrong | First Responsible Party | Escalation Path |
|---|---|---|
| Model output error causing business loss | Who judges it, who blocks it, who pays for it | To both sides' leads within 1 hour |
| Data source failure / upstream outage | The data-side point of contact | To vendor on-call within 4 hours |
| Incident caused by user error | Customer-side system administrator | To a joint post-mortem, same day |
| Cost overrun | The platform provider | To the project lead, in the weekly report |
Why is this table worth drawing at all? Back to the line that opened this chapter. When Frost wrote "good fences make good neighbours," he wasn't talking about suspicion — he was talking about drawing the boundary before anything goes wrong. Put the fence up in advance, and the neighbors stay friends. Wait until the sheep have trampled the garden to draw the line, and no matter how you draw it, it damages the relationship. A responsibility handoff table is the fence you put up before going live: whose fault it is when the model breaks, whose fault it is when the data breaks, whose fault it is when someone fat-fingers an action — write it down in black and white, and when something breaks, you find the right person off the list without burning the relationship. It's the front-end slice of the full handoff covered in Chapter 17 — what's handed off at go-live is "first response"; long-term operational responsibility is a bigger table of its own.
With all three documents in hand, run them past the review table and you'll find out exactly how solid they are. The following review-meeting minutes aren't a real case, but every item in them comes straight off the checklists above:
Go-Live Review Minutes (illustrative simulation) Passed: all five elements of the acceptance sheet are complete; the "who decides" field names the actual business owner, not "the business side" in the abstract; the rollback line for phase three of the gradual rollout is quantified. Stuck: the degradation path — "fall back to manual" — is written down, but has never been rehearsed. A rehearsal date got scheduled on the spot before sign-off proceeded. The on-call roster only covers the first two weeks after go-live; after that, it's blank. Disputed: responsibility for incidents caused by user error. The vendor argues it's entirely the customer-side administrator's fault; the customer argues the interface should have been designed to prevent the error in the first place. After a standoff, it went in as "shared responsibility, plus a joint post-mortem rule." One-third passed, one-third stuck, one-third disputed — that's roughly what a real review meeting's minutes look like.
One last look at the two ways this degrades when responsibility is left hanging — both happen exactly where nobody drafted a contract. One is "run it free for three months first" — responsibility has no owner; everyone's happy while it runs well, and the relationship blows up the first time something breaks, because no page anywhere ever said whose fault it was. The other is "the results are great, just keep using it" — the vendor is effectively refusing the handoff, and the system stays permanently stuck at "trial," never growing into a production system. The two degrade in opposite directions, but the root cause is the same: an absent contract.
V. Both Sides Have Homework to Do
This contract shouldn't wait to be remembered until the go-live review. Both sides have homework due earlier than that.
The customer's self-check list. Flip the vendor's own customer-screening filter over, and it becomes the customer's self-check list: is your project, internally, the kind of "no-man's-land" a vendor dreads — pushed down from a group-level directive, led by an IT department, with no business department willing to put its name on it? Does your procurement process structurally force "look, but don't touch" — demanding the vendor prove its capability first while handing over exactly zero real data? Does what you wrote into the RFP amount to a universe-scale demand — "AI-empower the entire group"? The verdict is straightforward: that last kind of requirement scares off exactly the vendors who know what they're doing, and attracts exactly the ones bold enough to bluff. If any one of these three is a "yes," fix your own house before you talk about going live.
The vendor's homework. A good proposal is written backward: the first layer has to be the business outcome, stated in one paragraph, in the customer's own language — "cut average anti-money-laundering investigation time from 4 hours to 15 minutes within eight weeks." The second layer is the value-validation path — acceptance metrics, measurement method, baseline data, an exit mechanism for each phase — and the signal it sends is "we're not afraid to be checked." Only the third layer is the delivery method, spelling out what the customer needs to contribute. Last comes risk and countermeasures — "here's how we've thought this could die"1. That second layer is the shadow this chapter's three-piece set casts back into the pre-sales stage: the acceptance sheet isn't homework you make up at go-live — it's a commitment planted on page one of the proposal. This playbook has a precedent a full decade older than the AI era: a data-analytics company has done heavyweight pre-sales implementation during free trials since 2013, on exactly one principle — the demo is the proof of concept. Always ask a prospective customer for a real dataset. Never perform on a carefully prepared sample.
The reality at the negotiating table is: when the customer side doesn't understand metrics, they tend to fall back on one of two postures — "we need 100% accuracy" (which turns the acceptance sheet into an impossible task) or "just run it free for three months" (which leaves responsibility hanging). The fix for both postures is the same sentence: turn "results" into a measurable number. This is also the vendor's own sales logic — trade the acceptance sheet for trust: the more clearly the acceptance terms are written, the more confidently the customer signs, and the faster the deal closes. In an enterprise market that's been burned by over-promising before, the willingness to be checked is the scarcest form of sincerity there is.
VI. Going Live Is a Ramp, Not a Switch
The contract's signed — but going live still isn't a switch you flip. It's more like climbing a ramp: release traffic or user groups in stages, and every stage has a rollback line — if the metric falls below the agreed value, drop back a stage, find out why, then try again. Go-live day isn't the finish line. It's the first step of the ramp.
Buried inside the length of that ramp is exactly the stretch of time most schedules leave out. A bank-project post-mortem circulating in the Chinese community gives a strikingly lopsided pair of numbers: the technical work was done in six to eight weeks, and the pilot phase plus building trust took another four months4. The first half is progress the vendor controls; the second half is time the customer's organization needs — and most schedules only draw the first half. So the full answer to "when can this go live" has to set aside an explicit time budget for building trust: the stretch it takes front-line users to go from "afraid to use it" to "can't work without it" is exactly as real as development time, and how to earn it is the subject of the chapter on adoption.
Closing out this stretch of the book's road: an open-source methodology document once marked reference paces for this kind of deployment — the Minimum Viable Deployment validated in weeks (2–6 weeks), the full deployment measured in months (1–4 months), and one warning attached — time-to-value that keeps getting longer is the first sign that the methodology or the platform itself is broken1. Check it against the three-piece set: the acceptance sheet is signed, the readiness checklist has passed review, the responsibility table is up on the wall, and the ramp has climbed its last stage — only then can the words "ready to go live" actually be said. And what those words really announce isn't that the system succeeded. It's that the contract just took effect.
But the contract taking effect is only the ticket to get in the door. Turning "a validation that runs" into "a production system you can actually trust" still means getting past eight walls — data, permissions, evaluation, approval, cost... Every one of them is real, and every one of them can be climbed. That's what Part Three is about.
Footnotes
-
Fan Bing, Forward Deployed Engineer (FDE) (XDash open-source book) ↩ ↩2 ↩3
-
"Inside an Applied AI Company" (Pace long-form post) ↩
-
Hacker News, "A Week in the Life of a Forward Deployed Engineer" (10 customers, 50 hours) ↩
-
"FDE Engineer 101: Bridging the Gap Between AI and Business" (Chinese-language compilation video) ↩