TeknonOSTeknonOS
← The journal
Playbook · Jun 2026 · 6 min read

The handover checklist for AI systems

Documentation, evals, admin controls, monitoring, and training the client team needs before launch day.

DM
Dr. Dereck Mush, MD, MBA
Founder, TeknonOS
The handover checklist for AI systems

A handover is not a slide deck and a Zoom call. It is a set of artefacts and access rights that let your team keep the system running after the build team leaves. This is the checklist we run before we mark any engagement complete.

Why most handovers quietly fail

The average AI handover looks fine on the day it happens and falls apart six weeks later. Not because the system broke, but because nobody on the client side had the context, tools, or authority to keep it healthy when the small drifts started.

A drift is not a bug. A drift is the model provider updating a response format, a prompt threshold slipping out of the sweet spot, or a workflow starting to see inputs that were rare during discovery. Each drift is fixable in an hour if someone is watching. Unfixed, drifts compound into a system that has technically been live for six months but is producing garbage for the last three.

The checklist below is what we run before we mark any engagement complete. It is deliberately boring, because boring is what survives a change of personnel on either side.

Documentation that survives contact with reality

A one-page architecture diagram that a new hire can understand in five minutes. If the diagram needs a legend, redraw it. The purpose is not to impress a board; it is to let a support engineer who has never seen the system trace a bug from symptom to source without a synchronous call.

A runbook covering the top ten failure modes, each with a symptom, a check, and a fix. Written for the operator who will actually get paged, not the engineer who built it. If the runbook says check the logs without specifying where, what to search for, and how to interpret the result, it is not a runbook, it is a wish.

A change log with every prompt, threshold, and integration change, timestamped and reversible. This is the single most valuable artefact in the whole handover, because it is what lets you answer the question what changed yesterday when today's outputs look wrong.

A glossary of the terms unique to your business that appear in the prompts. Six months from now, a new operator will read a prompt that says apply the standard renewal treatment and have no idea what that means unless the glossary tells them.

Evals the client can run themselves

Golden dataset of at least two hundred examples for every critical path, with expected outputs and a pass criterion. The dataset needs to include the messy cases, not just the clean ones. If your evals only cover happy paths, they will pass the day the system breaks in the real world.

A one-click eval run that produces a diff against the previous version. If it takes an engineer to run, it will not get run. The whole point of the eval harness is to lower the cost of asking is this still working to something an operator can do on a Tuesday morning without booking a meeting.

A rollback path that reverts to the last passing version in under five minutes. Rollbacks that require a code deploy, a database migration, or a cache flush will be avoided in practice, which means the system will be run in a broken state instead. Design the rollback as a one-command operation from day one.

A quarterly commitment to expand the golden dataset with real production examples that surprised the operators. The dataset from launch day is a starting point, not a finished asset. If it does not grow, the evals slowly stop reflecting reality.

Admin controls and permissions

Named admin accounts on the client side, with the build team removed within thirty days of handover. Every extra day the build team keeps admin access is a day the client is not actually operating the system, and a day where an incident could pull the build team back in for the wrong reason.

Scoped permissions for each operator role, tested against real access patterns. A permissions model that looks correct on paper but blocks the operators from doing their job on day one is worse than no permissions at all, because it teaches everyone to work around it.

Cost ceilings and rate limits configured at the platform level, not just in code. Ceilings in code can be silently bypassed by a bad deploy. Ceilings at the platform level cannot.

Audit logs on every admin action, retained for at least a year. When something goes wrong, the first question is what changed and by whom, and the answer needs to be available without opening a support ticket with the platform vendor.

Monitoring, alerts, and on-call

Alerts on latency, error rate, cost, and eval regression, each routed to a real human channel. Not a shared inbox that nobody watches. A named channel with a named owner and an expectation of response within a stated window.

A weekly system health report emailed to the process owner automatically. The report includes throughput, cost, top failure modes, and any eval regressions caught during the week. Five minutes to read, one page long, no jargon.

An on-call rotation, even if it is only one person, so nobody has to hunt for who owns the system when it breaks. On-call for an AI system is usually much lighter than for a live production application, but the phone number needs to exist.

A quarterly incident review, even in quarters where nothing broke. The review looks at near-misses, drift indicators, and any changes to upstream systems. The purpose is to catch the slow-moving risks before they become fast-moving incidents.

Training the operators actually use

One forty-five minute training session per role, recorded, with a written companion sheet. The recording matters because half the operators will miss the live session, and the companion sheet matters because half the operators will not watch the recording either.

A thirty-day feedback loop where the operators can flag surprising outputs and get answers within a day. This is the loop that catches the subtle failure modes the evals will not. It also builds the trust that decides whether operators use the system as intended or route around it.

A named point of contact on the build team for the first quarter, so the handover has a graceful tail. Hard cutovers on day one leave the client stranded the moment anything unexpected happens. A graceful tail costs the build team a few hours a month and dramatically improves the odds that the system is still in daily use a year later.

A ninety-day review meeting at which the process owner walks the build team through what has changed in the business since launch, and both sides agree the next set of adjustments. This is the meeting that turns a one-off engagement into a durable relationship, or ends it cleanly if the business has moved past what we were hired to build.

For a deeper look at the operating discipline that makes all of this handover work stick, see our note on why AI readiness is mostly an operations problem.

DM
Written by
Dr. Dereck Mush, MD, MBA

Founder, TeknonOS · Physician-operator writing on AI systems for real businesses. If any of this rings true for your business, connect on LinkedIn or book a call and we will walk through it with you.

Follow on LinkedIn
/ Ready to build?

Book your free 30-minute scope call.

Pick a time below. We'll map what you already run and show you the system that ties it together. No commitment, no pitch deck.

No commitment
required
Reply within
24 hours
Serving clients
worldwide