Delivery
From AI pilot to production: a four-week delivery framework
A focused delivery sequence for moving one bounded agent from process map to monitored production use.
- Published
- 17 September 2026
- Reading time
- 6 min read
- Topic
- Delivery
Organise discovery, connection, testing and release around one operational job.
Why pilots stall
Pilots often optimise for a compelling demonstration rather than a supportable operation. They use curated data, broad developer credentials and hand-picked examples, then encounter ownership, integration and risk questions too late.
A production path begins with the job and its controls. Four weeks is suitable for a bounded use case with available interfaces and empowered owners; complex procurement, inaccessible data or high-risk decisions may require more preparation.
Week one: define the operating contract
Map the trigger, inputs, decisions, actions, outputs, owner and exceptions. Agree what the agent will not do. Capture a baseline for time, quality and delay, using ranges if needed.
Review real cases with the people doing the work. Establish access paths and select a representative evaluation set before implementation begins.
- One-page job charter
- System and data map
- Authority boundaries
- Success measures
- Representative test cases
Week two: connect and make the path visible
Build the smallest end-to-end route through real interfaces. Establish the agent identity, permissions, schemas, logging and a safe destination such as a draft queue.
Keep decisions observable. Operators should be able to see each stage and distinguish data, integration and model failures. This foundation matters more than polishing a demonstration.
Week three: evaluate and govern
Run normal, edge, missing-data and adversarial cases. Compare results with agreed criteria and inspect failure clusters. Add deterministic validation where rules can catch an error more reliably than a model.
Exercise approvals, escalation, retry and shutdown. Confirm privacy, security and risk owners understand the evidence they will receive after launch.
Week four: release with an owner
Start with a controlled audience or workload and monitor every run. Train operators on exceptions and define response expectations. Record the released versions of prompts, models, tools and schemas.
Handover includes dashboards, runbooks and a review cadence. The build is complete when the team can operate it, not when the agent produces its first good output.
NEXT STEP
Put this guide into practice
See how this applies to a defined agent build, including its systems, controls and operating owner.
See our four-week build process