Managed Services
Two Years of Continuous Production Support for a Managed Services Platform
A managed services provider runs ServiceNow as the platform its own customers depend on, where inbound alert email drives automated case creation and routing correctness feeds contractual SLA clocks directly. We have owned that subsystem in production for over two years.
- 2+ years Continuous production support
- 15+ Root cause analyses delivered
- 3 to 1 Monitoring sources unified
The situation
A managed services provider delivers services to its own downstream customer base on ServiceNow. Inbound alert email from multiple monitoring sources drives automated case creation for many customers at once, and whether a case lands in the right queue at the right priority feeds contractual SLA clocks directly.
That makes correctness a commercial problem, not just a technical one. The complexity comes from three things stacked together: domain separation between customers, per-customer routing rules, and multiple upstream monitoring sources that each describe an alert differently.
What we do
This is an ongoing engagement rather than a project, spanning multiple ServiceNow platforms, each with its own dev, test, and production environments.
- Alert email routing framework. A custom scoped table with template-driven per-customer routing, built at the start of the engagement and still being extended today.
- Inbound email action architecture covering new, reply, and forward paths, so a customer replying to a notification updates the right case instead of opening a second one.
- Bi-directional ticket sync with a third-party system, including attachment handling.
- Platform work alongside it: automated test tooling, portal access and SSO, and approval workflow integrity.
The outcome
More than two years of continuous production support. Over that period we have delivered more than a dozen formal root cause analyses and a library of client-facing knowledge base articles, across multiple platforms and their full dev, test, and production environments.
The more useful measure is ownership rather than volume. The alert routing framework we built at the start of the engagement is the same one we were extending this afternoon. That is two years of continuous responsibility for a subsystem the client’s own SLAs depend on, not two years of unrelated tickets.
If your situation is similar, we’d be happy to talk.