Rollback Strategies - What Works and What Doesn't
Rollback Reality¶
Rollback sounds simple, but network rollback is rarely transactional.
Challenges:
- State changed outside the workflow during execution
- Protocol convergence side effects are time-dependent
- "Undo" commands may not restore prior behaviour exactly
- Traffic patterns and dependencies may have shifted
Rollback is a strategy portfolio, not a single button.
Practical Rollback Methods¶
Common patterns:
- Configuration snapshot restore
- Reverse-change command sets
- Feature-level disable to contain impact
- Route-policy or path steering fallback
- Human-guided recovery runbook
Each method has different speed, certainty, and risk.
Decision Matrix¶
Choose rollback path by context:
- Fast containment needed: disable or isolate impacted feature
- Known deterministic change: reverse-change may be sufficient
- Broad uncertain impact: restore snapshot with validation gates
- High ambiguity: pause automation and switch to human-led recovery
Why Automatic Rollback Can Be Unsafe¶
Auto-rollback can worsen incidents when:
- Root cause is unknown
- Rollback target is stale
- Partial changes already improved stability
- Multiple workflows interact on the same devices
Automatic rollback should be policy-bounded and evidence-based.
Production Checklist¶
- Rollback strategy is defined before rollout starts
- Pre-change snapshots are captured and validated
- Rollback triggers are explicit and measurable
- Post-rollback verification is mandatory
- Human takeover criteria are documented
Anti-Patterns¶
- Assuming rollback always restores previous behaviour
- No pre-change snapshot strategy
- Triggering rollback on any warning signal
- Running rollback and forward remediation concurrently
Key Takeaway¶
Safe rollback is controlled recovery under uncertainty. Prefer recoverability and containment over blind reversion.
Between the Lines
Kintsugi is the Japanese practice of repairing broken pottery with gold, so the break becomes the most visible part of the object rather than the thing you hide. The bowl is not pretending it was never dropped.
A rollback is not a confession either. The systems worth trusting are the ones that have been broken, recovered and documented — a considerably more forgiving standard than most engineers are willing to apply to themselves.
Continue the Series¶
- Series Index: Production-Grade Network Automation Principles
- Previous: Part 7 - Designing Automation That Can Safely Fail
- Next: Part 9 - Separating Read and Write Phases in Automation Workflows
Need help applying this in a live Cisco environment?
This guide is part of the Nautomation Prime Foundation and stays free to read, share, and reuse. If you want the pattern implemented, governed, or adapted for your estate, that is paid engineering work — start a discovery conversation or review how Nautomation Prime delivers engagements. If you are a registered UK charity or CIC, there is a free and low-cost route instead.