Skip to content

8. Operations and Incident Response

Move from symptom to scoped evidence, hypothesis, safe action, verification and handoff without destructive guessing.

Level: Practitioner · Time: 60 minutes · Prerequisite: modules 1–7

stabilize → define impact → establish timeline → inspect full path
→ compare healthy peer → form hypothesis → make smallest safe change
→ verify service + redundancy → preserve evidence → hand off

Always distinguish data-plane health, control-plane/configuration state and management-plane access. A management timeout does not prove I/O loss; a surviving I/O path does not prove redundancy.

Complete the daily health-check runbook, then use three Academy incidents: port shutdown, inactive zoning and iSCSI authentication. Record commands, observations, rejected hypotheses and final verification.

  1. Why should you compare against a healthy peer or baseline?
  2. What verification is required after service is restored?
  3. When should an operator stop and escalate rather than continue changing state?

Incident authority, evidence retention and change approval are organization-specific. Never copy lab remediation directly into production without scope, backups and rollback.

Review the incident triage and change rollback runbooks, then attempt the Practical Assessment.