iwantcoding.com
🔥 Daily 👥 Rooms 🏆 Top Log in Sign up

12.3 Incident, Problem, Change

Incident, Problem, and Change Management are three distinct ITSM practices. Incident restores service; Problem finds root cause; Change controls how production evolves.

12.3 Incident, Problem, and Change Management

Three practices, three goals

PracticeGoal
IncidentRestore service quickly
ProblemEliminate recurrence
ChangeKeep changes safe and reversible

Incident lifecycle

StepActivity
DetectMonitoring + user reports
LogTicket with severity + impact
ClassifyCategory + assignment
DiagnoseInitial triage
ResolveRestore service (workaround OK)
CloseConfirm with user; capture learning

Severity examples

SEVDefinitionCommunication
1Total outage; war roomCEO informed; status page
2Major degradationSenior engineers
3Minor; standard teamInternal channels
4Request or low impactTicket only

Change categories

CategoryRiskApproval path
StandardLowPre-approved
NormalMediumCAB review
EmergencyCriticalBypass + post-review

Blameless post-mortem template

SectionContent
What happened (timeline)Hour-by-hour or minute-by-minute
ImpactUsers, money, data
Root causeTechnical + organisational
What we did wellReinforce these
What we should improveHonest list
Action itemsOwner + date
Public summaryIf customers affected

Worked example - bank payment outage

ItemDetail
IncidentPayment fails 22 minutes; SEV2
ActionFailover to standby; queue drains in 8 min
ProblemRoot cause: connection pool exhaustion
FixPool sizing + circuit breaker via Normal change
ImproveChaos test; alert on pool > 80 %
Mentor’s tip: Incident restores; Problem prevents; Change controls. Blameless post-mortems convert pain into improvement. Tier your changes; let standard be standard, normal be normal, emergency be rare.

Discussion

Loading…