ssm.ro Docs
Security, Infrastructure & OperationsIncident Management

Incident Response Procedure

Classification, detection, response, escalation, and post-incident analysis

The incident response procedure defines how security or availability events on the SSM.ro platform are detected, classified, handled, and analyzed.

Incident classification

SeverityDescriptionExamples
Critical (P1)Total service unavailability or confirmed data compromiseProduction outage, credential compromise, data breach
Major (P2)Significant degradation or high riskSharp increase in errors, partial unavailability
Minor (P3)Limited impact, no data affectedIsolated malfunctions, non-critical errors

Detection

Incidents are detected through the monitoring stack:

  • New Relic — NRQL alerts on error rates and Ping Monitor (Critical alert on endpoint unavailability)
  • Sentry — email alerts on new runtime exceptions or regressions
  • Provider status — monitoring status.heroku.com and AWS service health

Details: Metrics and Alerts.

Response steps

  1. Identification — the alert is received and confirmed
  2. Triage and classification — severity and impact are established
  3. Containment/limitation — impact is limited (e.g. rollback of a defective release, credential rotation)
  4. Remediation — the fix is applied (data restoration, patch, service restart)
  5. Notification — tenants and/or authorities are informed per the notification procedure
  6. Verification — return to normal and data integrity are confirmed
  7. Post-incident analysis — written review and corrective actions

Response to specific scenarios

ScenarioImmediate action
Defective releaseOne-click Heroku rollback (< 30 min)
Credential compromiseImmediate rotation of all secrets (Heroku config vars, AWS keys, Postmark/e-signature provider tokens); log review; notification if customer data is affected (< 2 hours)
Data lossRestoration from PITR/snapshot (Postgres) or versioning/CRR (S3)
Personal data breachActivation of the GDPR data breach notification procedure

Escalation

P1/P2 incidents are escalated to the responsible technical team and, where applicable, to management and the DPO (for personal data). Communication to customers is done via email and status updates.

Post-incident analysis (post-mortem)

For any event with production impact:

  1. Written review — causes, impact, timeline, actions taken
  2. Corrective actions — identified and tracked through to implementation
  3. Procedure updates — lessons learned integrated into SOPs

Major incidents with public impact are recorded in the Incident History.