Lab 01 of 09Auto-remediation simulator

Self-Healing, With Brakes

Break a small cluster and watch a remediation controller fix it — until its guardrails decide a human should.

  • Auto-remediation
  • Policy gate
  • Kubernetes model

Autoplay: scripted incident. Click anything to take over.

Cluster · 3 nodes × 4 pods

  • ok
  • hot
  • crash
  • restart
  • down
node-aReady
  • api-----v--running
  • api-----v--running
  • checkout-----v--running
  • search-----v--running
node-bReady
  • api-----v--running
  • checkout-----v--running
  • checkout-----v--running
  • search-----v--running
node-cReady
  • api-----v--running
  • checkout-----v--running
  • search-----v--running
  • search-----v--running
Request success100.0%
Error rate0.0%
last 60s · dashed = 99% SLO

Remediation controller

bounded · policy.yaml
  1. Detect
  2. Diagnose
  3. Propose
  4. Policy gate
  5. Act
  6. Verify
  7. Close
gate.deny →Page human
verify.fail →Stop & page

Watching. Nothing to do.

Budget / incident
0/3
Rollbacks / 6h
0/1
Restarts
0
Flapping
0
Gate denials
0
Pages
0
restart cooldown (120s)
apiclear
checkoutclear
searchclear

Event log

sim time · 4× real

    Built by Melih Kızmaz · runs entirely in your browser

    What you are looking at

    A toy cluster of three nodes running three services, and a remediation controller watching a single health signal. Each button breaks the cluster in a different way. With guardrails on, the controller walks the same state machine described in the article: detect, diagnose, propose one action, pass it through a policy gate, act once, then verify for 20 seconds that things are actually better before closing. A denied proposal or a failed verify never retries quietly; it pages a human.

    Why the brakes matter

    A killed pod is an honest failure: one restart fixes it, and both loops handle it (the naive one faster, since the verify hold is a real cost). Injected latency is a deceptive signal: the pods are fine, a dependency is not, so the guarded loop restarts once, sees no improvement, and stops. A bad deploy gets one gated rollback per window. A node failure has a node-sized blast radius, which policy says is propose-only. Turn guardrails off and replay the same faults to watch flapping, restart storms and a bad release spreading to every replica, all logged as "success" and none of them paging anyone.

    The policy numbers (3 actions per incident, 120-second restart cooldown, one rollback per window, 20-second verify hold) come from the lab repo behind Melih Kızmaz's write-up, where the naive loop restarted a healthy service 13 times in 106 seconds and the guarded one escalated to a human at 20.4 seconds. This page is a simulation that runs in your browser, not a real cluster.

    Read the long versionWhen auto-remediation makes incidents worse: blast radius and guardrailsNaylaLabs