AI incident response tools risk eroding engineer expertise

Original: AI handles incidents, engineers lose touch with their systems

Why This Matters

AI-driven automation in SRE may widen the skill gap for complex incidents as routine practice disappears.

As AI-powered incident response tools handle routine system failures automatically, SRE professionals warn that engineers are losing hands-on troubleshooting experience. When complex, novel incidents arise that AI cannot resolve, human responders may lack the practice needed to respond effectively.

Sylvain Kalache, a former LinkedIn SRE and current Rootly employee, argues that while AI-assisted incident response tools—often called 'AI SREs'—are effective at resolving routine failures autonomously, they risk degrading human engineers' operational intuition over time. These tools inspect alerts, query telemetry, correlate deployments, and implement fixes without waking on-call staff. However, routine incidents are traditionally how engineers develop a feel for how systems behave and fail. Human-factors researcher Lisanne Bainbridge identified this dynamic in her 1983 paper 'The Ironies of Automation,' noting that automation reduces operators' chances to practice routine tasks while leaving them responsible for novel, abnormal situations requiring greater skill. Kalache predicts average MTTR will fall for most incidents, but resolution times for complex cases will rise as engineers lose familiarity with their systems. He draws a parallel to aviation: pilots rarely encounter engine failures in real flights—fewer than one per 100,000 engine flight hours—yet must train for them rigorously every six months under FAA rules. The 2015 TransAsia Airways Flight 235 crash, where a crew misidentified an engine fault and the aircraft stalled within 117 seconds of the first warning, illustrates the cost of inadequate rare-event preparation. Kalache advocates for software incident simulators as a solution, citing Rootly's partnership with Uptime Labs, which places engineers in realistic simulated outage scenarios using actual observability tools and LLM-powered stakeholder interactions in Slack.

Source

sylvainkalache.com — Read original →