Why AI development creates a reliability blind spot for humans, and what to do about it

Part 1 of The Intellyx Building Agentic Resilience Series by Jason English, for Gremlin

Application development and operations teams are adopting AI coding tools at an exponentially increasing rate, from a rounding error of 6% of code output by AI in 2023, to as much as 51%-75% for a majority of enterprises.

Agents are making pull requests (PRs) faster than human developers could have ever dreamed. As features are pushed to market faster, there’s a sharp increase in production incidents, with 80% of development shops specifically tracing production outages to AI.

If AI coding agents were expected to deliver software 10 times faster, those results haven’t shown up in the industry surveys, as another useful benchmark notes only a 20% boost in successful PR velocity, with change failure rates up 30%.

Looking across these reports from leading software quality, observability, and security testing firms, we can see an industry-wide drag on the expected AI productivity boost for software development, as the rate of failures increases linearly with the increased volume of AI code delivered, and the time spent responding to production outages increases.

But despite the unknown risks of an adolescent market, companies have already opened the garage doors for AI coding agents and handed them the keys to the Ferrari.

Clearly we are missing something: the ability to reliably keep AI-driven software releases on track and avoid costly incidents that will slow our delivery progress in the future.

Looking for root cause through observability and SRE lenses

The software industry has made great strides in advancing observability, which looks at streams of system telemetry such as logs and metrics to understand the inner workings of software as it runs, and comparing real-time data to historical data to spot anomalies and alert human teams about concerning trends.

For years, the holy grail of observability was using the golden signals of latency, traffic, errors, and saturation to discover the root cause of any perceived failure. A new role of SRE (site reliability engineering) emerged, consisting of highly skilled individuals who were very well-versed in the architectural details of a system, so they can drop into an on-call incident and hopefully use an esoteric suite of analytics and tools to find that root cause and implement a fix.

One early shortcoming of observability was being able to “cause the cause” and reproduce the conditions under which a failure occurs. Observability and testing vendors built real user monitoring (RUM) and synthetic data to help reproduce these failure conditions in pre-production environments.

A new class of “AI SRE” solutions has now entered the market focused on the detection, triage, and AI-assisted remediation of issues, so teams can respond to them as quickly as possible. Every team wants to be responsive and reduce all the MTTs…

Read the whole story on Gremlin.com here: https://www.gremlin.com/blog/why-ai-development-creates-a-reliability-blind-spot-for-humans-and-what-to-do-about-it

 

SHARE THIS:

Principal Analyst & CMO, Intellyx. Twitter: @bluefug