Using AI to strengthen system reliability
Agency
Federal program
Engineering teams can't improve what they can't measure. Maintaining reliable government services depends on quickly identifying recurring issues, understanding why they occur, and focusing engineering effort where it will have the greatest impact.
For one federal customer, Ad Hoc applied AI-assisted analysis to transform a labor-intensive, manual software reliability process into an automated workflow. The result was clearer operational insight, more efficient maintenance, and more time for engineers to focus on delivering new capabilities.
The challenge
The team regularly monitored the reliability of its software delivery process, but investigating failures had become increasingly time-consuming. Over years of supporting the platform, engineers had seen failures come from many different sources, including:
Identifying the root cause required manually reviewing logs across dozens of software delivery runs, making the investigation process both repetitive and difficult to scale.
Each month, the team manually reviewed dozens of software delivery runs to understand why failures occurred and determine which problems appeared most frequently during that reporting period. The process involved multiple manual steps and could take the better part of half a day for an engineer to complete.
Because the types and frequency of failures could change from month to month, engineers needed a clearer way to compare results across many runs and identify the problems having the greatest effect on reliability. That evidence would help the teams focus their effort on the highest-value improvements to their continuous integration and delivery processes and other automated systems.
Our approach
Rather than continuing to rely on manual investigations, Ad Hoc developed an AI-enabled workflow that automated both the collection and analysis of software delivery data.
The workflow automatically:
Gathers operational data.
Measures delivery performance.
Analyzes failed runs to identify the most frequent problems.
Synthesizes findings into structured recommendations to prioritize future improvements.
Throughout the process, engineers remained responsible for reviewing the results, validating recommendations, and determining which improvements to implement.
By turning operational data into actionable insights, we help engineers reduce maintenance effort, improve system reliability, and focus more time on delivering new capabilities.
Outcomes
The new workflow transformed a manual monthly process into an automated capability, significantly reducing recurring maintenance effort while improving visibility into system reliability.
Additionally, this work:
- Reduced the monthly analysis process from approximately half a day of manual work to about one hour, saving approximately 75% of an engineer's time.
- Identified the most common failure patterns across 69 software delivery runs, giving engineers clear evidence to prioritize improvements with the greatest impact.
- Created a structured record of recurring issues, making it easier to track trends, prioritize technical debt, and share knowledge across the engineering team.
- Reduced reliance on manual monitoring by creating a more reliable, automated process for collecting and analyzing operational data.
Rather than replacing engineering experts, AI helped teams navigate an increasingly complex set of failure scenarios by automating repetitive analysis and more quickly uncovering likely causes. That reduced the time spent searching through logs and allowed engineers to instead focus their expertise on fixing issues and improving the platform for the customer.