Skip to main content

Posts

Three Truths About AI SRE

The current enthusiasm for AI in Site Reliability Engineering (AI SRE) is well-founded. We are seeing incredible advancements in agents that can ingest alerts, parse logs, and propose rapid solutions to outages. They are becoming increasingly confident when it comes to suggesting bug fixes, patching vulnerabilities, and helping coordinate around incidents. But there is a dangerous blind spot that every engineer knows about, but many businesses are ignoring because they are tempted by the increase in velocity…the simple truth that recovery is not the same as reliability. While your AI SRE excels at diagnosing an incident, it often lacks the context, tools, and proactive mindset to prevent that incident from recurring in the first place. If your strategy relies solely on reactive recovery, you aren’t building a more resilient system, you are simply building a faster way to apply duct tape. Here Are Three Critical Truths About Your AI SRE Solution 1. Fixing a Symptom Instead of ...
Recent posts

HackerOne Report Surfaces Massive Increase in Vulnerability Backlogs

An analysis of the backlog of vulnerabilities published today by HackerOne finds that while the rate at which issues are being resolved has increased 54% in the past year there has also been a 131% over the last two years in the total number of known validated issues that have not been resolved. On average, organizations have cut average resolution time from 135 days to 62 days, but the pace at which vulnerabilities are being discovered in the artificial intelligence (AI) era continues to increase. In fact, a separate survey of 111 security leaders finds 70% are seeing validated findings being added to their backlogs faster than they are being remediated. HackerOne CEO Kara Sprague said that as researchers make use of AI to discover vulnerabilities, it’s apparent that DevSecOps teams are starting to be overwhelmed. As a result, it’s clear that many of those teams now need to start applying AI to remediate vulnerabilities faster by, for example, formally addressing best vulnerability ...

DevOps Has Always Been Hard to Define. Does a Standard Help?

Ask a room full of technology professionals to define DevOps and you can still start an argument. Is it a culture? A collection of engineering practices? An organizational model? Automation? A job description? Somewhere along the way, we managed to make it all of those things, depending on who was talking and what they were selling. That ambiguity has been with DevOps almost from the beginning. It helped the movement travel across organizations with very different needs. It also allowed companies to rename an infrastructure team, buy some pipeline tools and declare the transformation complete. Now PeopleCert and DevOps Institute have published The DevOps Standard , and the familiar questions have returned. Can you standardize something people still struggle to define? Does a movement need a standard? And what happens when that standard arrives just as AI agents are changing who, or what, does the work? I have a connection to this discussion. I am the founder and editor-in-chief of ...

Federated Query at Petabyte Scale: A Deployment Pattern for a Governed AI-Agent Data Layer

Enterprises with data spread across ERP, CRM, cloud warehouses and observability platforms face a recurring problem: As data volume grows into the tens of petabytes, traditional ETL-and-centralize approaches produce unacceptable query latency and duplicated infrastructure cost. This case study describes an automated deployment framework for a distributed SQL query engine (federated query architecture) that queries data at its source rather than moving it, built and operated as the lead infrastructure engineer at a health care technology company. It further describes how that same data layer was extended into a governed, agent-accessible interface for enterprise GenAI, using a semantic layer and tool-exposure protocol (MCP-style) to connect a conversational AI interface to real-time business data. This architecture fully preserves the role-based access control and audit requirements appropriate for regulated (health care-adjacent) data. The contribution is a reusable deployment and gov...

Getting to Reliable AI-Driven Development

To truly transform software delivery with AI, organizations must embrace a spec-driven approach. By grounding AI in clear specifications, this approach prevents AI from generating hallucinations while optimizing token costs by 8X–12%. The use of AI in software development is now practically universal. According to the 2026 Software Lifecycle Engineering Decision Maker Survey from Futurum, 97% of surveyed organizations are already using or planning to use AI for software development, with more than three-quarters actively using AI in development workflows. However, there’s a world of difference between using AI for development and maximizing its value. Currently, most conversations about AI in software delivery focus on things like faster autocomplete or chatbot debugging. Many organizations struggle to move beyond individual productivity improvements and create repeatable enterprise workflows that yield durable results. In fact, a Stack Overflow developer survey found that only 33...

QA Automation Frameworks for Enterprise Applications

Modern enterprise applications demand testing strategies that can keep pace with continuous delivery, distributed architectures, and increasingly complex deployment pipelines. Traditional manual testing approaches often become a bottleneck as engineering teams scale their applications and release software more frequently. This article explores the essential building blocks of an enterprise-grade QA automation framework, focusing on maintainability, scalability, reliability, and integration with modern DevOps workflows. Why Enterprise QA Requires a Different Approach Unlike small applications, enterprise systems typically consist of multiple services, APIs, web applications, background workers, and third-party integrations. A successful automation strategy must support:  Repeatable and deterministic test execution  Parallel execution across multiple environments  Continuous validation within CI/CD pipelines  Cross-browser and cross-platform compatibility  Reliable API and e...

AI Is Breaking Developers’ Traditional Game Testing Playbook

Game testing has always been an unpredictability problem for developers. Players move in the wrong direction, trigger events out of order, or combine mechanics no designer anticipated. Even so, developers could usually define the rules governing the world and check whether the game followed them. AI is making that harder. NPCs now react to changing conditions, and animation systems select behaviors based on surroundings rather than a fixed sequence. Take-Two Interactive, Rockstar Games’ parent company, holds a patent for a virtual character locomotion system that uses modular animation building blocks and selection criteria to control how characters move through a 3D environment, one sign of character behavior becoming more responsive to context instead of scripted. That can make for a more convincing game, but it leaves developers wondering: How do you test software when you can’t list every way it might behave? The Test Matrix is Getting Harder to Define Games already...