Skip to main content

Posts

The Agent Failure Your Approval Gate Can’t Catch

An agent that does the wrong thing leaves evidence — a bad diff, a malformed record, a customer complaint. An agent that does nothing and reports success leaves a green checkmark, and green is the one signal your incident process is built to trust. Worry more about the second failure, because everything a DevOps team owns for catching failure — exit codes, run status, alert rules, the dashboard itself — is instrumentation built for software that fails loudly. None of it fires when the failure is an absence. Andrew Filev made the strongest version of the reliability argument in these pages in July ( https://devops.com/reliability-comes-from-the-system-not-the-agent/ ): reliability is a property of the system, not the agent. “Reliability has rarely come from any single component in isolation,” he wrote. “It comes from how systems handle failure.” Aviation does not assume perfect pilots and hospitals do not assume perfect surgeons; both wrap imperfect actors in app...
Recent posts

OllyGarden Extends AI Agent to Identify Instrumentation Gaps

OllyGarden this week revealed it has extended the capabilities of its artificial intelligence (AI) for optimizing the collection of telemetry data to now also discover where no existing instrumentation exists and what instrumentation should be applied to address that gap . Fresh off raising an additional $4 million in funding, OllyGarden founder Juraci Paixão Kröhling said a Minimum Viable Instrumentation (MVI) capability that has been added to the Rose AI agent the company previously developed makes it possible to identify the most relevant sources of telemetry data that should be instrumented using OpenTelemetry, an open-source framework for collecting that data that is being advanced under the auspices of the Cloud Native Computing Foundation (CNCF). That capability extends the scope of an AI agent that was created to help DevOps teams reduce the amount of telemetry data they need to collect and store by identifying which logs, traces and metrics are the most relevant. The overall...

IBM and Red Hat Disclose Discovery of More Than 400 Java Vulnerabilities

IBM and Red Hat this week reported they have identified and remediated more than 400 previously unknown vulnerabilities in Java libraries since launching a Lightwell initiative earlier this year. Additionally, Lightwell Clearinghouse, a program that enables IT organizations to submit specific open source software dependencies for priority review and remediation, is now generally available. Ben Bread, a senior principal product manager for Red Hat, said the 400 unknown vulnerabilities represent twice the number that was expected to be uncovered and there will undoubtedly be more to come as artificial intelligence (AI) tools are used to analyze more legacy code. Additionally, DevSecOps teams should expect a similar number of vulnerabilities to be discovered in libraries created using other programming languages, he added. While IT teams are contracting with IBM and Red Hat to fix vulnerable code in their IT environments, the code fixes developed are being contributed back to upstrea...

Attaching Evidence to Alerts: an Enrichment Sidecar for Alertmanager

Every on-call engineer knows the ritual. A page lands in Slack: the 5xx rate on a service is above the threshold. The message tells you that something is wrong and nothing about what. So you open Kibana or Loki in another tab, set the time window to the last few minutes, filter by the service, and start reading. In my team this took two or three minutes on every alert, and these are the worst minutes of the incident, because nothing is being diagnosed yet. The obvious request is to put the last few error lines into the alert message itself. I tried to do this inside Alertmanager, and it cannot be done there. It helps to understand why before building anything around it. Why Alertmanager Cannot Do It Alertmanager renders notifications through Go templates, and a template can only use what is already attached to the alert: the labels and the annotations which Prometheus put there when the rule fired. There is no mechanism which would call Elasticsearch, Loki or any other HTTP API at t...

Secrets Sprawl and Rotation: A Practical Vault Management Playbook

In 2025 alone, security firm GitGuardian detected 28.65 million new hardcoded secrets exposed in public GitHub commits — a 34% increase year over year and the largest single-year jump the company has recorded in five years of publishing its annual State of Secrets Sprawl report. Since 2021, that number has grown 152%, far outpacing the 98% growth in GitHub’s active developer base over the same period. Secrets sprawl isn’t a hypothetical DevOps hygiene problem anymore. It’s accelerating faster than most security teams can respond to it, and AI-assisted development is a large part of why. The financial stakes keep climbing alongside it. IBM’s newly released 2026 Cost of a Data Breach Report puts the global average cost of a breach at $4.99 million, a 12% increase over the prior year and a new record high. In IBM’s 2025 edition , breaches where compromised credentials were the initial access vector carried a $4.67 million average cost on their own, and organizations took an average o...

Coding Agents Broke Git’s Scaling Math. GitHub Is Rebuilding to Keep Up

For most of Git’s history, a commit marked a decision. A developer finished a change, checked it, and pushed. Hosting platforms were built around that rhythm. AI coding agents don’t work that way. “An agent in a tight loop commits or checkpoints after nearly every action,” GitHub engineer Brian Celenza wrote in a GitHub blog post explaining why the company is rebuilding the Git infrastructure behind its platform. The numbers show how fast the load has shifted. GitHub recorded 7.38 billion commits in September 2026, more than five times the total a year earlier. Monthly Git activity climbed from 218.2 billion events in September 2025 to 473.3 billion in August 2026. Pushes grew 4.9x year over year, from 0.69 billion to 3.35 billion per month. Pull request merges are up nearly 4x. GitHub Actions ran 3.26 billion times in September, also 4x the prior year. The busiest single repository handled roughly 1 billion requests in August. Nobody planned capacity for curv...

AI Is Taking on Entry-Level Engineering Work. Who Trains the Young Engineers?

When I started my career on a support desk, I dealt with failed logins, missing permissions and recurring system errors every day. The work was repetitive, but repetition taught me to spot patterns, test assumptions and look beyond the symptom a user reported. Over time, I was resolving around 90% of incidents at first contact. That experience still shapes how I approach infrastructure work. As AI takes on more of these routine tasks, I worry about how the next generation of engineers will develop the same judgement. That kind of repetition is becoming harder for early-career engineers to get. AI tools are already embedded in software development, with 84% of developers surveyed by Stack Overflow saying they use or plan to use them, while 51% of professional developers use them daily. SignalFire’s 2026 State of Tech Talent report found that entry-level hiring has fallen roughly 65% at major technology companies and 76% at early-stage startups compared with 2019. AI is increasing pro...