An agent that does the wrong thing leaves evidence — a bad diff, a malformed record, a customer complaint. An agent that does nothing and reports success leaves a green checkmark, and green is the one signal your incident process is built to trust. Worry more about the second failure, because everything a DevOps team owns for catching failure — exit codes, run status, alert rules, the dashboard itself — is instrumentation built for software that fails loudly. None of it fires when the failure is an absence. Andrew Filev made the strongest version of the reliability argument in these pages in July ( https://devops.com/reliability-comes-from-the-system-not-the-agent/ ): reliability is a property of the system, not the agent. “Reliability has rarely come from any single component in isolation,” he wrote. “It comes from how systems handle failure.” Aviation does not assume perfect pilots and hospitals do not assume perfect surgeons; both wrap imperfect actors in app...
OllyGarden this week revealed it has extended the capabilities of its artificial intelligence (AI) for optimizing the collection of telemetry data to now also discover where no existing instrumentation exists and what instrumentation should be applied to address that gap . Fresh off raising an additional $4 million in funding, OllyGarden founder Juraci Paixão Kröhling said a Minimum Viable Instrumentation (MVI) capability that has been added to the Rose AI agent the company previously developed makes it possible to identify the most relevant sources of telemetry data that should be instrumented using OpenTelemetry, an open-source framework for collecting that data that is being advanced under the auspices of the Cloud Native Computing Foundation (CNCF). That capability extends the scope of an AI agent that was created to help DevOps teams reduce the amount of telemetry data they need to collect and store by identifying which logs, traces and metrics are the most relevant. The overall...