KCD Porto x DevOpsDays Portugal 2026 · 30-minute talk
Autonomous DevOps: Engineering Systems That Don't Need Our Attention
Automation took execution off our hands. It didn't take attention off them. This talk is about which operational work can safely run without a human looking at it, and how to decide that with a score instead of a feeling.
- When
- (exact slot to be announced)
- Where
- Super Bock Arena, Porto
- Track
- AI + ML Platform Engineering · 30 min
- Speaker
- Ilia Dubovskii, Beamreach
Speaker page on the KCD Porto schedule → DevOpsDays Portugal 2026 →
The talk in one paragraph
Autonomy isn't a property of the agent. It's a property of the operation. An operation can run without human attention when it's verifiable (the system can check it did the right thing), reversible (there's a cheap, known way back) and contained (the worst case is known in advance and small). So the engineering work isn't building a smarter agent; it's building the machinery that decides which operations qualify, and making that decision auditable. I call that machinery Trust Engineering. In my own system it isn't a slide: it's a scoring function and a threshold, and the talk shows the real one.
Abstract
For decades, DevOps has focused on automation. We built CI/CD pipelines, Infrastructure as Code, observability platforms and deployment systems that execute work reliably and consistently. But automation still depends on human attention. Someone reviews the dashboards. Someone investigates the alerts. Someone decides whether an infrastructure change is safe. Someone maintains the honeypots. Someone keeps Infrastructure as Code aligned with reality.
This talk introduces Autonomous DevOps: an approach to engineering operational systems that no longer require our limited focus for routine decisions while remaining observable, auditable and trustworthy. Autonomy does not mean giving AI unrestricted control. It means identifying operational tasks that are verifiable, reversible and contained, and letting them run within clearly defined boundaries while humans stay responsible for strategy, architecture and exceptions.
Using practical examples, including autonomous honeypot management, Infrastructure as Code coverage, infrastructure rightsizing and observability noise suppression, we'll look at how DevOps teams can move from automation to autonomy gradually, without compromising reliability or security, and why Trust Engineering is the foundation that makes it possible.
Key takeaways
- What Autonomous DevOps is and why it extends traditional automation.
- Why human attention has become the limiting factor in operating cloud-native systems.
- A practical framework based on verifiable, reversible and contained operations.
- Real-world examples of autonomous operational workflows.
- How Trust Engineering provides the guardrails that make autonomy practical.
What the talk covers
The framework is code
In Compass, the system I build, every proposed operation gets a risk score made of five parts: blast radius, data loss, reversibility, operational load and compliance. Each part is computed from facts about the infrastructure (which resources, which types, which environment, what live metrics say), not from the model's opinion. A single autonomy dial sets how much risk may run unattended, and the default lets only the lowest-risk operations through. Raising it is a deliberate configuration change. The talk puts the actual scoring function on screen.
Containing the model separately
The score decides whether an operation may proceed. A second, independent layer decides what the model can physically do: it works inside a repository checkout with an allow-listed set of tools, its loops stop after a fixed number of turns, and its only output is a pull request. It has no write path to live infrastructure. A prompt is a preference; a missing permission is a boundary.
Four examples, one shape
Honeypot management, IaC coverage, rightsizing and noise suppression, each scored the same way. One of them fails the test, and saying so is the point: rightsizing touches live, load-bearing resources and shouldn't run unattended.
The honest gap
Of the three properties, verifiability is the hardest to check mechanically, and that includes in my own system. The talk is open about where that leg stands and what closing the gap takes.
What's opinion
The talk marks which claims are running code and which are my opinion. Two opinions I'd most like challenged: that attention, not compute or headcount, is now the binding constraint on running cloud systems; and that verifiability will turn out to be harder than reversibility or containment.
Who it's for
Platform and DevOps engineers, SREs and engineering leads who are being asked to "add AI" to operations and want a way to decide where it's safe. It's a grounded point of view, not a product launch: no pricing, no demo pitch. If you run this kind of system and want to try the ideas in a real environment and break them, I'd like to hear from you.
Materials
Slides and the recording will be linked here after the talk.
Slides
Published after the talk on 19–20 November 2026.
Recording
Published after the talk on 19–20 November 2026.
Read more
- Autonomous DevOps guide — the framework from this talk, written up.
- Autonomous CloudOps
- Obstacles for autonomous software
- IaC coverage and Terraform drift — one of the four examples.
- Honeypot security and canary assets — another of the four.
- Compass — the system the examples come from.
- One Fake Password Is Worth More Than 10,000 Alerts — my Ignite at DevOpsDays Barcelona, 14 Nov 2026.
- All talks
Beamreach