AI Doesn't Need Autonomy—It Needs Discipline

Applying SRE principles to AI agents through a controlled audit-plan-apply workflow with policy and human permission.

Audit-plan-apply interface with a policy and human-approval gate before execution

Technical illustration

Every AI agent demo looks impressive.

The assistant installs packages, fixes issues, restarts services, cleans disks—all without asking. It feels powerful. Magical, even.

And then you imagine it running sudo on your own machine. That’s where the excitement quietly turns into discomfort.

I like AI. I use it daily. I build with it. But when I tried applying autonomous AI agents to something very real—system maintenance on my Linux machine—I ran into a hard truth:

AI doesn’t need more autonomy. It needs discipline.

This article is about why I built a terminal-based assistant called Housekeeper, and what it taught me about designing AI that people can actually trust.

Code and details at https://github.com/pain459/housekeeper. Feel free to contribute ❤️

Terminal session showing Housekeeper audit, plan, and apply commands with approval prompts for system actions

End-to-end execution of Housekeeper

Why Autonomous AI Breaks Down in Real Systems

Autonomy sounds attractive because it removes friction.

But real systems are full of properties that autonomous agents struggle with:

  • Side effects
  • Partial failures
  • Irreversible actions
  • Context that’s hard to infer
  • Commands that are “safe” only in the right situation

Restarting a service might fix one issue and trigger another. Cleaning logs might delete information you need tomorrow. Upgrading packages might silently break drivers.

In production environments, we don’t tolerate this kind of blind action. We require audits, plans, approvals, and rollbacks.

So why would we lower the bar just because the operator is an AI?

Thinking Like an SRE, Not an Agent Designer

I come from an SRE mindset, and one habit is deeply ingrained:

Observe first. Then plan. Then act.

No SRE logs into a system and starts “fixing things” immediately. They collect signals, understand the state, propose actions, and get approval.

So instead of building an autonomous agent, I built Housekeeper—a terminal-based assistant designed to behave like a careful SRE.

Not an executor. A reviewer.

Housekeeper’s Core Model: Audit → Plan → Apply

Housekeeper is intentionally boring in its design:

audit → plan → apply

And that’s exactly why it works.

Audit: Read-Only, Always

The assistant starts by collecting system signals.

(.venv) ravik@STING:~/src_git/housekeeper$ housekeeper audit
╭──────────────────────────────────── housekeeper ────────────────────────────────────╮
│ Audit saved: /home/ravik/.local/share/housekeeper/audits/audit_20251227_200202.json │
╰─────────────────────────────────────────────────────────────────────────────────────╯

Behind the scenes, this gathers things like:

  • package upgrade health
  • disk usage and journal growth
  • failed systemd units
  • firewall status
  • reboot-required signals
  • SSD trim state

Nothing mutates the system. No sudo. No side effects.

For most personal machines, this alone surfaces issues people didn’t know existed.

Plan: AI That Explains Itself

Next, the AI analyzes the audit and produces a plan.

(.venv) ravik@STING:~/src_git/housekeeper$ housekeeper plan
╭────────────────────────────────── housekeeper ───────────────────────────────────╮
│ Plan saved: /home/ravik/.local/share/housekeeper/plans/plan_20251227_200226.json │
╰──────────────────────────────────────────────────────────────────────────────────╯

A plan isn’t just commands—it’s structured reasoning.

Actions 3:
• Simulate package update
 Rationale: Ensure system is ready for upgrades
 Risk: safe
• Clean up journal logs
 Rationale: Reduce disk usage from system logs
 Risk: safe
• Run fstrim to optimize disk usage
 Rationale: Improve SSD performance
 Risk: safe

Crucially:

The AI cannot execute anything.

It makes a case. It does not take action.

Apply: Policy + Permission

This is where discipline really shows up.

(.venv) ravik@STING:~/src_git/housekeeper$ housekeeper apply --plan /home/ravik/.local/share/housekeeper/plans/plan_20251227_200226.json

Each command is evaluated against a strict policy and shown for approval:

╭─────────────────────────────────── housekeeper ───────────────────────────────────╮
│ Loaded plan: /home/ravik/.local/share/housekeeper/plans/plan_20251227_200226.json │
│ Actions: 3 │
╰───────────────────────────────────────────────────────────────────────────────────╯

--------------------------------------------------------------------------------
Simulate APT Update (risk: safe)
To ensure the system is up to date and check for any potential issues with package management.
 $ sudo apt update -y
 -> policy: safe (sudo apt operation (approval required))
Run this command now? [y/N]: y
 $ sudo apt upgrade -s
 -> policy: disruptive (Package changes require approval)
Run this command now? [y/N]: y

Even safe commands require confirmation.

Disruptive commands are clearly flagged:

Review UFW rules (risk: safe)
To ensure the firewall is configured correctly and no unnecessary ports are open.
 $ sudo ufw status
 -> policy: disruptive (Firewall changes require approval)
Run this command now? [y/N]: y

Nothing runs “just because the AI said so”.

Why Permission Is Not Friction

A common argument against human-in-the-loop systems is that they’re slow. That assumes speed is the primary goal. In system operations, the real goal is trust.

When an assistant explains:

  • why an action is suggested
  • what exactly will happen
  • how risky it is

…you start trusting it.

Ironically, Housekeeper feels more capable than autonomous agents because it’s predictable. I always know what it’s about to do.

A Practical Critique of Autonomous AI Agents

This isn’t an anti-AI argument. It’s a rejection of blind autonomy in high-impact environments.

Autonomous agents work well when:

  • actions are reversible
  • failure is cheap
  • the blast radius is small

System maintenance is the opposite. Here, intelligence without accountability is a liability.

AI shines when it:

  • analyzes complex state
  • surfaces hidden risks
  • proposes well-reasoned actions
  • explains trade-offs clearly

It fails when it:

  • assumes context
  • executes without consent
  • optimizes for completion over correctness

Housekeeper intentionally avoids that trap.

Policy Beats Prompting

One of the most important lessons from building this system:

Safety must live in code, not in prompts.

No prompt can replace:

  • command allow-lists
  • hard blocks for destructive operations
  • explicit sudo rules
  • risk classification enforced by logic

The AI can suggest anything. The policy decides what’s allowed.

This separation is what makes the system reliable.

What Building Housekeeper Changed for Me

Working on this project shifted how I think about AI entirely.

AI is not a replacement for operators. It’s a force multiplier for judgment.

The best role for AI in system operations isn’t:

  • autonomous executor
  • fixer
  • self-directed agent

It’s:

  • planner
  • reviewer
  • advisor
  • second brain

Exactly like a Senior SRE sitting next to you, saying:

“Here’s what I see. Here’s what I’d recommend. You decide.”

The Bigger Lesson

We don’t need smarter AI agents. We need better-designed ones.

AI that:

  • knows its boundaries
  • respects permission
  • logs everything
  • accepts oversight
  • optimizes for correctness, not autonomy

Discipline isn’t a limitation. It’s what makes intelligence usable.

Final Thought

Housekeeper is not flashy.

It doesn’t surprise me. It doesn’t act on its own. It doesn’t pretend to be in charge.

And that’s exactly why I trust it.

AI doesn’t need autonomy—it needs discipline.


Originally published on Medium on December 27, 2025.