Anyone who knows me knows that I am big on self-hosting.
Over the years, my homelab has grown from a few containers into an environment running services such as my password manager, media server, recipe manager, DNS, analytics, monitoring tools and many other applications.
As the environment grew, one maintenance task became increasingly repetitive: keeping up with container image updates.
For a long time, Watchtower handled that job. For some containers, I was comfortable allowing automatic updates. For others, Watchtower would simply notify me that a newer image was available and I would investigate it myself.
That worked, but it still left an important question unanswered:
A new image being available does not necessarily mean that it is safe to install.
Before upgrading, I often wanted to know:
What changed in this release?
Are people reporting problems after upgrading?
Are there breaking changes?
Does the application require a database migration?
Is the new version compatible with the database version I am currently running?
Are there configuration changes I need to make first?
Is there anything specific about my environment that makes this upgrade risky?
Watchtower could tell me that something had changed.
What I really wanted was something that could understand what changed and what I should do about it.
That became one of the main reasons I introduced Hermes Agent into my homelab.
The idea was fairly simple.
Instead of receiving an update notification and doing all the research myself, I wanted an AI operator that could investigate the update for me.
When a new image became available, Hermes could identify the application and version, investigate release notes and changelogs, look for breaking changes, check migration requirements, research known issues, consider compatibility with my existing environment, assess the risk and recommend whether I should proceed.
That sounded useful.
But it immediately created a much more important problem.
I did not want an AI agent with unrestricted access to my infrastructure.
Giving an LLM the Docker socket, root privileges, unrestricted SSH access or broad sudo permissions would certainly make automation easier.
It would also create exactly the kind of security model I did not want.
So the project became less about giving AI access to my homelab and more about answering this question:
How can I give an AI agent enough authority to be useful without allowing the AI itself to become the authority?
The answer eventually became a four-part architecture.
The architecture separates governed execution from durable event context.
I separated the system into four responsibilities:
Observe → Reason → Govern → Act
Each layer has a different job, and deliberately separating those responsibilities became one of the most important design decisions in the project.
Watchtower detects container-image changes.
Originally, that resulted primarily in notifications. That works well when the consumer is a person who sees the notification when it arrives, but notifications are a weak source of operational history for an agent.
Logs rotate. Containers restart. Notifications are transient. Hermes may not be processing an event at the moment it occurs.
So I added a Watchtower Event Service.
The Event Service captures Watchtower notifications and preserves them as structured, append-only events in persistent history.
The path became:
Watchtower → Event Service → Persistent Event History → homelabctl → Hermes
Hermes can now ask for events from a bounded time window and reason about something that happened earlier.
The agent does not have to be present when an event occurs in order to know that it occurred.
That turns a transient notification into durable operational context.
Hermes provides the intelligence layer.
When an update is detected, Hermes can investigate the application and release by examining things such as:
release notes,
changelogs,
GitHub issues,
upgrade documentation,
breaking changes,
database requirements,
migration notes,
and community reports.
It can then interpret that information in the context of my environment and recommend a course of action.
This is where an AI agent adds considerably more value than conventional automation.
Traditional automation works extremely well when the rule is deterministic:
If X happens, do Y.
Software upgrades are often less predictable.
The better question is:
X happened. Given everything we know about this application, this release and this environment, what should we do?
That is a reasoning problem.
But reasoning is not authorization.
That distinction became one of the central principles of the entire project.
I built a deterministic control layer called homelabctl to sit between Hermes and the infrastructure.
Hermes may determine what it thinks should happen.
homelabctl determines whether it is actually allowed to happen.
The policy layer defines things such as:
which services Hermes is allowed to operate,
which operations are available,
whether a particular operation is permitted,
and whether that operation requires my approval.
For example, a service might permit Hermes to request a restart while requiring my explicit approval for an update.
Conceptually:
beszel:
permissions:
restart: true
update:
mode: approval-required
This configuration grants two separate capabilities. Hermes may request a restart of Beszel, while an update requires explicit approval. Restart permission does not automatically grant update permission, and the update policy cannot be bypassed by the agent.
A broader section of the control file might look like this:
mealie:
permissions:
restart: true
update:
mode: disabled
beszel:
permissions:
restart: true
update:
mode: approval-required
it-tools:
permissions:
restart: true
update:
mode: automatic
Another service could have updates disabled entirely, while another could eventually be eligible for automatic updates.
Importantly, making a service eligible for automatic updates does not itself enable unattended execution. A separate global automation gate controls whether automatic execution is enabled at all.
The important point is not the YAML syntax.
It is where the authority lives.
Adding more intelligence to Hermes does not expand its permissions.
Changing the AI model does not expand its permissions.
Improving its prompts does not expand its permissions.
Authority changes only when I deliberately change the control configuration.
If Hermes believes something should be restarted but policy says it cannot restart that service, the answer is no.
If an operation requires approval, Hermes cannot reason its way around that requirement.
Once an operation has passed the policy checks and received approval where required, the infrastructure can execute it.
But even execution is not enough.
A successful Docker command only tells me that Docker accepted the request. It does not tell me that the application recovered successfully.
After an authorized operation, homelabctl waits for the container to return and evaluates its resulting state. Depending on the service, that can include Docker health status and verification of an external HTTP endpoint.
The operation is not considered successful merely because the command ran.
The expected outcome must also be verified.
The control layer also records structured audit information about the requested operation, target, authorization decision, execution and result.
If something goes wrong, I do not want the explanation to be:
The AI probably did something.
I want to know what was requested, whether it was permitted, what executed and whether the service actually recovered.
The four-layer model also dictated how Hermes connects to the homelab.
The connection is intentionally constrained.
Hermes uses a dedicated operating-system account called agent-operator. That account has its own SSH key, but the key does not open a normal interactive shell.
The SSH configuration applies a ForcedCommand and disables facilities the agent does not need.
Every connection is directed through an SSH-facing gateway and then a second execution wrapper before homelabctl is invoked.
The actual path is:
Hermes
↓
SSH as agent-operator
↓
sshd ForcedCommand
↓
agent-homelabctl-ssh
↓
agent-homelabctl-exec
↓
homelabctl
↓
service and operation policy
↓
Docker
It is deliberately not:
Hermes
↓
SSH shell
↓
Whatever command the model generated
Those wrappers are not merely deployment helpers. They are part of the security boundary.
Hermes does not receive the Docker socket, root access, unrestricted sudo or a general-purpose command environment.
I did not just restrict which commands the AI could run.
I removed the shell from the equation entirely.
The same principle applies to the commands exposed to Hermes.
I do not want the agent constructing arbitrary infrastructure operations.
Conceptually, I want this:
homelabctl restart beszel
not this:
ssh server "docker compose down && docker pull ... && docker compose up ..."
The first expresses intent through a controlled interface.
The second gives the agent control over implementation.
With the controlled interface, homelabctl can determine whether the service exists in the allowlist, whether the requested operation is permitted, whether approval is required and how the operation should actually be executed.
The AI does not need to know or control those implementation details.
There was another problem I had to solve once Watchtower events became operational inputs.
A Watchtower event may identify an image, but the operational target I care about is normally a service or container.
Those are not always equivalent.
Multiple containers can use the same image. Names can differ between Watchtower, Docker Compose and the current inventory.
Automatically selecting the first apparent match would create exactly the kind of hidden uncertainty I wanted to avoid.
Instead, homelabctl correlates an event against the authoritative inventory and returns one of three outcomes:
Resolved — exactly one current service matches.
Ambiguous — more than one service could match.
Unresolved — the event cannot be mapped safely.
The deterministic layer does not manufacture certainty for the AI.
If it cannot establish the target reliably, it says so.
Hermes can investigate ambiguity and explain it.
It cannot silently convert ambiguity into authority.
With all of these pieces together, a typical upgrade workflow now looks closer to this:
New image detected
↓
Event Service captures and persists the event
↓
homelabctl correlates it with the inventory
↓
Hermes investigates the release
↓
Breaking-change and migration analysis
↓
Risk assessment
↓
Approval if required
↓
homelabctl validates policy again
↓
Controlled execution
↓
Health verification and audit record
↓
Result returned to Hermes
And at any point in that workflow, the correct outcome can be:
No action taken.
That is an important feature, not a failure.
One of the best validations of the architecture happened after the system moved beyond testing and started operating regularly.
Hermes now performs periodic image-update checks and can investigate available upgrades.
More importantly, it does not treat:
“New version available”
as synonymous with:
“Install it.”
Recently, I asked Hermes to investigate and upgrade a service in my environment.
Instead of simply carrying out my request, it researched the release and surfaced a concern that made proceeding questionable.
The workflow stopped.
That was exactly the behaviour I wanted when I started this project.
The success condition was never:
The AI successfully upgraded the container.
Sometimes the better outcome is:
I investigated the upgrade, found a compatibility or risk concern, and recommend that we do not proceed yet.
That is where the distinction between automation and an operator becomes meaningful.
Automation that always executes is easy.
An operator that knows when it should stop is considerably more useful.
Building and actually operating the system changed how I think about connecting AI agents to infrastructure.
The biggest lesson was the importance of separating intelligence from authority.
AI provides intelligence.
It can research, interpret, compare, investigate, reason and recommend.
Policy provides authority.
It determines what exists, what can be operated, which actions are permitted and when approval is mandatory.
Automation provides execution.
Once an operation has been validated and approved where necessary, the infrastructure performs it.
Those responsibilities do not need to belong to the same component.
In fact, I believe they are safer when they do not.
I have always been a strong believer in least privilege.
Connecting AI to infrastructure did not change that principle. If anything, it made it more important.
An agent should have access to the smallest interface necessary to accomplish its task.
Not:
Here is Docker. Be careful.
But:
Here are the specific operations you are allowed to request.
The architecture limits blast radius at multiple levels: no direct Docker access, no unrestricted shell, constrained SSH, explicitly allowlisted services, explicitly permitted operations and human approval for higher-risk actions.
No individual control is the entire security model.
The combination creates the boundary.
Another lesson was that notifications and logs are not enough when an agent needs to reason about past events.
Persisting events separately from the component that generated them means operational history survives restarts, log rotation and periods when Hermes is not actively processing an event.
That has applications well beyond container updates.
This may be the biggest thing I would do differently if I were designing a similar system again.
It is tempting to start with the AI.
Instead, I would first define:
What operations should exist?
Which services are in scope?
Which actions should never be available?
What requires approval?
What can safely happen automatically?
What is the narrowest interface the agent needs?
What happens if the AI makes the wrong decision?
Then I would put intelligence on top of that interface.
That produces a very different architecture from starting with a powerful AI agent and gradually trying to remove privileges afterward.
Least privilege is much easier to maintain when it is part of the architecture from the beginning.
I also plan to publish a sanitized open-source homelabctl repository.
The repository will include the deterministic control layer, the agent-homelabctl-ssh and agent-homelabctl-exec wrappers, example service policy, restricted SSH and sudo configuration, and the Watchtower Event Service.
The operational configuration from my own homelab will not be published.
The goal is to share the reusable boundary and safe examples without exposing environment-specific details or secrets.
That matters because the central claim of this architecture should be inspectable.
Readers should be able to see how unsupported commands are rejected, services and operations are allowlisted, approvals are enforced and the AI is prevented from becoming the authority.
Once the repository is ready, I will add the link here so readers can inspect the boundary for themselves.
I originally started this because I wanted something smarter than a Docker update notification.
Today, the system can monitor for changes, investigate what changed, evaluate potential risk, interact with me when necessary and operate parts of the homelab through a controlled interface.
But the part I am happiest with is not that it can perform an upgrade.
It is that it can choose not to.
And even when the AI recommends doing something, another layer still decides whether it is authorized to happen.
That is the behavior I wanted:
not an AI with the keys to the homelab, but an AI operator working inside defined boundaries.
After building and actually using this system, the principle I keep coming back to is simple:
AI can reason.
Policy decides.
I stay in control.
What started as a way to save time maintaining Docker containers has evolved into a small governed AI operations layer for my homelab.
And now that it is actually investigating changes, making recommendations, stopping questionable upgrades and operating successfully within the boundaries I created, I think it is finally doing the job I originally imagined.
Building an AI agent that needs access to real systems?
I help teams think through agent architecture, permissions, approval boundaries and controlled execution.