A night watch for
31 production services
When a site goes down, the Ops Agent investigates the live estate, writes a frozen list of commands, and waits. Approve means "run exactly this". The model never gets a second turn.
Platforms we run and connect for colleges


Investigate freely.
Change only what a person has read.

Approve never re-enters the model
This is the design rule everything else follows. The Approve button is bound to the incident, the proposal round and a hash of the plan. Execution is a plain bash runner over the stored list. There is no step where the AI reinterprets "yes".
- Validated deterministically, then preflighted against the live system, then independently reviewed, before anyone sees a button
- Only configured Slack accounts can approve. Everyone else is ignored for anything that changes production
- A deny list runs again at execution time: no forced deployments, no destructive patterns
- "Suggest another action" forces a new proposal, capped so it cannot loop
What it does
Pingdom says a site is down. Within seconds the agent has acknowledged it in Slack, waited 75 seconds to see whether it was a blip, checked whether a deployment is already rolling, and only then started an investigation against live AWS evidence.
- Reads the containers, load balancers, hosts and logs for the named site, from a map of 31 services
- Explains what it found in a Slack thread, in plain language
- If a fix is justified, writes a frozen plan: up to 16 exact shell commands and up to 8 success checks
- A second model pass, with no tools, must review that plan and return PASS before an Approve button appears
- A person presses Approve. The executor runs the stored commands, stops at the first failure, and checks the result
- Every morning a read-only sweep of the whole fleet posts one exception report. It never repairs
#ops-agent · incident thread · what an engineer sees
moodle.inneall.net
What we saw: load balancer target unhealthy for 4 minutes. The web container on one host is out of memory and has been killed twice.
Suggested action: restart that one container and re-check the health endpoint.
Approve Suggest another action DeclineThe full command list stays on the incident record, not in the channel. If approved, one line reports the outcome.


The ten-minute veto
One narrow class of repair can run without a click: the site is still down, the plan is low risk, it uses only Systems Manager commands, and both the review and the live preflight passed. Then the agent waits ten minutes from first detection, reminds at five, and acts if nobody has objected.
- Nothing that touches git, pipelines, databases, DNS, IAM, secrets or server lifecycle qualifies
- The clock starts at detection, not when the card was posted
- Most incidents never get this far. They end as false alarms or deploy blips with no production change
- The IAM role denies the dangerous actions outright, so the policy is a second line of defence, not the only one
What it is, and what it is not
It is deliberately not a chat model with the AWS keys. It is an investigation and proposal tool with hard execution gates, bound to a map of known sites and fenced in by code and policy.
- It is a first responder with hard gates. It is not a replacement for on-call; people still own the judgement calls
- Approve means "run this exact script". It does not mean "use your judgement"
- Most "downs" end as false alarms or deploy blips with no production change at all
- It can run frozen, validated commands on a live container. It cannot rewrite access rights, databases, DNS or secrets, however it is asked
- It is bound to a service map and our own operational skills. It is not a general assistant over the account
- Backed by 179 automated tests across approval policy, the executor, queue routing, plan review, deploy watch and recovery
What we learned
The interesting engineering was not the model. It was everything around it.
- Confirm before you think: a 75-second re-check kills most alerts before any model is called
- Watch deploys: never repair a service while a pipeline is rolling
- Separate the queues: operators should get "status" back in a second
- Freeze the plan: the thing a person approves must be the thing that runs
- Keep evidence off the channel: one outcome line in Slack, the full command record on the incident
- The worker costs tens of dollars a month. Model spend is capped by the plan and fails closed, never silently
Why we built it
We run dozens of client sites on a shared container platform. Outages look identical from the outside and need a different safe fix every time: the wrong container on a shared host, a process out of memory, a deploy still rolling, or nothing at all. The first twenty minutes of every incident was the same work: is it really down, which container, which command, is a pipeline running. We built something to do that work in seconds, so an engineer wakes up to a diagnosis and a plan, not a blank page.


DISCOVER
Scope of work
Integration mapping
Risk and dependency analysis
Finalised project plan
DESIGN
Concept proposals
Stakeholder reviews Design development
Final specifications
Approval to progress
BUILD
System integration
Customisations
Content and data migration
Front end interface
TEST
Browser/app performance testing
Accessibility
System and security
DEPLOY
Final configurations and domain go-live
Support and monitoring
Bug-fix support
Performance tuning
SUPPORT
Responsive support
Planned maintenance
Quality control
Performance tuning
We love API environments
and the challenge of integration
We’ve integrated systems others refused to touch!
“Inneall’s impact on our ability to deliver reliable, scalable technology has been absolutely enormous. What sets them apart is their deep knowledge of our business, they understand our systems, our teams, and our priorities. That means we’re not dealing with a technology supplier; we’re working with a partner who genuinely cares about getting it right.”
Edward Ormonde
Head of IT, Dublin Business School
