Offer · Operate Early access

Your tools already know what broke.
The hour after that is the problem.

Incident Neutralisation takes a correlated problem from whatever you already run, gives it one ticket per root cause, remediates it only from a runbook your team wrote and your change process authorised, and then reads your monitoring for a full window to decide whether the service is actually back.

Six steps, every time

We do not correlate your events — your monitoring or your correlation tool already does that well. We start where that stops and carry the problem to a proven recovery.

1

Own it

One ticket per root cause. Symptoms attach to it; a cause that already has a ticket never gets a second one.

2

Diagnose

The specialist whose scope it is reads your monitoring, your logs and your configuration data, and shows the evidence.

3

Propose

A runbook from your catalogue, matching this cause on this class of asset — with its bounds and its backout.

4

Execute

Through the automation you already own, under an authorised change, checked again in the second before launch.

5

Verify

Your monitoring, across the whole window, against how the service looked before the incident.

6

Record

Ticket, evidence, timings, refusals and the graph markers — the trail your service review and your auditors read.

Back means your monitoring says back

A job that ends green says the job ran. A quiet console can mean the tool that would have told you is down. And the agent that performed the action is the last one who should vouch for it — so the verdict is computed from the readings, the same way every time, and it is kept.

Recovered

Every signal stayed inside its pre-incident band for the whole window, and nothing covering the service is still firing.

Not recovered

A reading left the band, or an alert is still firing. The ticket stays open and the action is backed out through its declared backout.

Unverified

A source could not be read, or the window has a hole in it. Said plainly, escalated — and never counted as an incident neutralised.

One agent per scope, because that is what makes autonomy safe

A generalist agent carries the blast radius of everything it could touch, so it can never be trusted with any of it. Each specialist has one scope, one set of runbooks and its own record.

Linux

Filesystems, memory, a service down or flapping, a log that filled a partition, a scheduled job that failed.

Windows

Services, system volumes, an update that left a device pending, a task that failed, a device that stopped reporting.

Virtualisation

A guest down, a guest running with nobody home, contention on a host, a datastore filling, a replication that stopped.

Storage

Volumes and aggregates filling, quotas reached, snapshot reserve gone, a replication broken or lagging.

Containers

Crash loops, pods nothing can schedule, evictions, a rollout that never became ready, a node under pressure.

Databases

Read and explained, never changed: blocking, replication lag, slow statements, vacuum and backup state — as an argued recommendation.

What the agents may not do when they start

Never, at entry

  • Act on a cause no runbook in your catalogue covers
  • Approve a change, or launch without one
  • Touch a service you called critical without a person approving that change
  • Change a database, at any hour
  • Repeat an action that already failed on the same cause
  • Widen a job's targets to finish it, or exceed the hourly action limit
  • Close an incident because a job succeeded, or because an alert went quiet

Always

  • Only one agent can launch anything, and it re-reads the record immediately before every launch
  • Waking a person is a correct outcome — out of hours, often the right one
  • A source that could not be read is reported, never treated as good news
  • A scope whose configuration data is stale runs in observation until it is fixed
  • Every action, refusal and verdict leaves a record

The incident you can still prevent

Most of what pages you at 03:00 was visible on Thursday afternoon: a volume climbing to its ceiling, free space falling to a floor, a queue drifting, a pool of connections shrinking. The trend is measured and projected, never guessed — and a weak trend is reported as weak.

What you get

  • The projected crossing time, with the readings behind it
  • How well the trend actually fits — a straight line through noise is a guess with a decimal point
  • The runbook that would prevent it, ready to raise as a change

What it is not

  • Not an authorisation: at entry, acting before impact is a recommendation a person decides
  • Not a promise about anything nobody measures
  • It earns more autonomy the way everything else does — on the record of projections that proved right

Autonomy earned per kind of action

A class is one runbook, on one kind of asset, through one tool. Each climbs on its own verified record, and drops back on its first failure.

Read & diagnose

Reading your monitoring, owning the ticket, diagnosing and recording runs on its own from day one. Nothing changes on your estate.

Catalogued runbooks

Start with a person confirming each execution, and run unattended once that class of action has a verified-recovery record behind it.

Critical services

One change, one approval, every time — until that class earns a pre-approved path of its own.

Before impact

Recommended first: the projection and the runbook go to a person, and the record of what proved right is what opens the next step.

Nothing is autonomous by default, and nothing is capped forever: the verified record decides, and your change board signs off each promotion. One failed verification, one out-of-scope attempt or one incident traced to an action puts a class back, at once.

What is live, what early access adds

Live on the platform today

  • Read connectors to your monitoring, on-call platform, service-management tool and configuration data
  • Execution through your own automation, with approval gates and a person's decision recorded
  • Recovery judged from monitoring against a pre-incident baseline, computed the same way every time
  • Trend projection with its fit quality, for the incidents you can still prevent
  • Tamper-evident audit trail, and installation on your premises

Built with early-access customers

  • Your runbook catalogue and its pre-approved classes as a first-class object, with promotion per class
  • Autonomy gated automatically on the freshness of your configuration data
  • Service levels for the agents themselves: auto-resolution, justified escalations, time to human pickup
  • Every action bound to a ticket by the platform, not only by the agents' own checks

Priced per governed estate band, plus per incident neutralised with a verified recovery — never per alert, and never for a recovery we could not prove. Your own operations team runs it, in or out of hours. Remediations take the same governed path as every other production change through Change Guard; for the storage, backup and virtual-desktop loops behind them see agentic infrastructure operations, and for proof that a business service can be restored at all, Recovery Proof.

See every offer · Fix Squad