Find your biggest AI opportunities in under 30 minutes.Book a consultation
Back[ Coventa ]Case Study

[ AGENTIC AIOps ]LIVE

Resolve™

Run IT operations as outcomes, not headcount. A governed fleet of agents detects, triages, decides and acts across service desk, applications, infrastructure, network and security — while humans supervise the exceptions and the incumbent ITSM stack stays the system of record.

RESOLVE // CLOSED LOOP

// Problem

The Problem

Level-one and level-two operations consume the majority of an IT organisation's people budget while producing none of its differentiation. The work is repetitive, well-documented and still slow, because a runbook is only as fast as the human reading it — and out of hours, it isn't read at all. Automation projects address the top ten ticket types and stall, because the eleventh requires judgment about an unfamiliar system state.

  • Mean time to restore is bounded by human availability, not by how hard the problem is.
  • Runbook knowledge is written down but not executable, so it degrades as staff turn over.
  • Traditional automation covers the head of the distribution and leaves the long tail untouched.
  • Out-of-hours incidents wait for a person regardless of severity.

// Overview

Resolve operates as a closed loop over the existing operations estate. Agents ingest signals and tickets, triage against a retrieval layer grounded in the organisation's own runbooks and prior resolutions, decide on an action within a bounded-autonomy policy, execute against the relevant system, verify the restored state, and write the outcome back as new grounding for the next occurrence. Incumbent ITSM, monitoring and configuration tooling are retained — Resolve acts through them. Autonomy is scoped per action class, so read and diagnose are broadly permitted while change actions are gated according to blast radius.

// AI System

Why AI

The long tail is the entire argument for agents. Scripted automation must anticipate its trigger; the tail cannot be anticipated, which is why it stays manual. A retrieval-grounded agent can read an unfamiliar state, find the closest prior resolution in the organisation's own history, and reason to an action — and where it cannot reach confidence, escalate with a diagnosis already attached, so the human starts from a hypothesis rather than a blank ticket. The learning loop is what compounds: every resolution, human or agent, becomes grounding for the next one.

// Specs

Specifications

SCOPE
Service desk, applications, infrastructure, network, security
AUTONOMY
Bounded per action class; change actions gated by blast radius
GROUNDING
Retrieval over the organisation's own runbooks and resolution history
COVERAGE
24/7
SYSTEM OF RECORD
Incumbent ITSM retained; no rip-and-replace

// Features

Features

  1. 01Closed loop from detection through verification, not a triage assistant that hands off.
  2. 02Retrieval grounded in the organisation's own runbooks and prior resolutions, with citations.
  3. 03Autonomy scoped per action class, so diagnosis is free and change is gated.
  4. 04Escalations arrive with a diagnosis and evidence attached, not as a raw ticket.
  5. 05Every resolution written back as grounding, so coverage compounds with use.
  6. 06Operates through incumbent ITSM, monitoring and configuration tooling.

// Architecture

Architecture

RESOLUTION LOOP

Runtime · one item, left to right


  1. 01Signal / Ticket Intake
  2. 02Grounded Triage
  3. 03Decision + ConfidenceService DeskApplicationsInfrastructureNetworkSecurity
  4. 04Action Execution
  5. 05State Verification
  6. 06Write-Back to Grounding

dashed = the inference step, where the system exercises judgment

System stack

Data in · decisions out

01

Sources

Signals across the estate

Observabilitylogs, metrics, tracesITSMServiceNow / JiraCI/CD & change recordsInfra & network APIsSecurity tooling

02

Ingestion

Correlate before anyone reads

Alert correlationdedupe, group, timelineRunbook ingestionPDF, wiki, scriptsChange-event linkingTopology graph

03

Ontology

Incident with its causes

IncidentService · DependencyChange · DeployRunbook · StepAuthority

04AI

Intelligence

Diagnose, act, verify

Diagnosis agentLLM over logs, metrics, changesRunbook matcher & executorVerification loopSLO-checkedDomain agentsdesk · app · infra · network · securityConfidence gating

05Human

Human control

Authority is configured, not assumed

Pre-approved action classesEscalation with diagnosisChange approvalKill switch

06

Actions

Written back

Infra & app actionsITSM ticketsPost-incident notes

Observability

Every model call traced; evals run on real cases, not anecdotes.

Governance

Entitlements enforced at retrieval; rules versioned by the organisation.

Write-back

Systems of record are written only through the approval gate.

Below-confidence items escalate with diagnosis attached; the resolution returns to grounding either way.

// Impact

Impact

60–80%
Autonomous L1/L2 resolutionindicative target
2–3×
Faster mean time to restoreindicative target
24/7
Coverage without shift coverdesign intent

Interested in Resolve?

Let's scope a bounded pilot against your ticket distribution.

Get in touch
// End of case studyResolve™