What the term is actually doing
The term AIOps has accumulated a certain amount of noise. In vendor usage it can mean anything from log clustering to a language model that drafts an incident summary. Stripped of marketing, the useful definition is narrower: applying statistical and machine-learning methods to operational data in order to detect, explain, prioritise or predict conditions that matter.
That definition is deliberately modest. It excludes the fantasy of fully autonomous operations and includes the unglamorous work of sorting through telemetry.
Observability as the substrate
Machine learning needs data, and operational data is uniquely awkward. It is high-volume, high-dimensional, time-ordered, and frequently correlated in ways that are not obvious from any single signal. Metrics, logs, traces and events each tell part of a story that only becomes legible when read together.
Three properties matter more than volume:
- Cardinality. Modern systems produce labels and dimensions that multiply quickly. A model that performs well on a hundred series can behave unpredictably on a hundred thousand.
- Non-stationarity. Operational baselines drift. Traffic patterns shift, deploys change behaviour, dependencies evolve. A model trained on last quarter's normal is frequently wrong about this week's.
- Label scarcity. There is rarely a clean answer key. Most incidents are not labelled, and the ones that are labelled are often labelled inconsistently across teams.
Any serious approach to AI operations has to treat these as first-order constraints rather than data-cleaning details.
Anomaly detection and its limits
Anomaly detection is the most common entry point, and the most commonly oversold. Statistical methods can flag deviations from a learned baseline, but a flag is not a diagnosis. The failure mode is familiar: alert volume increases, signal-to-noise declines, and the system becomes something operators learn to ignore.
The more useful framing is that detection should be judged by what happens after it fires. If a detection reduces time-to-understanding, it is doing work. If it merely moves the moment of confusion earlier, it is not.
Prioritisation over detection
In practice, the harder problem is not finding things — it is deciding which of them to look at first. A production environment may surface hundreds of concurrent signals, most of which are benign. Ranking them requires context that goes beyond the signal itself: recent changes, dependency topology, business criticality, historical incident patterns, and the current state of on-call.
This is where machine learning has more room to be genuinely useful, and where the design question becomes organisational rather than purely technical. A prioritisation model encodes a theory of what matters. That theory should be legible to the people who have to live with its consequences.
Human in the loop
The phrase "human in the loop" is often used as a fig leaf — a way of saying that a system is technically supervised without saying anything about how the supervision actually works. A more honest version asks specific questions:
- What information does the human see, and in what form?
- What is the cost of overriding the system's suggestion?
- How does the human's response feed back into the model?
- Who is accountable when the system is wrong?
Answering these questions well is not a constraint on AI operations. It is most of the design work.
What to watch
The credible near-term direction is not autonomy. It is compression — reducing the distance between a signal appearing and a competent human understanding what it means, and reducing the number of decisions that require a human at all. That is a smaller claim than the category usually makes, and a more useful one.
This page describes a field of exploration. ConvexOps does not currently operate a product, service or platform in this area.