Territory 02

Cloud & Infrastructure

Distributed systems, observability and the operating discipline required to keep a modern estate reliable, efficient and explainable as it grows.

Cloud operations as a discipline

Cloud infrastructure was supposed to reduce operational burden. In some respects it did — provisioning that once took weeks now takes minutes, and capacity can be elastic rather than forecast. In other respects it relocated the burden rather than removing it. The work moved from managing physical machines to managing the abstractions layered above them, and the abstractions accumulated.

A modern estate is rarely one cloud or one architecture. It is a collection of services with different lifecycles, different owners and different failure modes, connected by networks whose behaviour is only partly under the operator's control. The operational challenge is not any single component. It is the seams.

The observability gap

Distributed systems fail in distributed ways. A latency increase in one service may originate in a dependency three hops away, in a shared resource, in a configuration drift, or in a change deployed by a different team an hour earlier.

Traditional monitoring was built around the assumption that a failing component announces itself. That assumption no longer holds. What replaces it is observability: the ability to ask new questions of a system without having predicted them in advance. That requires high-cardinality data, consistent instrumentation, and a willingness to treat telemetry as a product rather than an exhaust stream.

Efficiency as an operational concern

Cloud cost is frequently treated as a finance problem. It is more usefully treated as an operational signal. Where cost concentrates, complexity usually concentrates too. Idle resources, over-provisioned capacity, redundant pipelines and orphaned storage are often symptoms of processes that have drifted rather than decisions that were made.

Efficiency work is rarely glamorous. Its value is in the compounding: fewer moving parts, clearer ownership, less to reason about when something breaks.

Reliability has a cost

Reliability is not free, and it is not uniformly valuable. A system that supports real-time trading and a system that generates weekly reports do not warrant the same investment in redundancy. The operational discipline is to make that distinction explicit and to accept the trade-offs that follow, rather than defaulting to the highest available target.

The harder version of this question is what happens when reliability targets are set by expectation rather than analysis. Systems accumulate defences against failures that have not occurred, and the defences themselves become a source of complexity.

The distributed default

For most organisations, distribution is no longer a choice. Multi-region deployment, third-party dependencies and edge compute have become defaults rather than deliberate architectural decisions. Each addition is individually justified. Collectively they produce a system whose behaviour cannot be predicted from the behaviour of its parts.

That is the environment cloud operations has to work in. The tools have adapted — service meshes, distributed tracing, infrastructure as code — but the underlying tension has not. More distribution means more capability and more ways to be surprised.

What good looks like

Good infrastructure operations tend to share a set of unremarkable properties. Changes are reversible. State is declared rather than discovered. Alerts correspond to user-visible conditions. Ownership is unambiguous. The system can be explained by someone who did not build it.

None of these are novel. None are achieved by tooling alone. The consistent pattern is that operational quality is a consequence of decisions made early and maintained deliberately — and that it degrades quietly when it is not.

This page describes a field of exploration. ConvexOps does not currently provide cloud or infrastructure products or services.