DevOps & SRE

Most platforms do not fail. They stay empty. The landing zone is running, the pipeline template exists, and teams still deploy the way they did three years ago, because the old way is the one they know.

We work inside your teams until the platform is the path of least resistance, and we keep what runs on it fast and boring once it is there.

TRUSTED BY TEAMS AT:

  • Deutsche Telekom logo: white stylised letter T on a magenta background
  • Uniper logo in blue, showing the word "uniper" split across two lines.
  • GOLDBECK logo in bold black uppercase letters on a white background
  • PwC logo featuring the lowercase letters "pwc" in black with two orange diagonal shapes above
  • Vattenfall logo with the name in dark grey bold letters and a circle split into yellow upper half and blue lower half on the right
  • Schwarz-produktion logo on a white background reading “SCHWARZ PRODUKTION” in white text inside a dark blue square.
  • Cornelsen logo — white bold wordmark on a red background
  • Meridiam logo with tagline "for people and the planet" in dark green on a white background.

What exactly is DevOps & SRE?

Platform engineering builds the foundation. DevOps and SRE decide whether anyone benefits from it. Two disciplines, and we offer both, because one without the other leaves you with either a platform nobody uses or teams optimising on ground that cannot carry them.

This page is about everything that happens on top of the foundation. If the foundation is what you are missing, we build it as Platform Engineering, or you take KumoOps and skip building one altogether. We work in your teams, on your services, with your people watching, because enablement that happens in a ticket queue is not enablement.

Not sure which one you need?

  1. Platform EngineeringNo platform yet and you want to build one.
  2. KumoOpsA platform without building one.
  3. DevOps & SREThe platform exists and your teams are not on it. You are here.

Golden paths, not documentation

A team adopts a platform when deploying on it is faster than the alternative. We build a template per framework you actually run, with the pipeline, the Helm chart, the base image, the policies and the dashboards already wired in. A new service goes from repository to production on day one instead of week six. Target coverage is the handful of service types that account for most of what you build.

Several people seated along a long wooden table in a wood-paneled room, attending a workshop with laptops and drinks on the table.

Pipelines and packaging, rebuilt for what you ship now

AI has made writing code faster. It has not made your build, test and release path faster, which is where the extra throughput now queues up. We modernise pipelines, containerise applications that were never designed for it, and fix the packaging underneath: base images with a maintenance owner, SBOM per build, signed artefacts, promotion between stages that does not rebuild the image. Faster delivery and a supply chain an auditor can follow are the same piece of work.

Two colleagues sitting at a wooden table, each focused on their own laptop.

Reliability you can measure, not assert

Most SLOs fail because they measure the load balancer instead of the user, or p50 instead of p99. We define objectives from real user journeys, instrument the services to match with OpenTelemetry, and cut alerting down to what someone should actually be woken up for. When something does break, the goal is the same every time: find the cause in minutes, not in a war room with fourteen people on a call.

Three people sit on green armchairs in a bright lounge, chatting beside a laptop and a small table.

Why this changes what your teams can ship

  • Enablement, with an end dateWe pair with your engineers on their services rather than handing over a wiki. The measure of success is that you stop calling us for the next one.
  • Golden paths from a tested library55 Terraform modules and 40 Helm charts, already running in production elsewhere. Your templates start from proven building blocks, not from a blank file.
  • We fix the numbers you are judged onDeployment frequency, lead time for changes, change failure rate and recovery time, measured before we start and again when we hand over. Plus SLOs per service and a toil inventory, so improvement is visible instead of claimed.
  • Regulated environments are the normal caseKRITIS, healthcare, carriers, financial services. Policies, evidence and segregation of duties stay intact while delivery gets faster, because they are part of the pipeline rather than a gate in front of it.
  • You choose who holds the pager afterwardsYour teams take on-call with the runbooks and dashboards we built with them. Or KumoOps does, our ready-made cloud platform with CloudOps included, at a fixed monthly price with incident response under 15 minutes

Four steps. From baseline to boring.

  1. Step 01

    Measure what you have

    The four delivery metrics, a toil inventory, and one value stream mapped end to end from commit to production. Usually the slowest part is not where anyone expects it.

  2. Step 02

    Build one golden path

    The service type you create most often, complete: pipeline, packaging, policies, observability, rollback. One real team, one real service, in production.

  3. Step 03

    Roll it out by pairing.

    The next teams migrate with us alongside them, not from a handbook. Every migration makes the template better and the next one shorter.

  4. Step 04

    Close the reliability loop

    SLOs, error budgets, on-call rotation and blameless postmortems that actually change something. Then your team runs it, or KumoOps does.

Areas we support in

Telecommunications

Carrier-grade delivery where a failed rollout is visible to millions within minutes.

Financial Services

Change control and separation of duties inside the pipeline, not as a manual approval in front of it.

Logistics & E-commerce

Peak-season reliability with SLOs tied to order flow rather than to server uptime.

Healthcare

Regulated release paths with full traceability from commit to deployed version.

Working with the team is genuinely enjoyable from a customer’s perspective. Requirements are thoughtfully challenged to ensure the highest quality of the software products. The teams are masters of their craft, ensuring the seamless operation of the containerized cloud application.
uniper logoLars DonatUniper Energy Sales · Uniper SE

What we work with.

Nothing here is exotic, and that is the point. Your teams have to run this after we leave, and hiring for it should be easy. We use the tools with the largest ecosystems and the best documentation, and we spend our effort on wiring them together well rather than on picking something clever.

DELIVERY
  • GitLab CI
  • Argo CD
  • Progressive Rollouts
  • Automated Rollback
  • Trunk based Delivery
PACKAGING & RUNTIME
  • Helm
  • Kustomize
  • Kubernetes
  • Container Images
  • Base Images
  • Signed artefacts
RELIABILITY
  • SLOs from user journeys
  • OpenTelemetry
  • Grafana
  • On call design
  • Error budgets
  • Blameless postmortems

Your platform is running. How many teams are actually on it?

Two weeks, one value stream, your four delivery metrics and the list of what is in the way. No obligation.

AVG. RESPONSE < 1 WORKING DAY