Partial Outage across multiple systems

Major incident Web Application EU data center
2026-07-15 08:45 CEST · 1 hour, 20 minutes

Updates

Post-mortem

Incident Summary

On 15 July 2026, between 08:44 and 09:45 CEST, our ERP experienced a partial outage lasting approximately one hour,
affecting the availability of the web application for customers in the EU. The root cause was a surge of contention on a
small number of database records, which caused a growing backlog of stalled requests to exhaust the available
request-handling capacity of a backend service. As that capacity became saturated, errors began appearing across many
unrelated parts of the application, rather than being limited to the specific workflows that triggered the original
contention.

Impact on Customers

During the incident, customers may have experienced:

  • Slow or failed page loads across many parts of the Magicline web application
  • Errors on actions unrelated to the underlying trigger, such as checking in members, logging in, or loading
    translations and locale data
  • Intermittent failures across a broad range of requests, rather than a single isolated feature

Our Response

Our response followed the status page timeline:

  • 08:28 CEST — Underlying condition begins

    A small number of database records began experiencing heavy, concurrent load, causing operations touching those
    records to take longer than expected to complete.

  • 08:32–08:44 CEST — Detected

    Automated monitoring detected a rising error rate and triggered internal alerts.

  • 08:44–09:00 CEST — Escalates

    As stalled operations accumulated, the affected service’s capacity to handle new requests became fully saturated,
    producing a broad wave of errors across many unrelated parts of the application.

  • ~09:15 CEST — Recovery begins

    The underlying contention started to subside naturally, and error volumes began declining.

  • 09:35 CEST — Active remediation

    The affected service instances were restarted to relieve pressure on request-handling capacity.

  • 09:40–09:45 CEST — Resolved

    The remaining stuck database operations were manually cleared, and error rates and response times returned to normal.

Resolution

The incident was resolved through a staged approach:

  • Containment — Restarted the affected service instances to relieve pressure on shared request-handling capacity.

  • Remediation — Manually cleared the database operations holding the contended records, allowing normal processing
    to resume.

  • Verification — Confirmed error rates, response times, and backend capacity utilization returned to baseline.

  • Monitoring — Reviewing and strengthening early-detection monitoring for this class of database contention going
    forward.

Lessons Learned

A small number of highly contended database records can, under concurrent load, exhaust shared backend capacity and
cause errors well beyond the workflows directly involved. We are evaluating measures to contain this “blast radius,”
such as isolating capacity per workflow and circuit breaking.

Retry behavior on contended operations should use backoff rather than immediate retries, to avoid compounding contention
during a spike.

Earlier, more targeted detection for this type of database contention will help us identify and intervene in similar
situations faster in the future.

July 21, 2026 · 09:20 CEST
Resolved

The incident ‘Partial Outage across multiple systems’, which occurred between 2026-07-15 08:44 CEST and 2026-07-15 10:05 CEST, has been resolved.

Our engineering team will review the issue and implement additional measures to prevent similar incidents in the future.

If you continue to experience any problems, please open a ticket with our support team.

We apologize for any inconvenience caused.

July 15, 2026 · 10:06 CEST
Update

We found one cause and are fixing this issue currently. We will then continue investigation for potential additional causes.

July 15, 2026 · 09:59 CEST
Investigating

We are currently experiencing performance degradation and intermittent availability across several systems. Some features may be slow or unavailable. Our team is investigating the root cause.

July 15, 2026 · 09:39 CEST
Investigating

We are currently seeing high error rates and failing request across multiple systems and are analysing the error with highest priority currently

July 15, 2026 · 09:38 CEST
Investigating

We are currently experiencing performance degradation and intermittent availability across several systems. Some features may be slow or unavailable. Our team is investigating the root cause.

July 15, 2026 · 09:12 CEST

← Back