How to interpret a service incident
| Phase | What it means | What users should do |
|---|
| Investigating | A pattern is confirmed but the cause or scope is not yet established | Avoid repeated changes and preserve timestamps and errors |
| Identified | The affected component and likely cause are known | Follow any stated workaround and avoid risky retries |
| Mitigating | A fix, rollback, reroute, or capacity action is being applied | Expect partial recovery and continue to record failures |
| Monitoring | Service has improved and stability is being observed | Test normal workflows once and report reproducible exceptions |
| Resolved | The incident is closed based on current evidence | Retry the original action and open support if the issue persists |
Platform incident or individual server problem?
A platform incident usually affects many users, a shared dependency, or a specific dashboard capability. An individual server problem can affect only one software build, world, plugin set, account, network, or configuration. Compare the public status information with your server console and a second independent test before assuming every failure has the same cause.
What to record during an interruption
- Exact time and timezone.
- Dashboard state and requested action.
- HTTP or browser error when relevant.
- Console output before and after the failure.
- Whether other servers, accounts, or networks are affected.
- Whether the action later completed without another change.
Post-incident validation
After a resolution, confirm account login, dashboard loading, server state, start and stop actions, console access, player connection, world saving, and any operation that failed during the incident. Do not perform unnecessary software upgrades at the same time; recovery validation should test the original workflow with as few additional variables as possible.
Transparency standard
Status communication should distinguish confirmed facts from investigation, identify the affected capability, use concrete timestamps, avoid unsupported promises, and close with a clear recovery state. Longer or higher-impact incidents may lead to documentation changes, monitoring improvements, or an internal post-incident review.