Halyard Cloud status.halyard.dev
Uptime · 90d
99.982%
p95 latency
112ms
Incidents · 90d
4
Mean time to resolve
27min

Service uptime — last 90 days

hover a column for that day

Core platform

Developer surface

Regional edge

Operational Degraded performance Partial or full outage Scheduled maintenance No data

Incident history

all times UTC

Elevated build queue times in the CI cache layer

DegradedBuild PipelineMonitoring
  • Monitoring10:40 UTC

    Cache hit rates are back above 92%. We are leaving the incident open while the warm pool refills.

  • Identified09:38 UTC

    A bad eviction policy shipped with runner image 4.19 caused the shared layer cache to thrash. Rolling back to 4.18.

  • Investigating09:12 UTC

    Builds are queuing 3–6× longer than baseline in us-east-1. Investigating.

Scheduled reindex of the documentation search cluster

MaintenanceDocs & SearchCompleted
  • Completed03:15 UTC

    Reindex finished with no data loss. Search latency improved from 240ms to 96ms at p95.

  • In progress02:00 UTC

    Search on docs.halyard.dev falls back to keyword matching for the duration of the window.

Object Store returned 503s for multipart uploads in ap-southeast-1

OutageObject Storeap-southeast-1Resolved
  • Resolved17:29 UTC

    All upload paths healthy for 20 consecutive minutes. A full write-up is published in our postmortem archive.

  • Identified16:58 UTC

    A metadata shard failed to rejoin quorum after a routine host replacement. Draining traffic to the remaining two shards.

  • Investigating16:44 UTC

    Multipart uploads over 8 MB are failing in Singapore. Single-part uploads and all reads are unaffected.

Increased error rate on Compute Runners autoscale events

DegradedCompute RunnersResolved
  • Resolved11:52 UTC

    Scale-up requests are succeeding normally. We have added a rate limit to the placement controller.

  • Investigating11:05 UTC

    Roughly 2% of scale-up requests are returning a placement error. Existing workloads are not affected.

No incidents reported between 02 May and 29 May

All clear27 days

Get notified before your users notice

One email per incident, plus a weekly digest. Webhook, Slack and RSS delivery are available from the same preferences page.