HD

Site Reliability Engineer

Halcyon Data

Actively hiring🇺🇸 Seattle, WA, United StatesHybridFull-time$170k – $250k /yr5–9 yrs expDevOps Engineer
Posted 21d agoBe an early applicant

About the role

Halcyon Data is hiring a platform engineer in Seattle, WA. This role owns the path from a merged pull request to a running, observable service, and the on-call experience that follows it. Because a mistake here reaches every team at once, we care about production experience and blast-radius thinking far more than about certifications.

The US market for this role is deep and well paid, and candidates rightly compare offers carefully. The band below is the real band — we do not hold a number back for whoever negotiates hardest.

What the work actually looks like

You will work on CI/CD, infrastructure as code, and the observability stack, and you will spend real time removing the toil the rest of engineering has quietly learned to live with. Incidents are reviewed blamelessly and the actions that come out of them are actually scheduled. The team works a hybrid week out of Seattle, WA, usually three days together and the rest wherever you work best.

What we look for

This is a senior individual-contributor role. You will be expected to push back on requirements that do not make sense and to set the technical direction for your part of the product. We are looking for someone fluent in Kubernetes and Terraform who can also write decent code, and who treats a runbook as a product with users. Experience defining SLOs and making them mean something is a strong signal.

Working at Halcyon Data

Core hours are set by the team rather than by the company, and we are deliberate about protecting long stretches of focus time. The range shown is the annual gross salary band for this role.

Responsibilities

  • Own CI/CD pipelines and keep the path to production fast and reversible
  • Manage infrastructure as code with reviewed, repeatable changes
  • Run the observability stack: metrics, logs, traces and alerting that people trust
  • Reduce operational toil and make the on-call rotation humane
  • Lead incident response and turn reviews into scheduled work

Requirements

  • Production experience with Kubernetes and a major cloud provider
  • Terraform or an equivalent infrastructure-as-code tool
  • Comfortable scripting and writing small services, not only configuration
  • Experience with Prometheus, Grafana or a comparable observability stack
  • Calm, methodical approach to incidents and post-incident follow-up

Skills

Benefits

  • Medical, dental and vision cover from your first day
  • 401(k) with company match
  • Equity in a company with real revenue
  • Paid time off with a mandatory two-week minimum