Principal Staff Site Reliability Engineer

location_onBoca Raton, Florida, United States; New York, United Statesschedule5 天前
trending_up經驗等級:Principal
history最低經驗:10+ 年

職位描述

You will own reliability for critical services and set technical direction for hybrid cloud infrastructure. You will build infrastructure automation, improve observability, lead AI-assisted incident-response capabilities, strengthen production security, participate in on-call, and mentor senior engineers.

Responsibilities

  • Own end-to-end reliability for business-critical services, including SLOs, error budgets, capacity planning, disaster recovery, and incident command
  • Design and evolve multi-cloud and on-premises infrastructure across AWS, Azure, and colocated environments
  • Build and maintain Terraform, Ansible, and CI/CD infrastructure
  • Advance observability through metrics, logs, traces, and profiling
  • Lead AI-assisted SRE capabilities for triage, incident summaries, runbooks, and root-cause analysis
  • Harden production systems with security, data platform, and application teams
  • Participate in on-call, run blameless postmortems, and drive systemic fixes
  • Mentor senior engineers and lead infrastructure code and production-readiness reviews

Requirements

  • 10+ years building and operating production infrastructure at scale
  • Hybrid cloud and on-premises infrastructure experience
  • AWS and Azure expertise
  • Linux expertise
  • Terraform
  • Ansible
  • CI/CD
  • Python
  • Go
  • Kubernetes
  • Service mesh
  • Container security
  • Datadog
  • Prometheus
  • Grafana
  • OpenTelemetry
  • ELK
  • Splunk
  • Incident command

提到的技能


Boca Raton, Florida, United States; New York, United States

申請這份工作

Use the application link supplied with this listing to apply to DigitalBridge Group, Inc.. Check the destination before entering personal information.