Principal Staff Site Reliability Engineer
location_onBoca Raton, Florida, United States; New York, United Statesschedule5 hari lalu
trending_upTahap pengalaman:Principal
historyPengalaman minimum:10+ tahun
Penerangan kerja
You will own reliability for critical services and set technical direction for hybrid cloud infrastructure. You will build infrastructure automation, improve observability, lead AI-assisted incident-response capabilities, strengthen production security, participate in on-call, and mentor senior engineers.
Responsibilities
- Own end-to-end reliability for business-critical services, including SLOs, error budgets, capacity planning, disaster recovery, and incident command
- Design and evolve multi-cloud and on-premises infrastructure across AWS, Azure, and colocated environments
- Build and maintain Terraform, Ansible, and CI/CD infrastructure
- Advance observability through metrics, logs, traces, and profiling
- Lead AI-assisted SRE capabilities for triage, incident summaries, runbooks, and root-cause analysis
- Harden production systems with security, data platform, and application teams
- Participate in on-call, run blameless postmortems, and drive systemic fixes
- Mentor senior engineers and lead infrastructure code and production-readiness reviews
Requirements
- 10+ years building and operating production infrastructure at scale
- Hybrid cloud and on-premises infrastructure experience
- AWS and Azure expertise
- Linux expertise
- Terraform
- Ansible
- CI/CD
- Python
- Go
- Kubernetes
- Service mesh
- Container security
- Datadog
- Prometheus
- Grafana
- OpenTelemetry
- ELK
- Splunk
- Incident command
Kemahiran yang dinyatakan

Boca Raton, Florida, United States; New York, United States
Mohon kerja ini
Use the application link supplied with this listing to apply to DigitalBridge Group, Inc.. Check the destination before entering personal information.
