Lead Site Reliability Engineer

location_onJersey City, New Jersey, United Statesschedule2 days ago
sync_altWork style:Hybrid
trending_upExperience level:Lead
badgeEmployment:Full-time
Apply Nowopen_in_new

Job description

As a Lead Site Reliability Engineer at JPMorgan Chase within the Asset and Wealth Management, Tech Production and Infrastructure Delivery team, you will be responsible for improving reliability, resilience, and operational performance across a hybrid technology environment spanning modern distributed platforms and mainframe systems.

Job Responsibilities

  • Lead adoption and operationalization of SRE practices, including SLIs/SLOs, error budgets, reliability reviews, and blameless post-incident processes.
  • Design, implement, and continuously improve monitoring and observability capabilities across metrics, logs, traces, and event telemetry to support faster detection and diagnosis.
  • Establish actionable alerting standards, dashboards, and runbooks to improve operational readiness and reduce noise.
  • Drive automation initiatives (self-service, self-healing, automated remediation, CI/CD operational controls, and standardized tooling) to reduce manual effort and improve consistency.
  • Identify, measure, and reduce operational toil through process optimization, tooling enhancements, and platform improvements.
  • Improve incident management practices, including incident response coordination, escalation paths, and continuous improvement based on root cause analysis.
  • Support capacity planning, performance engineering, and resilience testing to strengthen availability and service stability.
  • Partner with application, infrastructure, and operations teams across distributed and mainframe domains to standardize reliability patterns and operational controls.
  • Contribute to governance and operational excellence, including documentation, control evidence where applicable, and operational health reporting.

Required qualifications, capabilities and skills

  • Relevant experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, or a similar reliability-focused role, including leadership of technical initiatives.
  • Strong knowledge of operating and supporting distributed systems in production (e.g., Linux, networking, middleware, containers and/or cloud platforms).
  • Hands-on experience with monitoring/observability platforms and practices (metrics, logs, traces), including dashboarding and alert engineering.
  • Demonstrated ability to automate operational workflows using one or more scripting/programming languages (e.g., Python, Go, Shell) and standard automation approaches (CI/CD, infrastructure-as-code).
  • Experience supporting or integrating mainframe systems into enterprise operations (monitoring, incident response, operational processes).
  • Strong communication and stakeholder management skills, with ability to lead cross-team reliability improvements.

Preferred Qualifications

  • Experience establishing SLIs/SLOs and reliability reporting at service or platform level.
  • Familiarity with ITSM/incident tooling, on-call operations, and operational maturity improvements.
  • Experience with resilience patterns (graceful degradation, failover, rate limiting) and reliability testing (chaos testing, load/performance testing).
  • Exposure to regulated or high-control environments and operational risk management practices.

Skills mentioned


Banking, Financial Services
Jersey City, New Jersey, United States

JPMorgan Chase & Co. is an American multinational financial-services firm headquartered in New York City and the largest bank in the United States by assets. It operates across consumer and community banking, corporate and investment banking, commercial banking, and asset and wealth management.

Apply for this job

Use the application link supplied with this listing to apply to JP Morgan Chase. Check the destination before entering personal information.

Apply Nowopen_in_new