AI Fleet Platform Software Engineer
location_onSunnyvale, California, United States; Toronto, Canadaschedule前天
sync_alt工作作風:油電混合
history最低經驗:12+ 年
職位描述
You will build and operate software that manages large fleets of AI clusters. You will create services, integrations, operational tools, and user-facing applications that help operators monitor health, capacity, performance, and incidents. You will lead projects from design through production, automate operational workflows, and improve reliability as the fleet grows.
Responsibilities
- Build and operate software for managing large fleets of AI clusters
- Provide operators with actionable views of cluster health, capacity, performance, and issues
- Develop services and integrations across infrastructure systems
- Automate incident investigation and service-restoration workflows
- Design reliable systems that withstand component and site failures
- Gather platform-user needs and make practical product and engineering decisions
- Lead projects from design through production and use operational feedback to improve them
Requirements
- 12+ years of industry experience building and operating production software for distributed systems or large-scale infrastructure
- Strong Go or Python skills
- Experience designing services and APIs
- Expertise in control planes, fleet management systems, or operational platforms
- Experience with Linux, containers, Kubernetes, and distributed-system failures
- Experience designing for asynchronous work, retries, and partial failures
- Experience with event streaming, workflow automation, or time-series telemetry
- Strong judgment in reliability, security, and observability
- Ability to lead ambiguous projects and collaborate across engineering and operations teams
提到的技能
申請這份工作
Use the application link supplied with this listing to apply to Cerebras Systems, Inc.. Check the destination before entering personal information.
