AI Fleet Platform Software Engineer
location_onSunnyvale, California, United States; Toronto, Canadaschedule一昨日
sync_alt働き方:ハイブリッド
history必要経験年数:12+ 年
仕事内容
You will build and operate software that manages large fleets of AI clusters. You will create services, integrations, operational tools, and user-facing applications that help operators monitor health, capacity, performance, and incidents. You will lead projects from design through production, automate operational workflows, and improve reliability as the fleet grows.
Responsibilities
- Build and operate software for managing large fleets of AI clusters
- Provide operators with actionable views of cluster health, capacity, performance, and issues
- Develop services and integrations across infrastructure systems
- Automate incident investigation and service-restoration workflows
- Design reliable systems that withstand component and site failures
- Gather platform-user needs and make practical product and engineering decisions
- Lead projects from design through production and use operational feedback to improve them
Requirements
- 12+ years of industry experience building and operating production software for distributed systems or large-scale infrastructure
- Strong Go or Python skills
- Experience designing services and APIs
- Expertise in control planes, fleet management systems, or operational platforms
- Experience with Linux, containers, Kubernetes, and distributed-system failures
- Experience designing for asynchronous work, retries, and partial failures
- Experience with event streaming, workflow automation, or time-series telemetry
- Strong judgment in reliability, security, and observability
- Ability to lead ambiguous projects and collaborate across engineering and operations teams
記載されたスキル
この求人に応募する
Use the application link supplied with this listing to apply to Cerebras Systems, Inc.. Check the destination before entering personal information.
