Engineering Manager, Production Platform and Orchestration
Job description
Role Overview
Positron is seeking an Engineering Manager to lead Production Platform and Orchestration within our Upstack Engineering organization. This team owns the software and operating practices that provision, deploy, observe, upgrade, and reliably operate Positron systems in production. You will inherit a technically strong core team and help it grow into a durable organization capable of supporting a fleet that is expanding by several multiples.
This is a technical leadership role with real operational accountability. You will set direction, build the team, create clear ownership, and improve the systems and processes behind fleet orchestration, deployment lifecycle, observability, incident response, release automation, and production reliability. You will work closely with serving and API, model enablement, compiler and runtime, hardware, customer-facing, and data center partners.
The strongest candidate will combine systems depth with organizational judgment, moving comfortably between architecture, delivery, incidents, people development, and cross-functional planning. This description intentionally emphasizes outcomes and ownership over a fixed organizational chart. As the fleet and customer base grow, the function may develop dedicated groups for fleet orchestration and capacity, deployment lifecycle, reliability and observability, data center operations, customer production operations, and operational tooling.
Key Responsibilities
Team Leadership and Organizational Growth
- Lead, coach, and grow a team of engineers spanning production platform, fleet orchestration, reliability, and operational automation.
- Hire thoughtfully, develop emerging leaders, and create ownership boundaries that remain effective as the organization scales.
- Translate customer and business priorities into sequenced engineering work while protecting the team from reactive, unstructured operations.
Platform Strategy and Roadmap
- Establish a clear technical and organizational roadmap for provisioning, deployment, environment lifecycle, fleet health, capacity, upgrades, and rollback.
- Build reliable orchestration and control-plane capabilities for inventory, placement, configuration, health management, and multi-system operations.
- Prepare the production platform for a heterogeneous accelerator environment that may include FPGA, ASIC, and GPU infrastructure.
Reliability and Production Operations
- Define service-level objectives, operational metrics, alerting standards, and a sustainable on-call model for customer-facing production systems.
- Own the operating cadence for incidents, escalations, postmortems, corrective actions, launch readiness, and reliability reviews.
- Increase automation across deployment, upgrades, remediation, diagnostics, capacity planning, and common support workflows so that fleet growth does not require linear headcount growth.
Cross-Functional Partnership
- Partner with hardware and data center teams on rack bring-up, networking, firmware, sparing, failure handling, and platform transitions.
- Collaborate with Distributed Serving and API, Model Enablement, and Compiler and Executor teams to turn new capabilities into supportable production services.
Required Qualifications
- Demonstrated success managing and growing engineering teams responsible for distributed systems, cloud infrastructure, production platforms, SRE, or a closely related domain.
- Strong technical judgment across Linux systems, networking, orchestration, deployment systems, observability, and production reliability.
- A proven record of turning ambiguous operational demands into a coherent roadmap, explicit ownership, and measurable engineering outcomes.
- Ability to recruit, coach, and retain engineers across experience levels while maintaining a high technical bar.
- Comfort operating across software, hardware, data center, customer, and business boundaries.
- Excellent written and verbal communication skills, with sound prioritization and the ability to make tradeoffs visible to technical and executive stakeholders.
- A hands-on leadership style, close enough to architecture and operations to ask the right questions without becoming the team's bottleneck.
Preferred Qualifications
- Hands-on experience operating GPU, FPGA, ASIC, or other accelerator fleets in production.
- Ownership of services with meaningful availability expectations, including on-call, incident management, root-cause analysis, and reliability planning.
- Experience building control planes, schedulers, placement systems, capacity-management systems, or multi-rack orchestration.
- Background in data center deployment, hardware lifecycle, firmware coordination, sparing and RMA processes, or production networking.
- A track record of scaling infrastructure from early deployments to multiple sites or hundreds of systems.
- Exposure to customer-facing infrastructure where engineering teams participate in production escalation and service readiness.
- Demonstrated automation work that materially reduced operational toil, incident frequency, or recovery time.
Leveling & Scope
While this role is currently posted at a specific level, we are a growth-oriented organization and are open to hiring at a more senior level for the right candidate. Please note that this job description serves as a focused but generalized overview of the role; specific responsibilities and impact expectations will be tailored to the experience and seniority of the final hire.
What Success Looks Like
In the first six months, you will build trust with the team and partner organizations, clarify ownership, decision rights, and the near-term hiring plan, and baseline fleet health, incident load, deployment reliability, operational toil, and the largest single points of failure. You will establish a practical operating cadence for on-call, incident review, release readiness, and reliability prioritization, and produce an agreed roadmap that balances immediate production needs with platform investments and automation. Between six and twelve months, you will grow the team and create durable ownership for fleet orchestration, deployment lifecycle, observability, and production reliability, while improving automated provisioning, upgrades, rollback, health monitoring, and operational diagnostics. Avoidable incidents and manual intervention will decrease, deployment confidence and customer readiness will increase, and you will be developing engineers and technical leads who can independently own major platform and operational domains. By twelve to eighteen months, you will be operating a resilient production organization with clear specialties, healthy management span, and sustainable coverage, supporting a substantially larger and more diverse fleet without proportional growth in operational effort, and demonstrating measurable improvement in availability, deployment speed, upgrade safety, incident recovery, and operational efficiency.
Why Join Us?
- You will build the production platform that turns purpose-built inference silicon into services customers can depend on, with direct ownership of how a rapidly growing fleet is deployed, operated, and scaled.
- You will shape both the technology and the organization from an early stage, defining the orchestration, reliability, and automation foundations that Positron will operate on for years to come.
Compensation & Benefits
The base salary range for this role is $200,000 – $300,000.
Please note that the figures provided represent the base salary range only and do not include other elements of our total compensation package, equity, or comprehensive benefits.
At Positron AI, we value the unique expertise each candidate brings. While the range above reflects our typical expectation for the position, we reserve the flexibility to exceed this range for candidates whose specialized skills, significant experience, or unique qualifications fall outside the standard scope of the role. Final offers are determined based on a variety of factors, including internal equity, and individual impact.
Benefits & Perks
We want you to do your best work and feel confident that you and your family are taken care of. That means comprehensive coverage, real time to rest, and support for your future.
Health and wellness
- Fully company-paid medical, dental, and vision insurance for you and your dependents
- Company-paid life and disability coverage, with voluntary options to add more
- Supplemental hospital, critical illness, and accident coverage available
Time off and flexibility
- Unlimited paid time off, we encourage everyone to truly unplug and recharge
- 13 paid company holidays
- Remote-first culture with a company-provided computer and home office setup
Compensation and future
- Competitive salary and equity
- 401(k) with company matching, eligible from day one
Visa Support
This position is open to candidates currently authorized to work in the U.S. We cannot provide new visa sponsorship for this role but are open to facilitating H-1B visa transfers for eligible candidates.
Equal Opportunity Employer. If you're excited about the role but don't meet every bullet, we'd still love to hear from you.
Skills mentioned

Positron AI is a hardware company at the forefront of accelerating artificial intelligence, with a mission to make advanced machine learning more accessible and efficient. Founded in the spring of 2023, the company is dedicated to providing solutions that offer superior performance per dollar and enhanced energy efficiency. All of Positron's products are designed, fabricated, and assembled in the United States, emphasizing a commitment to domestic manufacturing and supply chain security. The team at Positron brings together over 400 years of combined experience in AI, systems, silicon, and cloud technologies. The company was established to address the growing problem of GPUs becoming a financial burden for organizations deploying AI. The core mission is to deliver inference that is highly efficient, affordable, and proudly American-made, with the ultimate goal of making GPUs an optional component in the AI technology stack. The company's first-generation product, Atlas, is the world's first accelerator designed specifically for LLM-inference. Developed and shipped within 18 months of the company's inception, Atlas was created to tackle the cost and energy constraints that are hindering growth in the AI sector. Positron is already developing its second-generation system, Titan, which aims to be even faster and more efficient. By leveraging the insights gained from Atlas, Titan is being designed to unlock virtually limitless context and enable the concurrent operation of numerous models and AI agents, supported by terabytes of memory per accelerator. Positron's innovative hardware architecture is engineered to resolve the power, memory, and scalability challenges inherent in legacy infrastructures, thereby offering the lowest total cost of ownership for transformer models. The company's systems are compatible with Hugging Face transformer models and serve inference requests through an OpenAI API compatible endpoint, ensuring seamless integration into existing AI workflows.
Apply for this job
Use the application link supplied with this listing to apply to Positron AI. Check the destination before entering personal information.
