Site Reliability, Staff - 18350

30 days ago
Experience level:Staff
Minimum experience:5+ years
Apply Now

Job description

We Are

Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to rapidly innovate AI-powered products. We deliver industry-leading silicon design, IP, simulation and analysis solutions, and design services. We partner closely with our customers across a wide range of industries to maximize their R&D capability and productivity, powering innovation today that ignites the ingenuity of tomorrow.

You Are

You have spent years keeping Linux infrastructure running under real pressure, the kind where a compute farm going down at 2 a.m. means a tape-out slips and someone loses sleep. You know the difference between a system that looks healthy in a dashboard and one that actually stays up when 10,000 simulation jobs hit it at once. You have debugged enough weird failures to know that the answer is rarely in the first log you check, and you are comfortable digging until you find it.

Working across cloud and on-prem does not rattle you. You have built things in Azure or AWS using Terraform and Ansible, not because it was trendy but because manual provisioning does not scale. You script in Bash and Python not to show off but to stop doing the same thing twice. When a designer in Singapore cannot access a license server, you do not guess, you pull tcpdump, check journalctl, trace the route, and fix it.

You care about uptime because you know the teams depending on your infrastructure are racing deadlines that matter. At Synopsys, you will support compute environments that power chip design work across the company, and what you keep running directly affects whether engineers can do their jobs.

What You'll Be Doing

Administer and maintain CentOS/ALMA Linux systems across Azure, AWS, and GCP cloud environments that support large-scale EDA compute farms and design workflows

Deploy and manage cloud infrastructure at scale using Terraform and Ansible, handling compute, storage, and networking for simulation, synthesis, and verification workloads

Monitor system health, capacity, and performance using Prometheus, Grafana, Elastic Search, or Splunk, catching issues before they affect design teams

Troubleshoot complex system, network, and application-level problems using tcpdump, netstat, journalctl, and other diagnostic tools to minimize downtime during critical tape-out windows

Manage license servers (FlexLM/FlexNet) and batch scheduling systems like LSF, SGE, Slurm, or Altair PBS that queue simulation and regression jobs

Coordinate tool version upgrades, freeware installs, and compatibility validation with CAD and EDA application support teams

Participate in rotational shift work (monthly rotation) and on-call weekend coverage, providing front-line support to internal design and verification teams

The Impact You Will Have

Keep compute infrastructure running so IC design, verification, and physical design teams can meet tape-out deadlines without infrastructure delays

Reduce incident response time and system downtime through proactive monitoring, automation, and structured root-cause analysis

Scale cloud infrastructure efficiently using Infrastructure as Code, enabling faster provisioning and more reliable deployments across global teams

Improve system reliability and performance tuning so simulation and synthesis jobs run faster and more predictably

Build runbooks and documentation that help junior admins resolve issues independently and reduce repeat escalations

Strengthen cross-team collaboration by translating complex infrastructure problems into clear updates for CAD, DevOps, security, and design engineering teams

Support the adoption of AI-powered tools and workflows by maintaining the underlying compute and cloud infrastructure they depend on

What You'll Need

5+ years of hands-on Linux systems administration experience, ideally supporting compute-intensive or EDA environments

Strong proficiency with CentOS, ALMA, or similar Linux distributions in production cloud environments

Hands-on experience with Azure (preferred), and working knowledge of AWS or GCP

Proven experience using Terraform and Ansible to deploy and manage large-scale compute and storage infrastructure

Strong scripting skills in Bash and Python for automation, troubleshooting, and workflow optimization

Experience managing batch scheduling systems such as LSF, SGE/Univa Grid Engine, Slurm, or Altair PBS

Familiarity with license management tools like FlexLM or FlexNet, and monitoring/logging platforms such as Prometheus, Grafana, Elastic Search, or Splunk is a plus

Who You Are

You can troubleshoot a network connectivity issue using tcpdump, netstat, traceroute, and dig without needing a runbook in front of you

You write scripts to automate repeated tasks because doing the same manual work twice feels like a waste of time

You treat internal design teams like customers, their productivity depends on your systems staying up, and you take that seriously

You can translate a complex infrastructure failure into a two-sentence update for a VP without losing the technical nuance that matters

You stay calm during on-call incidents and work methodically through root-cause analysis instead of guessing and hoping

You are comfortable coordinating across network, security, CAD, and DevOps teams to get a problem solved, even when it crosses multiple domains

The Team You'll Be Part Of

You will be part of a dynamic and collaborative Cloud Operations team that keeps our cloud environment running smoothly, securely, and reliably. Operating 24×5, we play a critical role in ensuring service availability, proactively monitoring infrastructure, and responding swiftly to incidents. We believe in teamwork, ownership, continuous learning, and a proactive approach to problem-solving. Every team member contributes to operational excellence by embracing automation, driving improvements, and sharing knowledge. Together, we don’t just keep the cloud running—we continuously evolve, innovate, and deliver reliable services that make a real difference to the business.

Rewards and Benefits

We offer a comprehensive range of health, wellness, and financial benefits to cater to your needs. Our total rewards include both monetary and non-monetary offerings. Your recruiter will provide more details about the salary range and benefits during the hiring process.

Skills mentioned


AI
synopsys.com

Founded in 1986 by visionaries Aart de Geus, David Gregory, and Bill Krieger, Synopsys started its journey in Research Triangle Park, North Carolina, initially named Optimal Solutions. The company's pioneering work in logic synthesis technology quickly positioned it as a leader in the field of electronic design automation (EDA). With a keen focus on innovation, Synopsys rapidly transitioned, changing its name and relocating to Mountain View, California. Through an initial public offering in 1992, Synopsys established itself as a prominent player in the semiconductor and systems design industry. Over the years, Synopsys has significantly expanded its portfolio, which now includes cutting-edge tools for chip design, intellectual property (IP) solutions, and robust software security offerings. The company has harnessed the power of artificial intelligence to enhance its software capabilities, launching products like DSO.ai for chip design automation and Synopsys.ai Copilot for semiconductor design optimization. With a workforce of around 20,000, and annual revenues exceeding $6 billion, Synopsys continues to be at the forefront of technological advancement. It aims to empower technology innovators globally, providing the tools necessary to tackle ever-increasing design complexities and accelerating the path to market for semiconductor solutions.

Apply for this job

Use the application link supplied with this listing to apply to Synopsys. Check the destination before entering personal information.

Apply Now