NRE/Network Architect
Job description
Job Summary
We are seeking a seasoned Network Reliability Engineer (NRE) with deep expertise in designing, operating, and optimizing hybrid & cloud network environments. You will focus on ensuring high availability, scalability, performance, and resilience of enterprise networks spanning on-premises data centers, SD-WAN, ACI, and major cloud platforms (AWS, Azure). This role combines network engineering with Reliability Engineering principle, emphasizing automation, observability, proactive incident prevention, and rapid recovery to meet stringent SLAs for mission-critical applications.
Responsibilities
· Design and implement highly reliable network infra.
· Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets.
· Lead major incident response,root cause analysis (RCA), and post-mortem reviews.
· Implement blamelesspost-mortems and drive corrective actions to prevent recurrence.
· Build automation scripts,tools, and self-healing mechanisms for network provisioning, configurationmanagement, monitoring, and failover.
· Develop and maintaincomprehensive dashboards, alerts, and logging using tools like Prometheus,Django, Grafana, Datadog, Splunk etc.
· Conduct network capacityplanning, performance tuning, and chaos engineering to validate resilience.
· Collaborate on network securityposture, zero-trust models, firewall policies, and compliance requirements.
· Work with DevOps, SRE, Cloud,Security, and Application teams to embed reliability into the developmentlifecycle.
· Guide junior engineers and contribute to knowledge sharing within the global team.
Requirements
· Bachelor Degree in Computer Science, Engineering, or related field (or equivalent experience).
· Preferred CCIE and/or CCNPcertification Core Networking
· Deep expertise in Routing (BGP,OSPF, EIGRP), Switching (STP, VLANs, VXLAN), Firewalls, Load Balancers, VPNs, and SD-WAN.
· Strong troubleshooting of complex Layer 2/3/4 issues.
· Experience applying SREprinciples (error budgets, toil reduction, automation).
· Proficiency in scripting(Python) and Infrastructure as Code (Terraform, Ansible, etc.).
· Prometheus, Grafana, Django,Datadog, etc.
· CI/CD & Automation -Jenkins, GitOps, Ansible.
· Packet analysis: Wireshark,tcpdump.
· Excellent problem-solving,communication, and stakeholder management
· Ability to work in a global,24x7 on-call rotation.
Must Have Skills
· Site Reliability engineering(SRE)
· Network Design and Engineering
· Cisco Application CentricInfrastructure (ACI)
· Network Automation
· Firewall
· Infrastructure as Code (IaC)
Skills mentioned
Apply for this job
Use the application link supplied with this listing to apply to AVATAR MODERN TECHNO SERVICES PTE. LTD.. Check the destination before entering personal information.
