Back to Jobs
Synopsys

Site Reliability, Staff

Synopsys
Noida, IndiaFull TimeSeniorPosted Today

General Information

Job Title Site Reliability, Staff Job ID 15510 Country India City Noida Date Posted 26-Feb-2026 Job Category Information Technology Job Subcategory Site Reliability Hire Type Employee Remote Eligible No

Descriptions & Requirements

Job Description and Requirements

We Are

We are a technology organization operating high-performance, large-scale Linux production environments that support critical platforms and engineering teams. Our focus is on operational excellence, service reliability, automation, and continuous improvement. We run 24x7 operations and partner closely with platform, network, security, and engineering teams to deliver stable, secure, and scalable infrastructure.

You Are

You are a senior individual contributor and technical leader with deep expertise in Linux operations and Site Reliability Engineering. You thrive in high availability, always-on environments, are comfortable leading complex production incident response as a hands-on expert, and are passionate about automation, reliability, and raising operational maturity across teams. You influence through technical direction, strong execution, and cross-functional partnership (not people management).

What You’ll Be Doing

  • Serving as a hands-on Staff SRE supporting and improving 24x7 Linux production operations (with participation in on-call / incident rotations as needed)
  • Acting as a senior technical escalation point during major production incidents, driving triage, mitigation, and root cause analysis (RCA)
  • Leading reliability improvements across Linux OS, bare metal and virtualized platforms, and monitoring/logging systems
  • Driving adherence to SLAs, OLAs, and operational KPIs such as availability, MTTR, incident volume, and toil reduction—using data to prioritize work
  • Designing and implementing automation using Ansible, Bash, and Python to reduce manual effort and improve consistency
  • Defining, improving, and standardizing SOPs, runbooks, escalation procedures, and operational documentation
  • Partnering with platform, network, security, and engineering teams to improve system resilience, observability, and operational readiness
  • Identifying systemic risks and reliability gaps, proposing solutions, and leading execution across teams (technical leadership without direct reports)
  • Mentoring and enabling L1/L2 engineers through technical guidance, reviews, incident coaching, and knowledge sharing

The Impact You Will Have

  • Improving stability, reliability, and efficiency of 24x7 Linux/SRE production operations
  • Reducing incident recurrence through strong RCA, corrective actions, and reliability engineering practices
  • Increasing operational maturity through automation, standardization, and high-quality documentation
  • Improving incident response quality and speed via better observability, runbooks, and escalation paths
  • Enabling engineering and R&D teams through predictable, resilient, and well-operated platform services

What You’ll Need

  • 10–14+ years of experience in IT Infrastructure, Linux Operations, or Site Reliability Engineering
  • Strong hands-on background in Linux system administration and production support in large-scale environments
  • Proven experience leading incident response, on-call models, and operational processes in 24x7 environments
  • Advanced knowledge of Linux OS internals and troubleshooting (performance, networking, storage, kernel/user space behaviors)
  • Experience with virtualization platforms (VMware, KVM, OpenStack, oVirt) and production operations for virtualized/bare metal fleets
  • Knowledge of monitoring and logging tools (e.g., Nagios, ELK or similar observability stacks)
  • Experience with automation and configuration management (Ansible)
  • Scripting skills in Bash and/or Python

Who You Are

  • Calm and effective under high-pressure production scenarios
  • Highly structured and data-driven in driving operational excellence and prioritization
  • An effective communicator who can influence stakeholders across engineering and corporate functions
  • Passionate about reliability engineering, automation, and continuous improvement
  • Comfortable owning ambiguous, cross-team problems end-to-end and driving them to closure

The Team You’ll Be A Part Of

You will be part of a Linux Engineering / Site Reliability Engineering organization responsible for frontline production support and reliability improvements. The team works closely with L2/L3 engineering, platform, network, security, and R&D teams to ensure reliable and scalable infrastructure operations across the business.

Rewards and Benefits

  • Opportunity to drive mission-critical, large-scale Linux and SRE reliability initiatives as a senior IC
  • High visibility role with exposure to senior leadership and engineering stakeholders
  • Ability to shape operational strategy, automation, and reliability practices through technical leadership
  • Strong focus on career growth, learning, and technical leadership development

At Synopsys, we want talented people of every background to feel valued and supported to do their best work. Synopsys considers all applicants for employment without regard to race, color, religion, national origin, gender, sexual orientation, age, military veteran status, or disability.

Ready to apply? You'll be taken to Synopsys's application page.
Site Reliability, Staff at Synopsys