Full-Time Site Reliability Engineer (SRE/DevOps)
Arista Networks is hiring a remote Full-Time Site Reliability Engineer (SRE/DevOps). The career level for this job opening is Experienced and is accepting Bengaluru, India based applicants remotely. Read complete job description before applying.
Arista Networks
Job Title
Posted
Career Level
Career Level
Locations Accepted
Share
Job Details
Who You'll Work WithArista Networks is looking for a skilled professional for our Engineering Productivity (EngProd) team to help maintain and support our rapidly expanding infrastructure and internal user base. The ideal candidate is someone who can wear many hats, is versatile, and is enthusiastic about learning new technologies.
As a part of the software engineering team, you will work with other team members to design, build, and administer secure, scalable, and fault-tolerant tools and infrastructure in a hybrid cloud environment.
Working in the EngProd group, you will collaborate with other engineers to design, build, scale, and operate the systems used by Arista’s product development teams.
What You'll Do
- Build, deploy safely and incrementally, and operate critical production systems with a focus on scalability, reliability, observability, performance, and security.
- Monitor, support, and enhance developer experience across services.
- Build automation to remove toil and efficiently operate production systems.
- Proactively monitor, respond to, and enhance alerts and set up automated alert handling.
- Create and maintain the incident response runbooks.
- Build and deploy new systems with scalability, reliability, and observability as primary requirements.
- Triage platform/infrastructural issues and help Arista software engineers in their triages. Engage with 3rd party vendor support.
- Deploy new systems in a staged manner.
- Write postmortem documents and build solutions to avoid incidents from repeating.
- Plan and communicate maintenance windows on production systems.
- Work with Arista’s product development teams to identify infrastructural issues that are causing bottlenecks and limitations in their workflows. Design and implement solutions to resolve them.
- Survey and adopt best practices around infrastructure/platform to maintain secure, scalable, and fault-tolerant systems.
- Implement solutions to scale the systems.
- Implement fault-tolerance and performance to improve availability of the systems.
- Study the design and sufficient implementation details of OSS systems for better triage and fix resolution.