Full-Time Senior Site Reliability Engineer
Stack AV is hiring a remote Full-Time Senior Site Reliability Engineer. The career level for this job opening is Experienced and is accepting USA based applicants remotely. Read complete job description before applying.
Stack AV
Job Title
Posted
Career Level
Career Level
Locations Accepted
Share
Job Details
Stack is developing revolutionary AI and advanced autonomous systems designed to enhance safety, reliability, and efficiency of modern operations.
Stack's autonomous technology incorporates cutting-edge advancements in artificial intelligence, robotics, machine learning, and cloud technologies, empowering us to create innovative solutions that address the needs and challenges of the dynamic trucking transportation industry.
About the Role:Stack AV Site Reliability Engineers are responsible for enabling and ensuring our production systems meet their service-level objectives. Through the implementation of centralized observability and automation, the SRE team constantly ensures the health, reliability, scalability, and performance of Stack AV’s infrastructure.
Responsibilities:- Monitor and maintain mission-critical production services to ensure maximum uptime.
- Design and implement scalable distributed systems to facilitate the development of self-driving vehicles.
- Design and implement an incident management framework and build a culture of blameless postmortems and continuous learning.
- Scale the reliability and velocity of our systems and processes through increased automation.
- Document actions to build a comprehensive library of runbooks, which will act as a knowledge base and foundation for automation.
- Participate in an on-call rotation to uphold the SLOs and SLAs of production services.
- Expertise in at least one scripting language (e.g. Bash, Python).
- Fundamental understanding of Linux operating system internals, TCP/IP networking, and storage subsystems.
- Experience scaling and securing services in the cloud (AWS, GCP) or cloud native environments.
- Experience using infrastructure-as-code principles to automate the creation of infrastructure resources (e.g. Terraform, CloudFormation).
- Understanding of engineering design limitations and ability to provide guidance to teams to scale their services to achieve desired performance within budget.
- Strong experience implementing and debugging cloud native and open source tools such as Kubernetes, etcd, Prometheus, OpenTelemetry, and Istio.
- Strong communication skills and the ability to work effectively in a diverse and distributed team.