Full-Time Staff Site Reliability Engineer

Wikimedia Foundation is hiring a remote Full-Time Staff Site Reliability Engineer. The career level for this job opening is Experienced and is accepting USA, UK, Canada, Germany, France, India, Brazil, Australia, Estonia, Sweden, Greece, Belgium, Poland based applicants remotely. Read complete job description before applying.

This job was posted 4 months ago and is likely no longer active. We encourage you to explore more recent opportunities on our site. However, you may still try your luck using 'Apply Now' link below. We recommend focusing on newer listings available here.

Wikimedia Foundation

Job Title

Staff Site Reliability Engineer

Posted

4 months ago on 4th August 2025

Career Level

Full-Time

Career Level

Experienced

Locations Accepted

USA, UK, Canada, Germany, France, India, Brazil, Australia, Estonia, Sweden, Greece, Belgium, Poland

Salary

YEAR $129347 - $200824

Job Details

The Wikimedia Foundation is looking for a Staff Site Reliability Engineer (SRE) focused on Machine Learning Infrastructure. You will join a distributed team working across UTC -5 to UTC +3 (Eastern Americas, Europe, and Africa) and report directly to the Director of Machine Learning.

You will be responsible for:

Designing and implementing robust ML infrastructure used for training, deployment, monitoring, and scaling of machine learning models.
Improving reliability, availability, and scalability of ML infrastructure, ensuring smooth and efficient workflows for internal ML engineers and researchers.
Collaborating closely with ML engineers, product teams, researchers, SREs, and the Wikimedia volunteer community to identify infrastructure requirements, resolve operational issues, and streamline the ML lifecycle.
Proactively monitoring and optimizing system performance, capacity, and security to maintain high service quality.
Providing expert guidance and documentation to teams across Wikimedia to effectively utilize the ML infrastructure and best practices.
Mentoring team members and sharing knowledge on infrastructure management, operational excellence, and reliability engineering.

Skills and Experience:

7+ years of experience in Site Reliability Engineering (SRE), DevOps, or infrastructure engineering roles, with substantial exposure to production-grade machine learning systems.
Proven expertise with on-premises infrastructure for machine learning workloads (e.g., Kubernetes, Docker, GPU acceleration, distributed training systems).
Strong proficiency with infrastructure automation and configuration management tools (e.g., Terraform, Ansible, Helm, Argo CD).
Experience implementing observability, monitoring, and logging for ML systems (e.g., Prometheus, Grafana, ELK stack).
Familiarity with popular Python-based ML frameworks (e.g., PyTorch, TensorFlow, scikit-learn).
Strong English communication skills and comfort working asynchronously across global teams.

Skills

Docker ELK Stack Grafana Kubernetes Terraform

FAQs

What is the last date for applying to the job?

The deadline to apply for Full-Time Staff Site Reliability Engineer at Wikimedia Foundation is 3rd of September 2025 . We consider jobs older than one month to have expired.

Which countries are accepted for this remote job?

This job accepts [ USA, UK, Canada, Germany, France, India, Brazil, Australia, Estonia, Sweden, Greece, Belgium, Poland ] applicants. .

Apply Now

Related Jobs You May Like

Azure DevOps Engineer

Jersey City, NJ

2 days ago

.NET

Azure

DevOps

Derex Technologies Inc

Full-Time

Experienced

Lead Palantir Developer

Seattle, WA

2 days ago

CI/CD Pipelines

Data Engineering

Palantir Foundry

Logic20/20 Inc.

Full-Time

Experienced

YEAR $156750 - $173329

Cloud AppOps Engineer

Atlanta, GA

3 days ago

Application Support

AWS

Cloud Services (EC2, S3, IAM, ELB, VPC, VPN)

Sutherland

Full-Time

Experienced

Staff DataOps Engineer

Remote, India

3 days ago

AWS

CI/CD

DataOps

Nagarro

Full-Time

Experienced

Query Tuning Specialist - Database Performance - Postgre

Austin, Texas

3 days ago

Database Management

Performance Tuning

Problem-solving

ServiceNow

Full-Time

Experienced

DevOps Engineer, Playout

New York, New York

3 days ago

CICD

Cloud Services (AWS, GCP, Azure)

DevOps

NBCUniversal

Full-Time

Experienced

YEAR $90000 - $110000

Query Tuning Specialist - Database Performance - Postgres

Austin, Texas

3 days ago

Database Management

Performance Tuning

SaaS/PaaS/Cloud Development

ServiceNow

Full-Time

Experienced

Lead Palantir Developer

Seattle, WA

4 days ago

CI/CD Pipelines

Cloud ETL

Palantir Foundry

Logic20/20 Inc.

Full-Time

Experienced

YEAR $156750 - $173329

Cloud AppOps Engineer

Atlanta, GA

4 days ago

Application Support

AWS

Cloud Security

Sutherland

Full-Time

Experienced

Site Reliability Engineer

Stamford, Connecticut

4 days ago

Cloud Platforms (AWS, GCP, Azure)

Configuration Management

Monitoring And Alerting Tools

NBCUniversal

Full-Time

Experienced

YEAR $110000 - $145000

Senior Cloud Platform Engineer (Networking)

Berlin, Germany

5 days ago

AWS

Networking

Scalable GmbH

Full-Time

Experienced

DevOps Engineer

Texas

5 days ago

AWS

GitLab

Kubernetes

InfStones

Full-Time

Experienced

All Remote Jobs

Full-Time Staff Site Reliability Engineer

Wikimedia Foundation

Job Title

Posted

Career Level

Career Level

Locations Accepted

Salary

Share

Job Details

Skills

FAQs

What is the last date for applying to the job?

Which countries are accepted for this remote job?

Related Jobs You May Like

Azure DevOps Engineer

Lead Palantir Developer

Cloud AppOps Engineer

Staff DataOps Engineer

Query Tuning Specialist - Database Performance - Postgre

DevOps Engineer, Playout

Query Tuning Specialist - Database Performance - Postgres

Lead Palantir Developer

Cloud AppOps Engineer

Site Reliability Engineer

Senior Cloud Platform Engineer (Networking)

DevOps Engineer

Looking for a specific job?