Oh My JobFind Jobs
Company
  • About Us
  • Blog
  • Contact
Tools
  • Paycheck Calculator
Legal
  • Terms of Service
  • Privacy Policy
  • California Privacy Rights
For Employers
Post a Job

Site Reliability Engineer

Mistral AI
Mistral AI
New York
Jul 10, 2026
Salary not listed
FullTime

Job Description

About the Role

Mistral is seeking experienced Site Reliability Engineers to own the reliability, scalability, and performance of our AI platform and customer-facing applications. You will partner with software engineers and research teams to ensure systems meet and exceed expectations for internal and external customers across high-stakes industries including finance, manufacturing, defense, healthcare, and the public sector.

Responsibilities

Operations

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructure for web services and ML workloads
  • Ensure platform, inference, and model training environments remain highly available with seamless replication across HPC clusters
  • Operate production systems and troubleshoot issues including incident response, on-call rotations, infrastructure scaling, and root cause analysis
  • Implement and enhance monitoring, alerting, and incident response systems to optimize performance and minimize downtime
  • Build and maintain CI/CD pipelines, containerization, orchestration, monitoring, logging, and alerting systems for client-facing APIs and large training runs

Development

  • Drive infrastructure automation improvements using Kubernetes, Flux, and Terraform
  • Collaborate with AI/ML researchers to develop solutions enabling safe, reproducible model-training experiments
  • Build cloud-agnostic platform abstractions between science and infrastructure
  • Design and develop workflows, automation scripts, and tooling to improve system reliability, availability, and performance
  • Work with the security team to ensure infrastructure follows security best practices and compliance requirements
  • Document processes and share knowledge across the team
  • Contribute to open-source projects, research publications, and technical content

Requirements

  • Master's degree in Computer Science, Engineering, or related field
  • 7+ years of experience in DevOps or SRE roles
  • Strong experience with cloud computing and highly available distributed systems
  • Demonstrated expertise with site reliability in critical environments (root cause analysis, production troubleshooting, on-call rotations)
  • Experience working with reliability KPIs, observability, alerting, and SLAs
  • Hands-on experience with CI/CD, containerization, and orchestration tools (Docker, Kubernetes)
  • Knowledge of monitoring, logging, alerting, and observability platforms (Prometheus, Grafana, ELK Stack, Datadog)
  • Familiarity with infrastructure-as-code tools (Terraform, CloudFormation)
  • Proficiency in scripting languages (Python, Go, Bash) and software development best practices
  • Strong understanding of networking, security, and system administration
  • Excellent problem-solving and communication skills
  • Ability to thrive in a fast-paced startup environment

Preferred Qualifications

  • Experience in AI/ML environments
  • Background with high-performance computing (HPC) systems and workload managers (Slurm)
  • Familiarity with AI-oriented infrastructure solutions (Fluidstack, Coreweave, Vast)

Benefits

Mistral offers a comprehensive benefits package that varies by location and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal allowances, transportation benefits, and other location-specific perks. See the Benefits page for details specific to your location.

Mistral AI on Oh My Job

14 open positions right now, including 4 in California.

Apply now
Share: