Job Description
About the Role
Mistral is seeking experienced Site Reliability Engineers to own the reliability, scalability, and performance of our AI platform and customer-facing applications. You will partner with software engineers and research teams to ensure systems meet and exceed expectations for internal and external customers across high-stakes industries including finance, manufacturing, defense, healthcare, and the public sector.
Responsibilities
Operations
- Design, build, and maintain scalable, highly available, and fault-tolerant infrastructure for web services and ML workloads
- Ensure platform, inference, and model training environments remain highly available with seamless replication across HPC clusters
- Operate production systems and troubleshoot issues including incident response, on-call rotations, infrastructure scaling, and root cause analysis
- Implement and enhance monitoring, alerting, and incident response systems to optimize performance and minimize downtime
- Build and maintain CI/CD pipelines, containerization, orchestration, monitoring, logging, and alerting systems for client-facing APIs and large training runs
Development
- Drive infrastructure automation improvements using Kubernetes, Flux, and Terraform
- Collaborate with AI/ML researchers to develop solutions enabling safe, reproducible model-training experiments
- Build cloud-agnostic platform abstractions between science and infrastructure
- Design and develop workflows, automation scripts, and tooling to improve system reliability, availability, and performance
- Work with the security team to ensure infrastructure follows security best practices and compliance requirements
- Document processes and share knowledge across the team
- Contribute to open-source projects, research publications, and technical content
Requirements
- Master's degree in Computer Science, Engineering, or related field
- 7+ years of experience in DevOps or SRE roles
- Strong experience with cloud computing and highly available distributed systems
- Demonstrated expertise with site reliability in critical environments (root cause analysis, production troubleshooting, on-call rotations)
- Experience working with reliability KPIs, observability, alerting, and SLAs
- Hands-on experience with CI/CD, containerization, and orchestration tools (Docker, Kubernetes)
- Knowledge of monitoring, logging, alerting, and observability platforms (Prometheus, Grafana, ELK Stack, Datadog)
- Familiarity with infrastructure-as-code tools (Terraform, CloudFormation)
- Proficiency in scripting languages (Python, Go, Bash) and software development best practices
- Strong understanding of networking, security, and system administration
- Excellent problem-solving and communication skills
- Ability to thrive in a fast-paced startup environment
Preferred Qualifications
- Experience in AI/ML environments
- Background with high-performance computing (HPC) systems and workload managers (Slurm)
- Familiarity with AI-oriented infrastructure solutions (Fluidstack, Coreweave, Vast)
Benefits
Mistral offers a comprehensive benefits package that varies by location and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal allowances, transportation benefits, and other location-specific perks. See the Benefits page for details specific to your location.
Mistral AI on Oh My Job
14 open positions right now, including 4 in California.