Oh My JobFind Jobs
Company
  • About Us
  • Blog
  • Contact
  • Career Guides
Popular Roles
  • CNA Jobs
  • Data Analyst Jobs
  • Electrician Jobs
  • FIFO Jobs
  • Property Management
Tools
  • Paycheck Calculator
  • Job Market Data
Legal
  • Terms of Service
  • Privacy Policy
  • California Privacy Rights
For Employers
Post a Job
Illustration - Research Engineer, Code Agents Infra

Research Engineer, Code Agents Infra

Mistral AI
Mistral AI
Palo Alto
Aug 4, 2026
Salary not listed
FullTime

Job Description

Mistral AI is hiring a Research Engineer for Code Agents Infra in Palo Alto. The role centers on building and operating execution, training, and data infrastructure for the company’s agentic models and coding assistants. Mistral AI develops full-stack AI solutions for enterprises in finance, manufacturing, defense, healthcare, and the public sector, with teams distributed across Europe, North America, Asia, and the Middle East.

Responsibilities

  • Design, deploy, and operate a high-throughput sandboxing platform that runs LLM-generated untrusted code across more than one million isolated environments concurrently for model evaluation and interactive RL environments.
  • Architect and scale high-throughput pipelines for synthetic code generation, agent trajectories, rollouts, and self-play data collection to support post-training and RL loops.
  • Optimize agent training codebases and distributed execution runtimes using PyTorch, Ray, and SLURM/Kubernetes to reduce multi-step rollout overhead, raise GPU utilization, and remove scaling bottlenecks.
  • Lower sandbox cold-start times to sub-second levels through container warm pools, snapshot/restore technologies such as CRIU and microVMs, and optimized image delivery across hybrid clusters.
  • Implement Kubernetes-native custom controllers, CRDs, and queuing systems to route short-lived evaluation, synthetic data, and agent execution tasks across diverse hardware fleets.
  • Maintain strict multi-tenant network and process isolation for untrusted agent code using container and sandboxing runtimes including gVisor and Firecracker, along with default-deny network configurations.
  • Ensure high availability, telemetry, and automated self-healing for millions of transient jobs and participate in on-call rotations for critical agent training and execution pipelines.

Requirements

  • Four or more years of experience in systems engineering, distributed systems, cloud infrastructure, or MLOps supporting LLM or RL workloads.
  • Proven experience building high-throughput data processing and generation pipelines for large-scale datasets using tools such as Ray, Spark, or custom distributed queues.
  • Strong expertise writing custom Kubernetes operators and controllers, managing Linux cgroups and namespaces, and optimizing Docker image layers and distribution systems.
  • Advanced proficiency in Python, Go, C++, or Rust with a demonstrated record of profiling and optimizing high-performance ML or backend systems codebases.
  • Hands-on experience with lightweight virtualization, container runtimes, or WASM technologies including Docker, gVisor, and Firecracker.
  • Deep familiarity with task queue systems, resource schedulers, and low-latency queuing architectures for high-volume, short-lived workloads.

Benefits

  • Comprehensive benefits package supporting well-being, growth, and work-life balance, which may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
  • Benefits vary by country; current details for each location are available on the company Benefits page.

Mistral AI on Oh My Job

15 open positions right now, including 4 in California.

Apply now
Share: