Oh My JobFind JobsResources
Company
  • About Us
  • Blog
  • Contact
  • Career Guides
Popular Roles
  • Registered Nurse Jobs
  • Home Health Aide Jobs
  • Software Developer Jobs
  • Project Manager Jobs
  • Part-Time Jobs
Tools
  • Paycheck Calculator
  • Job Market Data
Legal
  • Terms of Service
  • Privacy Policy
  • California Privacy Rights
For Employers
Post a Job
  1. Home
  2. Jobs
  3. California
  4. Sr. Staff Production Engineer - Data Platform
Illustration - Sr. Staff Production Engineer - Data Platform

Sr. Staff Production Engineer - Data Platform

Databricks
Databricks
Mountain View, California; San Francisco, California
Apr 9, 2026
Salary not listed

At a glance

  • 10+ years of experience

Job Description

RDQ126R106

At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity to tackle technical challenges, from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And we're only getting started.

Our Data Platform organization is the backbone of the world’s leading Data and AI platform. We operate a massive, multi-cloud (AWS, Azure, GCP), multi-AI, multi-region stack that powers thousands of the world’s most demanding workloads. As we move into the era of Agentic AI, we are not just scaling infrastructure; we are building a new generation of Agentic Observability and Reliability systems.

As a Sr. Staff Production Engineer, you will lead the strategic vision for the operational stability of our internal "Databricks-on-Databricks" environment. You will transition our infrastructure from traditional SRE models toward an agent-driven, self-healing architecture, ensuring that our platform—and the agents operating within it—remain rock-solid for mission-critical customer workloads.

The Impact You Will Have

  • Architecting Agentic Reliability: Define and drive the design of future "self-healing" infrastructure at scale where AI agents proactively detect, diagnose, and remediate production incidents before they impact customers.
  • Data Platform Optimization: Own the operational integrity of the Data Platform that powers our internal AI models, ensuring 99.99% availability for the compute, storage, and control plane services used by thousands of Databricks engineers.
  • High-Scale Operational Excellence: Establish the next generation of "Change Safety" protocols, utilizing automation and agentic guardrails to manage complex deployments across 100+ global regions.
  • Leadership in Chaos & Scale: Serve as a technical bar-raiser for the team, evangelizing modern SRE practices (including Chaos Engineering) to navigate the structural transformation of the industry toward agentic, autonomous systems.

What We Look For

  • BS/MS/PhD in Computer Science, or a related field
  • Technical Depth: 10+ years of production-level experience as a Software Engineer or SRE in highly distributed, multi-cloud environments.
  • Engineering Persona: You write code to solve operational problems. You are

Databricks on Oh My Job

1,178 open positions right now, including 193 in California. Average salary across all roles: $192–$1.

Apply now
Share: