Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. We're seeking a Senior Site Reliability Engineer to lead our SRE team in building and scaling platform reliability practices across our Engineering organization. This role combines technical depth in infrastructure and systems with leadership responsibility for defining and implementing reliability frameworks that protect customer experience.
About the Role
The SRE team at Affirm helps Engineering partners "Operate What They Own" with excellence by defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. As Senior SRE, you'll own quarterly goals for your team, lead engineers through ambiguous problems, and ensure everyone is supported throughout delivery. You'll participate in ideation with infrastructure, product management, developer experience, and analytics teams—articulating technical constraints and partnering on decisions that balance risks and trade-offs.
Responsibilities
- Own and deliver quarterly goals for your team while providing leadership and support through open-ended problem-solving
- Define and guide development of SLOs, service-level indicators, and reliability metrics across applications
- Drive the incident management, post-incident analysis, and change management processes organization-wide
- Proactively identify technical solutions and operational processes that strengthen incident readiness, response, and resilience
- Recommend and implement observability and alerting configurations to provide visibility on application performance
- Create and monitor metrics for team artifacts; escalate issues and support on-call operations
- Engage in service and architectural conversations to inform platform decisions
- Foster a culture of quality and ownership by setting code review and design standards, and advocating for them across teams
- Develop talent on your team through feedback, guidance, and mentorship by example
- Participate in on-call rotation as a requirement of the role
Requirements
- Software and systems engineering experience building and iterating on reliability, incident lifecycle, and resilience practices
- Expertise across infrastructure, platform, and distributed systems
- Experience with capacity management, load testing, and chaos testing
- Proficiency in automation, observability, and configuration management
- Development and product experience that informs operational decisions
- Demonstrated ability to lead teams and mentor engineers
- Comfort with on-call responsibilities and incident response
Benefits
- Remote position based in Spain