Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. We're looking for a Senior Site Reliability Engineer to join our remote Poland-based team and help our engineering organization operate systems with excellence, protect customer experience, and build the reliability practices that scale across Affirm.
About the Role
Site Reliability Engineering at Affirm is a focused team responsible for helping our Engineering partners "Operate What They Own" with excellence. The SRE function defines frameworks and best practices for operating applications, builds tooling, and provides training and consulting across the organization. You'll work across infrastructure, platform, and distributed systems while contributing to incident management, observability, and resilience practices that protect customer experience.
Responsibilities
- Own and deliver quarterly goals for your team, leading engineers through ambiguous, open-ended problems with proper support and guidance
- Provide visibility to teams and leadership on application performance through data-driven metrics and SLO development
- Drive the incident management and analysis process, and proactively strengthen incident readiness, response, and post-incident analysis
- Steer the implementation of change management and deployment practices across the engineering organization
- Participate in service and architectural conversations with infrastructure, product management, and developer experience teams
- Recommend and implement observability and alerting configurations that enable effective operations
- Support the operations and availability of your team's systems by creating metrics, monitoring performance, and contributing to on-call efforts
- Foster a culture of quality and ownership by establishing code review and design standards, and advocating for them through writing and technical talks
- Develop talent on your team through feedback, guidance, and leading by example
Requirements
- 4+ years of experience designing, developing, and launching systems
- Background in infrastructure, platform, and distributed systems
- Experience with capacity management, load testing, and chaos testing
- Proficiency in automation, observability, and configuration management
- Demonstrated experience in incident management, reliability engineering, or similar operational domains
- Strong collaboration and communication skills across engineering teams
Benefits
Remote position based in Poland.