Oh My JobFind JobsResources
Company
  • About Us
  • Blog
  • Contact
  • Career Guides
Popular Roles
  • Registered Nurse Jobs
  • Home Health Aide Jobs
  • Software Developer Jobs
  • Project Manager Jobs
  • Part-Time Jobs
Tools
  • Paycheck Calculator
  • Job Market Data
Legal
  • Terms of Service
  • Privacy Policy
  • California Privacy Rights
For Employers
Post a Job
  1. Home
  2. Jobs
  3. Manager, Software Engineering (Resilience Engineering)
Illustration - Manager, Software Engineering (Resilience Engineering)

Manager, Software Engineering (Resilience Engineering)

Affirm
Affirm
Remote US
Jun 8, 2026
Salary not listed

At a glance

  • Remote

Job Description

```html

Affirm is seeking a Manager of Software Engineering to lead the Resilience Engineering team. This role focuses on building and maintaining systems that validate production reliability through load testing and chaos engineering practices. You will oversee a team developing platforms and tooling that allow engineers to safely test system behavior under stress and failure conditions in production environments, working remotely across the US.

Responsibilities

  • Define and advance the vision for resilience engineering at Affirm, positioning production load testing and chaos engineering as core engineering disciplines
  • Lead and mentor a team of engineers who build platforms and tooling for controlled production experimentation
  • Partner with infrastructure, product, and security leadership to integrate resilience validation into the software development lifecycle
  • Establish best practices for safely testing system limits and failure scenarios in production environments
  • Own the design and evolution of platforms enabling safe, controlled production load testing and fault injection
  • Implement safeguards including isolation boundaries, approval workflows, and automated rollback mechanisms to protect live systems and users
  • Build systems providing end-to-end observability, traceability, and auditability for resilience experiments
  • Drive reliability improvements by systematically identifying system weaknesses through load testing and chaos experiments
  • Establish monitoring, alerting, and incident response practices tailored to proactive resilience validation
  • Work with engineering teams to design and execute production load tests and chaos experiments safely
  • Partner with infrastructure teams to build guardrails around tests and experimentation activities
  • Enable teams to adopt resilience engineering practices and integrate them into their development workflows

Requirements

  • Demonstrated experience managing engineering teams focused on systems reliability, infrastructure, or platform development
  • Hands-on background building or operating production systems at scale
  • Experience with load testing, chaos engineering, or similar resilience validation methodologies
  • Knowledge of production monitoring, observability tools, and incident response practices
  • Experience designing safety mechanisms, control systems, and approval workflows for high-risk operations
```

Affirm on Oh My Job

309 open positions right now, including 5 in New York. Average salary across all roles: $78,769–$103,777.

Apply now
Share: