Sr Manager, Software Development & Engineering Lead
Your Opportunity
At Schwab, you’re empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us “challenge the status quo” and transform the finance industry together.
Schwab Technology Services enables the future of how clients manage their money by providing innovative and reliable technology products and services as part of our ongoing commitment to democratize access to investing and financial planning.
The SDE Site Reliability Engineering Lead is a highly experienced engineer responsible for improving the reliability, availability, scalability, and operational excellence of Schwab's digital platforms. This role applies Site Reliability Engineering (SRE) and DevOps principles to support mission-critical business transactions across web, mobile, APIs, and supporting technology services. The engineer partners closely with software development, platform engineering, infrastructure, security, and product teams to ensure systems are designed, deployed, monitored, and operated with reliability as a core requirement. This role leads efforts to automate operational processes, eliminate manual toil, improve observability, and accelerate the safe delivery of technology changes into production. The engineer provides technical leadership during production incidents, drives root cause analysis, and implements sustainable corrective actions to reduce operational risk. The role is expected to establish and mature reliability practices including monitoring, alerting, service level objectives (SLOs), resilience testing, operational readiness, and release governance. As a senior technical contributor, this position mentors engineers, influences engineering standards, and champions a culture of shared ownership, continuous improvement, and operational excellence. The engineer serves as a key partner in ensuring exceptional client experiences through resilient, secure, and highly available technology platforms.
What You Have:
- Design, implement, and support reliability engineering solutions for critical production services supporting Schwab web, mobile, and API platforms.
• Lead operational readiness efforts for new applications, platform enhancements, and major technology initiatives before production release.
• Build and maintain comprehensive monitoring, alerting, dashboards, and telemetry capabilities using enterprise observability tools.
• Define, implement, and evolve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability measurements for business-critical services.
• Automate operational processes and repetitive support activities to reduce manual effort, improve efficiency, and eliminate toil.
• Engineer self-service capabilities, tooling, and operational workflows that improve scalability and supportability.
• Design and enhance CI/CD pipelines to enable reliable, secure, and repeatable software deployments.
• Partner with development teams to embed reliability, resilience, observability, and operational standards throughout the software development lifecycle.
• Support production change management activities, including deployment planning, release validation, rollback strategies, and production verification.
• Participate in on-call rotations and provide technical leadership during service disruptions, major incidents, and critical production events.
• Lead incident triage, diagnosis, restoration, and coordination activities to minimize client impact and reduce recovery times.
• Facilitate blameless post-incident reviews and drive corrective actions that prevent recurrence and improve operational maturity.
• Perform trend analysis of incidents, alerts, capacity, availability, and performance data to proactively identify reliability risks.
• Develop automation, scripts, APIs, and tooling using modern engineering practices to improve platform operations and resiliency.
• Execute and support resilience, failover, disaster recovery, capacity, and performance testing exercises.
• Collaborate with infrastructure, cloud, security, networking, and platform engineering teams to strengthen overall system reliability and security posture.
• Establish and maintain operational standards, runbooks, dashboards, engineering patterns, and reliability best practices.
• Mentor engineers and technical teams on SRE principles, DevOps practices, observability strategies, incident management, and operational excellence.
What you have
- 8+ years of experience in software engineering, Site Reliability Engineering (SRE), DevOps, platform engineering, or related technical disciplines.
• 5+ years of experience supporting and operating large-scale, customer-facing production applications and services.
• Demonstrated experience leading reliability engineering initiatives across multiple applications, services, or technology domains.
• Deep experience supporting mission-critical systems with demanding availability, performance, scalability, and resiliency requirements.
• Experience designing and implementing enterprise monitoring, alerting, observability, and operational telemetry solutions.
• Experience leading major incident response activities, coordinating technical resolution efforts, and driving post-incident reviews.
• Strong experience developing automation, tooling, and self-service capabilities to reduce operational toil and improve operational efficiency.
• Experience designing, implementing, and supporting CI/CD pipelines and modern release management practices.
• Experience with distributed systems, microservices architectures, cloud-native applications, APIs, and highly available production environments.
• Demonstrated experience establishing operational standards, governance practices, runbooks, support models, or reliability frameworks.
• Experience defining and measuring service reliability through SLIs, SLOs, error budgets, and operational performance metrics.
• Strong software engineering and scripting skills using languages such as Python, Java, Go, JavaScript, Shell, or similar.
• Experience mentoring engineers and leading technical initiatives across teams.
• Strong problem-solving, analytical, communication, and stakeholder management skills.
• Experience participating in and leading on-call operations for critical production systems.
• Ability to influence technical direction and drive cross-functional collaboration across engineering, product, infrastructure, security, and operational teams.
Preferred Skills:
- Experience defining, implementing, and operationalizing enterprise-wide SRE practices and reliability standards.
• Experience establishing observability strategies across large application portfolios, including metrics, logs, traces, synthetic monitoring, and business transaction monitoring.
• Experience with enterprise observability and performance platforms such as Splunk, AppDynamics, Grafana, Prometheus, Dynatrace, OpenTelemetry, or similar technologies.
• Experience with public cloud platforms including AWS, Azure, or Google Cloud at scale.
• Experience with Kubernetes, OpenShift, service mesh technologies, or container orchestration platforms.
• Experience implementing Infrastructure as Code (IaC) and configuration automation technologies such as Terraform, Ansible, Helm, or similar tools.
• Experience designing or leading resiliency testing, capacity planning, chaos engineering, disaster recovery, or large-scale failover exercises.
• Experience supporting digital channels, mobile applications, API ecosystems, or high-volume e-commerce/customer-facing platforms.
• Experience leading operational readiness reviews for large-scale technology initiatives and production launches.
• Experience reducing Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR) through observability, automation, and operating model improvements.
• Experience leading cross-functional technical communities, reliability councils, or engineering governance forums.
• Industry certifications related to cloud engineering, Kubernetes, DevOps, observability, or Site Reliability Engineering.
• Experience working within highly regulated industries such as financial services.
• Experience supporting both web and mobile client platforms and the business transactions that underpin digital customer experiences.
Job Sub-Family Specific Competencies
- Fostering Innovation - Developing, sponsoring or supporting new and improved methods, products, procedures, or technologies
- Incident Response - Resolving reported incidents through streamlined processes, minimizing disruptions, and promptly restoring services
- System Design and Architecture - Implementing concepts for system design, ensuring compatibility with cloud architectures, and utilizing adaptive approaches for lifecycle models and methodologies
What’s in it for you
At Schwab, you’re empowered to shape your future. We champion your growth through meaningful work, continuous learning, and a culture of trust and collaboration—so you can build the skills to make a lasting impact. Our Hybrid Work and Flexibility approach balances our ongoing commitment to workplace flexibility, serving our clients, and our strong belief in the value of being together in person on a regular basis.
We offer a competitive benefits package that takes care of the whole you – both today and in the future:
- 401(k) with company match and Employee stock purchase plan
- Paid time for vacation, volunteering, and 28-day sabbatical after every 5 years of service for eligible positions
- Paid parental leave and family building benefits
- Tuition reimbursement
- Health, dental, and vision insurance