Manager, Software Development & Engineering
Your Opportunity
Schwab remains committed to providing increased visibility to career growth opportunities and job requirements. This posting announcement is part of increased transparency and while all qualified applicants will be reviewed and considered, this organization has a preferred candidate identified for this role.
At Schwab, you’re empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us “challenge the status quo” and transform the finance industry together.
Schwab Technology Services enables the future of how clients manage their money by providing innovative and reliable technology products and services as part of our ongoing commitment to democratize access to investing and financial planning.
We are seeking a talented Site Reliability Engineer who is a passionate problem solver, an automation aficionado, and a life-long learner willing to share knowledge and develop others. Our SRE team partners closely with product development teams, cloud infrastructure teams, and on-prem platform teams to achieve highly available and resilient applications.
You will be expected to balance solving problems, engineering solutions, and coding or scripting to benefit your team, your applications, and Schwab clients. Our development stack consists of Java, Angular, Python, MongoDB, Aurora MySQL, SQL Server, Kafka, and RabbitMQ. Our applications run both on-premises and in AWS.
What you have
•Practice Site Reliability Engineering principles to improve the availability, scalability, resiliency, recoverability, and performance of client-facing applications.
•Develop automation, scripts, tools, and repeatable patterns that reduce operational toil and improve reliability outcomes.
•Partner with application development, cloud infrastructure, platform engineering, security, and business teams to support reliable delivery of technology solutions.
•Support production operations, incident response, service recovery, root cause analysis, and post-incident improvement efforts for mission-critical applications.
•Build and maintain monitoring, alerting, dashboards, runbooks, and operational procedures that improve service health and accelerate issue resolution.
•Analyze production trends, identify reliability risks, and recommend engineering improvements to strengthen system performance and operational readiness.
•Contribute to disaster recovery planning, resiliency exercises, infrastructure modernization, vulnerability remediation, and security-focused operational initiatives.
•Participate in an on-call rotation and provide timely support for critical production systems.
•Mentor engineers, share knowledge, and promote a culture of accountability, learning, automation, and continuous improvement.
Required Qualifications
• 5+ years of experience supporting and administering enterprise technology platforms in large-scale environments.
•5+ years of experience with automation, scripting, monitoring solutions, alert management, and operational process improvement.
•5+ years of experience working within Software Development Lifecycle (SDLC) practices and supporting continuous improvement initiatives.
•Experience supporting high-availability distributed systems, production operations, and platform reliability initiatives.
•Experience leading incident response, root cause analysis, and service recovery efforts for mission-critical applications.
•Experience deploying, configuring, supporting, or migrating cloud-based applications and infrastructure.
•Development or scripting experience using one or more technologies such as PowerShell, Python, Java, .NET, or Bash.
•Experience working with database technologies such as Aurora, SQL Server, Oracle, MongoDB, or similar platforms.
•Experience supporting messaging and event-driven technologies such as Kafka, RabbitMQ, IBM MQ, or Solace.
•Experience using observability and monitoring platforms such as Splunk, AppDynamics, or equivalent tools.
•Ability to analyze complex technical issues, make sound operational decisions, and communicate recommendations effectively to technical and non-technical audiences.
•Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field
Preferred Qualifications
•5+ years of experience supporting large-scale, mission-critical platforms within financial services or other highly regulated industries.
•Experience implementing and scaling Site Reliability Engineering (SRE) practices, including Service Level Objectives (SLOs), post-incident reviews, observability, and reliability metrics.
•Strong background in production operations, availability engineering, and operational risk management.
•Experience leading infrastructure modernization initiatives, disaster recovery planning, vulnerability remediation, and security-focused operational programs.
•Experience partnering across engineering, infrastructure, security, vendor, and business teams to deliver complex technology solutions.
•Experience designing and implementing automation solutions that reduce operational overhead and improve reliability outcomes.
•Demonstrated ability to mentor engineers, establish operational standards, and promote a culture of accountability and continuous improvement.
•Experience with Amazon Web Services (AWS), Google Cloud Platform (GCP), Tanzu Application Service/Cloud Foundry (PCF), or similar cloud platforms.
Job Sub-FamilySpecific Competencies
- Analytical Thinking-Approaching a problem by using a logical, systematic, sequential approach
- Application Maintenance and Support-Delivering effective management and technical services to address technical issues and minimize disruption to application users
- Oral Communication-Expressing oneself clearly in conversations and interactions with others
- Incident Response–Resolving reported incidents through streamlined processes, minimizing disruptions, and promptly restoring services
- Fostering Innovation–Developing, sponsoring or supporting new and improved methods, products, procedures, or technologies
What’s in it for you
At Schwab, you’re empowered to shape your future. We champion your growth through meaningful work, continuous learning, and a culture of trust and collaboration—so you can build the skills to make a lasting impact. Our Hybrid Work and Flexibility approach balances our ongoing commitment to workplace flexibility, serving our clients, and our strong belief in the value of being together in person on a regular basis.
We offer a competitive benefits package that takes care of the whole you – both today and in the future:
- 401(k) with company match and Employee stock purchase plan
- Paid time for vacation, volunteering, and 28-day sabbatical after every 5 years of service for eligible positions
- Paid parental leave and family building benefits
- Tuition reimbursement
- Health, dental, and vision insurance