Skip to main content
Search Jobs

Search Jobs

Sr Specialist - Software Development & Engineering

Austin, Texas, United States Requisition ID 2026-126993 Category Engineering & Software Development Position Type Regular Pay range USD $44.86 - $68.85 / Hour Application Deadline 2026-09-21
Apply Now

Your Opportunity


Schwab remains committed to providing increased visibility to career growth opportunities and job requirements. This posting announcement is part of increased transparency and while all qualified applicants will be reviewed and considered, this organization has a preferred candidate identified for this role.

At Schwab, you’re empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us “challenge the status quo” and transform the finance industry together.

We believe in the importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).

Schwab Technology Services enables the future of how clients manage their money by providing innovative and reliable technology products and services as part of our ongoing commitment to democratize access to investing and financial planning.

Site Reliability Engineer responsible for reliability, scalability, and operational stability of enterprise middleware and API platforms across GCP/GKE/Kubernetes and PCF. Owns production readiness, observability, capacity and performance engineering, release governance, incident response, and continuous resilience improvements for Java/.NET microservices, API gateways, and supporting data/infrastructure services.

What you'll do:

  • Design and implement SRE practices for API gateways, service discovery, and microservices to improve uptime, latency, and recovery performance. 
  • Drive release planning and production readiness reviews including deployment, rollback, capacity, observability, and health check validation. 
  • Operate and support API gateway platforms (Apigee X, IBM DataPower) including ingress/egress routing, failover policies, and downstream resiliency. 
  • Implement and govern API controls: authN/authZ, TLS/HTTPS, rate limiting, quota, spike arrest, threat protection, routing policies, and error handling. 
  • Troubleshoot and stabilize Kafka ecosystem components (brokers, connect, schema registry), producer/consumer failures, lag, and message delays. 
  • Support PCF workloads: routes, service bindings, scaling, restaging, environment configuration, health checks, and blue-green deployments. 
  • Operate GCP foundations: GKE, VPC, firewalls, load balancers, logging, monitoring, secrets, compute, and BigQuery integrations. 
  • Plan and execute GKE platform maintenance: cluster/node pool upgrades, autoscaling tuning, certificate rotation, and workload migration. 
  • Manage CI/CD and IaC delivery using GitHub Actions, Helm/Kubernetes manifests, Terraform, and Git-based change workflows. 
  • Partner with security teams on PKI and TLS/mTLS for third-party integrations, gateways, load balancers, and data connections. 
  • Support database reliability and connectivity for MongoDB, Aerospike, and Oracle with focus on performance and resiliency. 
  • Define, track, and report SLIs/SLOs/error budgets, latency, failure rates, availability, and recovery objectives. 
  • Build and maintain observability dashboards and alerting in Grafana, Prometheus, Splunk, Moogsoft, Syswatch/Syshub, and Cloud Monitoring. 
  • Automate recurring operational tasks such as diagnostics, health checks, deployment validation, and JVM/.NET troubleshooting artifacts. 
  • Apply AI-driven anomaly detection for proactive identification of abnormal traffic, latency, error, and infrastructure usage patterns. 
  • Support compliance and resiliency programs: DR attestations, vulnerability remediation, certificate and service account attestations. 
  • Lead troubleshooting of critical incidents across APIs, Kubernetes, GCP, TLS/OAuth, networking, load balancing, Akamai, and deployments. 
  • Coordinate major incident management, root cause analysis, corrective actions, and blameless post-incident reviews. 
  • Produce high-quality incident documentation: impact, timeline, root cause, detection gaps, and preventive action tracking. 
  • Create and maintain runbooks, SOPs, operational guides, and knowledge articles for support and on-call excellence.

What you have


To ensure that we fulfill our promise of "challenging the status quo," this role has specific qualifications that successful candidates should have.

Required Qualifications:

  • Strong SRE/Production Engineering experience in distributed systems and enterprise middleware. 
  • Hands-on Kubernetes and GKE operations, including upgrades, autoscaling, resilience tuning, and workload troubleshooting. 
  • API gateway expertise with Apigee X and/or IBM DataPower plus ingress/egress routing and policy management. 
  • Solid understanding of API security and traffic governance: OAuth, TLS/mTLS, authN/authZ, throttling, quota, threat controls. 
  • Experience with PCF operations and lifecycle management. 
  • Kafka platform operations and troubleshooting (brokers, connect, schema registry, lag, throughput, reliability). 
  • Strong observability and alert engineering using Grafana/Prometheus/Splunk/Cloud Monitoring. 
  • CI/CD and IaC proficiency with GitHub Actions, Helm, Terraform, and GitOps-style workflows. 
  • Incident response leadership with RCA and postmortem discipline. 
  • Scripting and automation skills in Bash, PowerShell, or Python. 
  • Practical experience with GCP networking/security fundamentals (VPC, firewall, LB, secrets, logging/monitoring). 
  • Ability to define and operationalize SLOs, error budgets, and reliability KPIs

Preferred Qualifications: 

  • Experience supporting both Java and .NET runtime behaviors under production load. 
  • Knowledge of Akamai integration patterns and edge-to-origin troubleshooting. 
  • Familiarity with MongoDB, Aerospike, and Oracle reliability/performance operations. 
  • Exposure to AI/ML-based monitoring and anomaly detection tools and practices. 
  • Experience in regulated enterprise environments with compliance, attestations, and audit readiness. 
  • Strong documentation practices for runbooks, operational standards, and knowledge management. 
  • Cross-functional leadership skills working with platform, security, development, and operations teams

In addition to the salary range, this role is also eligible for bonus or incentive opportunities


What’s in it for you

At Schwab, you’re empowered to shape your future. We champion your growth through meaningful work, continuous learning, and a culture of trust and collaboration—so you can build the skills to make a lasting impact. Our Hybrid Work and Flexibility approach balances our ongoing commitment to workplace flexibility, serving our clients, and our strong belief in the value of being together in person on a regular basis.

We offer a competitive benefits package that takes care of the whole you – both today and in the future:

  • 401(k) with company match and Employee stock purchase plan
  • Paid time for vacation, volunteering, and 28-day sabbatical after every 5 years of service for eligible positions
  • Paid parental leave and family building benefits
  • Tuition reimbursement
  • Health, dental, and vision insurance
Apply Now