Senior Software Engineer, Reliability - Traffic
Responsibilities
- You will design, build, and operate the software powering client's traffic routing and load-balancing platforms (using systems such as Apache Traffic Server, HAProxy, or Envoy).
- You will own the overall health, performance, and capacity of one or more of client’s internet-facing traffic services, keeping them monitored, automated, and ready to scale.
- You will triage and resolve the operational ticket flow across our intake queues and inbound escalations to diagnose, resolve, and route work.
- You will engineer the traffic tier for reliability: design for failure, isolate blast radius in the software, build in graceful degradation and automatic failover, and make SLOs a first-class part of how systems are built.
- You will partner with internal customers and deliver solutions. Be the support contact for product and platform teams that depend on the traffic tier; answer integration questions and follow through to resolution.
- You will build the monitoring, alerting, observability, and automation tooling that lets us deploy, observe, and operate the traffic tier safely at scale.
- You will drive operational excellence: capacity planning, configuration management, and safe, repeatable, automated deployments for rapid and reliable rollout of changes.
- You will participate in a 24x7 on-call rotation for the systems you own, providing emergency response to production incidents, mitigating quickly, and leading incident response when the traffic tier is impacted.
- You will drive blameless post-incident follow-through, turning each incident into a code or design fix so the same failure does not recur, and maintain the runbooks and documentation the team relies on.
- You will partner with development and customer teams from early design through production so platforms are built with scale and operability in mind.
Basic Qualifications
- BA/BS Degree or higher in Computer Science or related technical discipline, or equivalent practical experience
- 4+ years of industry experience in software design and development
- 4+ years of experience programming in one or more of C/C++, Rust, Go, Java, or similar languages
- Experience building or contributing to large-scale, high-availability production systems
- Strong troubleshooting skills that span application, operating system, and network (TCP/IP) layers
- Proficiency working in a Linux/UNIX environment
- Participate in a 24x7 on-call rotation for the services you own
- Working knowledge of networking fundamentals such as TCP/IP, HTTP, and DNS
Preferred Qualifications
- 5+ years of experience in software design and development for large-scale distributed systems
- Hands-on experience building or operating traffic infrastructure, such as proxies, load balancers, CDNs, edge, and routing platforms (for example, Apache Traffic Server (ATS), HAProxy, Envoy, IPVS, Azure Front Door (AFD), or Cloudflare)
- Depth in traffic engineering and networking: DNS, TCP/IP, UDP, HTTP/HTTPS, HTTP3, TLS, load balancing, Anycast, BGP, and global traffic routing
- A track record of improving system reliability: SLOs and error budgets, failure isolation, graceful degradation, capacity planning, or observability
- Experience building monitoring, alerting, and automation tooling for large-scale systems, including scripting in Python for automation
- Experience with deployment automation and configuration management for safe, repeatable rollouts
- Deep Linux/UNIX systems proficiency, including performance analysis and systems-level troubleshooting
- Experience safely operating and evolving critical, high-traffic, internet-facing production systems and participating in an on-call rotation / incident response.
- Strong written and verbal communication — you can run a calm incident, give clear status to customers and leadership, and document as you go.
- Hands-on proficiency with AI tooling for engineering work — actively use AI assistants and/or agentic tools day-to-day, write effective prompts, and know how to verify and correct their output rather than trusting it blindly.
Other keywords
Some more to take from:
Responsibilities
- Triage and resolve the operational ticket flow across our intake queues and inbound escalations to diagnose, resolve, and route work.
- Partner with internal customers and deliver solutions. Be the support contact for product and platform teams that depend on the traffic tier; answer integration questions and follow through to resolution.
- Accelerate triage and investigation through AI tooling to summarize alerts and incidents, surface likely causes, query logs and metrics. Speed up root-cause analysis using AI to correlate signals, propose hypotheses, and explain unfamiliar errors, while you steer the investigation and confirm findings against real evidence.
- Reduce toil through automation using AI tooling - tune noisy alerts, maintain operational documentation, scripts, queries, and regular changes to automate repetitive operational work.
Qualifications
- Strong Linux troubleshooting in production — confidently diagnose issues across processes, networking, disk, memory, CPU, and logs on live hosts.
- Solid networking fundamentals across L4 and L7 — TCP/UDP, DNS, connectivity, load balancing, proxies/reverse proxies, connection handling, timeouts, and retries.
- HTTP and TLS/SSL fundamentals — certificates, handshakes, SNI, and the common failure modes. (TLS/SSL issues are a recurring source of traffic incidents — this is core, not optional.)
- A root-cause mindset and the debugging skill to back it up — you can take an ambiguous production symptom and methodically work it down to cause under time pressure.
- Experience operating large-scale production systems and participating in an on-call rotation / incident response.