Kong Inc. logo

Site Reliability Engineer 2

๐Ÿ’ฐ $123,000 โ€“ $150,000๐ŸŒ ๐Ÿ‡บ๐Ÿ‡ธ United StatesFull-time๐ŸŽฏ LeadPosted 3 days ago

Are you ready to unlock intelligence?

If you donโ€™t think you meet all of the criteria below but are still interested in the job, please apply. Nobody checks every box - weโ€™re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.

Are you ready to unlock intelligence?

If you donโ€™t think you meet all of the criteria below but are still interested in the job, please apply. Nobody checks every box - weโ€™re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.

About the Role:

As a Site Reliability Engineer, youโ€™ll join the global Platform SRE team responsible for building, operating, and scaling Kongโ€™s multi-region SaaS platform that powers the worldโ€™s API connectivity.

Youโ€™ll design, automate, and run production systems serving thousands of customers across AWS, GCP, and Azure. Youโ€™ll work on everything from multi-region Kubernetes clusters to service mesh and gateway architectures, ensuring the reliability, scalability, and security of Kongโ€™s SaaS offerings.

This is a hands-on role ideal for engineers who thrive on running production SaaS systems at scale, automating operations, and continuously improving performance, resilience, and deployment pipelines.

What Youโ€™ll Do:

  • Operate and scale Kongโ€™s global SaaS platform (Konnect), ensuring reliability, availability, and performance across regions and clouds.

  • Build, automate, and maintain Kubernetes-based infrastructure and deployment workflows using Terraform/Terragrunt, Helm, and ArgoCD.

  • Design, maintain, and optimize multi-region data and caching layers โ€” including PostgreSQL, Redis, ClickHouse, and Druid โ€” for high availability and low latency.

  • Operate and improve Kong Gateway and Kong Mesh environments supporting hybrid and distributed architectures.

  • Develop and maintain CI/CD pipelines and GitOps workflows to automate service delivery and ensure consistent infrastructure changes.

  • Enhance observability and incident response readiness through systems like Datadog, Prometheus, Grafana, and Thanos, defining and tracking SLOs.

  • Collaborate closely with development and security teams to ensure smooth operation of SaaS services in compliance with reliability, security, and regulatory standards.

  • Participate in a global 24/7 on-call rotation and drive continuous improvement of operational playbooks and postmortem practices.

  • Lead and contribute to scaling initiatives that improve elasticity, reliability, and cost-efficiency across the SaaS platform.

What Youโ€™ll Bring:

  • BS in Computer Science or equivalent practical experience.

  • Proven experience managing SaaS or PaaS systems at enterprise scale (multi-region, multi-tenant, secure environments).

  • Deep expertise in Kubernetes, including debugging cluster/networking issues and designing for fault tolerance and scalability.

  • Strong proficiency with Infrastructure as Code tools like Terraform or Terragrunt.

  • Experience with CI/CD pipelines and GitOps workflows (ArgoCD, Atlantis, Helm).

  • Proficiency in one or more programming languages (Go, Python, Bash) for automation and tooling.

  • Solid understanding of Linux/Unix systems, networking (DNS, TLS/SSL, HTTP), load balancers and distributed systems.

  • Experiencing working with API gateway and service mesh technologies

  • Familiarity with streaming systems like Kafka and observability platforms (Datadog, Prometheus, Grafana).

  • Experience working in a 24/7/365 production support environment.

Bonus Points:

  • Hands-on experience with Kong Gateway, Kong Mesh, or similar service connectivity technologies.

  • Experience operating ClickHouse, Druid, or other time-series and analytics databases.

  • Experience managing PostgreSQL and Redis in multi-region configurations.

  • Working knowledge of AWS networking (PrivateLink, Transit Gateway, VPC Peering, Firewalls), Azure VNet, or GCP NCC.

  • Strong understanding of disaster recovery, resiliency testing, and compliance-driven reliability practices.

#LI-KC1

About Kong:

Kong Inc., the AI Connectivity Company, is building the connectivity layer of AI. Trusted by the Fortune 500ยฎ and AI-native startups alike, Kongโ€™s unified API and AI platform enables organizations to secure, manage, accelerate, govern, and monetize the flow of intelligence across APIs and AI traffic โ€” on any model, any cloud. For more information, visit www.konghq.com.

Explore more remote jobs