Summary
Nexthink is the leader in digital employee experience management software. The company provides IT leaders with unprecedented insight allowing them to see, diagnose and fix issues at scale impacting employees anywhere, with any application or network, before employees notice the issue. As the first solution to allow IT to progress from reactive problem solving to proactive optimization, Nexthink enables its more than 1,300 customers to provide better digital experiences to more than 18 million employees.
Responsibilities
- Implement and manage cloud-native systems (AWS) using best-in-class tools and automation.
- Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes to support rapid delivery cycles.
- Design, build, and maintain the infrastructure powering our multi-tenant SaaS platform with reliability, security, and scalability in mind.
- Define and maintain SLOs, SLAs, and error budgets, and proactively address availability and performance issues.
- Develop infrastructure-as-code (Terraform or similar) for repeatable and auditable provisioning.
- Build internal platform tools and automation to support provisioning, monitoring, and operational efficiency.
- Monitor infrastructure and applications ensuring high-quality user experiences.
- Participate in a shared on-call rotation, responding to incidents, troubleshooting outages, and driving timely resolution and communication.
- Act as an Incident Commander during the on-call duty and coordinate cross-team responses effectively to maintain an SLA.
- Drive and refine incident response processes, reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).
- Diagnose and resolve complex issues independently, minimizing the need for external escalation.
Apply for this position