Partner Company logo
Partner Company

Site Reliability Engineer

🕐 1 dia atrás📍 Switzerland🌍 Remoto

Accountabilities

  • Define and implement the reliability strategy across the platform, including SLOs, SLIs, error budgets, incident practices, and reliability standards adopted by engineering teams.
  • Drive major architectural decisions as infrastructure evolves, evaluating technologies and designing systems that remain scalable, resilient, observable, and maintainable.
  • Design and own event-driven communication and messaging infrastructure, including the transition from synchronous patterns to durable asynchronous architectures.
  • Manage and evolve cloud infrastructure on AWS, using Infrastructure as Code to automate provisioning, configuration, deployment, and operational processes.
  • Ensure Kubernetes and containerized workloads scale reliably as transaction volumes and AI workloads increase.
  • Build and maintain comprehensive observability through monitoring, dashboards, alerting, application performance monitoring, and distributed tracing.
  • Serve as the senior escalation point for complex production incidents, leading incident response, root-cause investigations, and blameless postmortems.
  • Turn incident findings into permanent improvements through architectural changes, automation, operational controls, and resilience patterns.
  • Establish a continuous chaos engineering and resilience testing practice through fault injection, game days, and controlled failure experiments.
  • Mentor senior and mid-level engineers while raising the technical bar for reliability engineering and influencing engineering practices across teams.
  • Use AI-assisted tooling for automation, runbooks, incident analysis, and root-cause investigations, while helping establish effective AI-enabled engineering practices.
  • Within the first 6–12 months, establish the platform reliability strategy, lead at least one major architectural evolution, and drive adoption of the SLO and error-budget framework across engineering teams.

Requirements

  • Extensive experience in Site Reliability Engineering, Platform Engineering, DevOps, or a closely related discipline, with demonstrated ownership of production-scale systems.
  • Deep expertise in event-driven architecture and messaging systems such as Kafka, NATS, or RabbitMQ, including at-least-once delivery, consumer groups, dead-letter queues, backpressure, and migrations from synchronous to asynchronous architectures.
  • Strong AWS expertise across services such as EC2, VPC, IAM, S3, and RDS, combined with solid networking fundamentals.
  • Hands-on Infrastructure as Code experience using Terraform, Pulumi, or similar tools, with infrastructure managed through version-controlled workflows and code reviews.
  • Strong production experience with Kubernetes and Docker, including container lifecycle management, resource limits, health checks, and orchestration at scale.
  • Proven observability expertise using Datadog or equivalent platforms, including dashboards, monitoring, APM, distributed tracing, and alerting.
  • Demonstrated experience defining and operating SLOs, SLIs, and error budgets across multiple services.
  • Hands-on experience with chaos engineering, fault injection, game days, or resilience experiments using tools such as Gremlin, Chaos Mesh, AWS FIS, or similar technologies.
  • Strong distributed systems debugging skills, with experience diagnosing asynchronous workflows, cascading failures, and complex production incidents.
  • Ability to code for automation and engineering tooling using Go, Python, or a similar programming language.
  • Solid database knowledge across SQL and NoSQL technologies, particularly PostgreSQL, MongoDB, and Redis, including indexing, replication, and performance optimization.
  • Proven technical leadership experience, including setting reliability standards, influencing architecture across teams, and mentoring engineers.
  • Advanced written and spoken English communication skills.
  • Experience with AI or MLOps infrastructure, including model serving, LLM inference, GPU/resource management, or AI agent observability, is highly advantageous.
  • Familiarity with multi-tenant container platforms and customer workload infrastructure is a plus.
  • Experience with data pipelines and orchestration tools such as Airflow or Prefect, and data platforms such as Databricks, Snowflake, or BigQuery, is beneficial.
  • Familiarity with incident management platforms such as PagerDuty, Opsgenie, or incident.io is an advantage.
  • Experience in the payments industry is preferred.
  • Additional experience with ECS, s6-overlay, AI agent frameworks, or Spanish proficiency is a plus.

Benefits

  • Competitive compensation.
  • Fully remote working environment with the flexibility to work from different locations.
  • One-time home office allowance to help create an effective workspace.
  • Company-provided work equipment.
  • Stock options.
  • Health plan available wherever you are.
  • Flexible days off.
  • Access to language, professional, and personal development courses.
  • Opportunity to work on globally scaled infrastructure supporting complex payment and AI workloads.
  • Significant technical ownership and influence over reliability strategy, architecture, and engineering standards.
  • Collaborative international environment with opportunities to mentor engineers and shape organization-wide engineering practices.

🇧🇷 Essa vaga exige inglês. Você está pronto?

A DevSpeak Academy prepara desenvolvedores brasileiros para conquistar vagas internacionais. Domine o inglês técnico com professores que entendem o mundo dev.

Conheça a DevSpeak Academy