Accountabilities:
- Turn ambiguous infrastructure challenges into clear technical proposals and drive them through RFCs, architecture reviews, and cross-team alignment.
- Design and build self-service platform capabilities and APIs, primarily in Go, covering onboarding, provisioning, deployment, observability defaults, and day-2 operations.
- Establish reliable delivery standards using Terraform, GitOps with Argo CD, progressive delivery, automated testing, and continuous deployment practices.
- Evolve multi-tenant EKS infrastructure to improve reliability, security, scalability, and cost efficiency.
- Develop and improve ingress and traffic-routing capabilities, including Envoy Gateway and multi-region, cross-account connectivity.
- Strengthen SLOs, alerting, incident response, and operational follow-up through Grafana Cloud and improved observability practices.
- Measure success through outcomes for consuming engineering teams, including faster provisioning and deployment, greater self-service, and improved operational reliability.
- Develop AI-assisted operational workflows such as alert enrichment, incident context gathering, runbook-assisted diagnosis, remediation recommendations, and onboarding assistants.
- Maintain appropriate human oversight for AI-assisted operational actions, with an emphasis on safety, auditability, and responsible automation.
- Participate in the on-call rotation after onboarding and shadowing, while helping improve the overall health of on-call through better alerts, runbooks, automation, and blameless postmortems.
- Lead strategic platform initiatives from initial design through production adoption and establish durable technical patterns across engineering teams.
Requirements:
- 8+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Strong software engineering expertise in Go or a comparable programming language, including system design, testing, debugging, code review, and long-term maintainability.
- Proven experience designing, delivering, and operating cloud services or infrastructure platforms in production.
- Deep expertise in at least one area such as Kubernetes, networking, cloud platforms, reliability engineering, or developer platforms.
- Strong Linux, networking, and production operations fundamentals.
- Experience setting technical direction and leading initiatives that require alignment across multiple engineering teams.
- Strong written and verbal communication skills, particularly in remote environments and through RFCs, design documents, and incident writeups.
- Experience with EKS, ingress, CNI, or service-mesh technologies is valuable.
- Familiarity with OpenTelemetry, Prometheus, Grafana, CI/CD, progressive delivery, GitHub Actions, Argo CD, and canary deployments is a plus.
- Experience leading large-scale migrations, platform adoption programs, or cross-team infrastructure initiatives is beneficial.
- Strong systems judgment, curiosity, pragmatic decision-making, and the ability to develop deep expertise while navigating adjacent technical domains.
- Willingness to participate in an operational on-call rotation and contribute to improving its effectiveness.
- Visa sponsorship may be considered on a case-by-case basis depending on business needs.
Benefits:
- CA$238,250–CA$382,250 + equity for Canada-based candidates.
- Remote-first work arrangement.
- Flexible scheduling and autonomy in managing your working hours.
- Generous paid time off, quarterly wellness days, and an end-of-year wellness break.
- Home-office support to help create an effective remote workspace.
- Technology stipend equivalent to US$100 net per month.
- Annual learning and development stipend covering conferences, courses, certifications, and continued professional learning.
- 16 weeks of paid parental leave after six months of employment.
- Equity participation for full-time employees.
- Medical, retirement, and paid-holiday benefits, with details varying by country.
- Opportunities to lead major infrastructure initiatives involving self-service provisioning, multi-region networking, continuous deployment, Kubernetes, and platform engineering.
- Significant technical ownership within a small, growing infrastructure team.
- Exposure to AI-assisted and agentic operational workflows, with a focus on safe and auditable automation.
- Fully remote collaboration with distributed engineering teams.
- Offices available in Seattle and Paris for connection and collaboration.
🇧🇷 Essa vaga exige inglês. Você está pronto?
A DevSpeak Academy prepara desenvolvedores brasileiros para conquistar vagas internacionais. Domine o inglês técnico com professores que entendem o mundo dev.
Conheça a DevSpeak Academy