Accountabilities
- Design, build, deploy, and operate high-load distributed backend services and APIs that support machine learning infrastructure.
- Take end-to-end ownership of core ML services and associated data pipelines, from system design and implementation through deployment, observability, maintenance, and continuous improvement.
- Build reliable, scalable, and reusable infrastructure components that make machine learning workloads easier for product and ML teams to run and operate.
- Partner closely with ML and product engineers to understand their requirements and translate them into effective platform capabilities and services.
- Make and communicate technical decisions by evaluating architectural options, trade-offs, scalability, reliability, and operational requirements.
- Maintain a high standard for service reliability, performance, observability, and maintainability.
- Proactively identify technical and operational problems and take ownership of resolving them rather than allowing issues to remain unaddressed.
- Use modern AI-assisted development tools thoughtfully while maintaining strong ownership of system design, engineering decisions, and code quality.
- Contribute to a collaborative engineering culture by sharing knowledge, supporting teammates, and helping others solve technical challenges.
Requirements
- 5+ years of professional experience in backend engineering, platform engineering, or a closely related discipline.
- Extensive professional experience with Python, the primary programming language used in the environment.
- Hands-on experience developing and operating software on a public cloud platform such as AWS, GCP, or Azure, or working with self-managed Kubernetes; experience with AWS is particularly relevant.
- Strong experience designing and building distributed, high-load services and APIs.
- Solid understanding of data structures, algorithms, and the trade-offs involved in selecting and implementing them.
- Strong system-design mindset, with the ability to design solutions before implementation and understand the architectural implications of technical decisions.
- Experience working effectively with modern AI-assisted coding and development tools, combined with the judgment to understand their appropriate use and limitations.
- High level of ownership, initiative, and accountability, with a proactive approach to identifying and solving problems.
- Strong communication and collaboration skills, with a friendly and supportive approach to working with teammates and cross-functional partners.
- Experience with Rust, C, C++, or Go is a strong advantage.
- Previous experience developing or contributing to ML platforms or machine learning infrastructure is highly valued.
- Experience setting up and operating vector databases such as Qdrant, Milvus, Weaviate, OpenSearch, or pgvector is a plus.
- Experience with model serving or inference infrastructure, including LLM workloads, is advantageous.
- Experience with Infrastructure as Code tools such as Terraform is a plus.
Benefits
- Fully remote working environment, allowing you to choose where you live.
- Unlimited vacation time, with employees strongly encouraged to take at least three weeks of vacation each year.
- Home-office stipend to help you create a productive and comfortable remote workspace.
- Apple laptop provided for new employees.
- Annual training and professional development budget.
- Maternity and paternity leave for eligible employees.
- Competitive base salary of $80,000–$120,000 USD, depending on knowledge, skills, experience, and interview results.
- Stock options offered in addition to the base salary.
- Regular team offsites providing opportunities for collaboration and connection.
- Opportunity to work with experienced colleagues and contribute to meaningful, technically challenging projects.
🇧🇷 Essa vaga exige inglês. Você está pronto?
A DevSpeak Academy prepara desenvolvedores brasileiros para conquistar vagas internacionais. Domine o inglês técnico com professores que entendem o mundo dev.
Conheça a DevSpeak Academy