Accountabilities
- Own machine learning projects end-to-end, from problem exploration and experimentation through production deployment, monitoring, and ongoing iteration.
- Design and build reliable agentic LLM systems using multi-step workflows, tool calling, retrieval, and orchestration techniques suitable for high-stakes clinical environments.
- Develop robust evaluation strategies, including evaluation datasets, offline and online testing harnesses, LLM-as-judge pipelines with human review, and regression testing.
- Improve model performance using evidence-driven approaches such as prompt engineering, retrieval optimization, distillation, fine-tuning, or other appropriate techniques.
- Work across the full machine learning lifecycle, including data preparation, model adaptation, serving, monitoring, and production feedback loops.
- Partner with Product, Clinical, and Engineering stakeholders to translate clinical requirements into technical solutions and identify trade-offs early.
- Review code, share technical knowledge, contribute to engineering best practices, and mentor less experienced engineers.
- Continuously explore and apply AI-assisted approaches that improve individual productivity, team workflows, products, and engineering processes.
Requirements
- Proven experience shipping production machine learning systems that are relied upon by real users or customers.
- Hands-on experience developing and deploying LLM-based systems, including prompting, retrieval, tool calling, and agent-style workflows.
- Strong evaluation expertise, with experience building evaluation datasets, frameworks, testing methodologies, and mechanisms for distinguishing meaningful improvements from statistical or operational noise.
- Strong foundations in machine learning and the ability to select appropriate approaches based on technical requirements, evidence, and trade-offs.
- Experience working with ambiguous or loosely defined problems and transforming them into reliable production solutions.
- Strong software engineering skills, including writing production-quality code, working with distributed systems, and debugging complex machine learning pipelines.
- Clear communication skills and the ability to collaborate effectively with both technical and clinical stakeholders.
- Demonstrated AI fluency, with at least Level 1 proficiency: using AI regularly to improve personal productivity. More senior expectations may include building AI-enabled workflows or embedding AI into products and processes.
- Experience with fine-tuning or preference optimization techniques such as RLHF or DPO is a plus.
- Experience in healthcare AI or other high-stakes domains where system errors can have significant consequences is advantageous.
- Experience building agent frameworks or evaluation tooling from scratch is a plus.
- Open-source contributions, technical writing, or other forms of technical knowledge sharing are valued.
- Willingness to work within a distributed European environment, with the position open to candidates across Europe.
🇧🇷 Essa vaga exige inglês. Você está pronto?
A DevSpeak Academy prepara desenvolvedores brasileiros para conquistar vagas internacionais. Domine o inglês técnico com professores que entendem o mundo dev.
Conheça a DevSpeak AcademyCandidaturas encerradasVer outras vagas
