Sr Platform Engineer, ML Infrastructure
Skills
About this role
Accountabilities: • Design, build, and operate scalable ML infrastructure and platform capabilities supporting experimentation, training, deployment, and production operations.
• Develop developer tooling, services, automation, and infrastructure that help ML and engineering teams build and operate production systems more efficiently. • Lead complex technical initiatives independently, from problem definition and architecture through implementation, rollout, and operational ownership. • Make architectural decisions that balance immediate delivery needs with long-term scalability, reliability, maintainability, and developer experience. • Partner with ML engineers, infrastructure teams, and other stakeholders to understand needs and deliver effective platform solutions. • Identify and solve challenging infrastructure problems involving performance, reliability, scalability, and operational efficiency. • Drive adoption and continuous improvement by incorporating feedback from engineering teams using the platform. • Maintain high standards for software quality, production readiness, observability, and operational excellence. • Deliver platform capabilities that create measurable engineering and business impact across multiple teams and use cases.
Requirements:
• 5+ years of professional software engineering experience, particularly in platform engineering, infrastructure, or distributed systems. • Strong Python engineering skills, including experience developing production services, SDKs, automation, or platform tooling. • Proven experience designing, building, and operating production platforms used by multiple engineering teams. • Solid understanding of ML platform architecture and the end-to-end machine learning lifecycle, including experimentation, distributed training, model deployment, and production operations. • Experience building and operating applications on Kubernetes and cloud platforms, with AWS experience preferred. • Strong understanding of production reliability, observability, scalability, and operational best practices. • Strong technical judgment and the ability to independently drive complex initiatives from discovery through production while collaborating across technical teams. • Experience with developer platforms, internal tooling, or services that improve engineering productivity and reduce operational complexity is preferred. • Familiarity with workflow orchestration or distributed computing technologies such as Airflow, Kubeflow, Ray, Spark, or similar systems is a plus. • Experience designing or optimizing distributed, GPU-intensive compute platforms for ML training, inference, or large-scale image processing is preferred. • Experience supporting production ML platforms in computer vision, robotics, or related technical domains is advantageous. • Demonstrated technical leadership through architecture, mentoring, or influencing technical direction across teams.