Senior Software Engineer - Platforms & Infrastructure
About this role
About the Role:
The Infrastructure team at Tubi builds and operates the core platforms that power our services at scale. We provide reliable, scalable, and developer-friendly systems for compute, networking, observability, and deployment.
As a Senior Software Engineer on the Infrastructure team, you will design, build, and operate the core platforms that all Tubi services run on: compute, networking, traffic, deployments, and observability. You will manage large-scale Kubernetes environments and use Infrastructure as Code (IaC) to keep cloud resources consistent and traceable. Your work can span the full platform stack, from the edge and service mesh to release tooling and reliability systems.
In this role, you’ll collaborate with cross-functional teams to deliver scalable, high-performing cloud solutions, working closely with application developers and gaining exposure to a wide range of technologies, including live-streaming, customer customization, and large-scale video transcoding pipelines.
This is a hybrid role (2 days/week) based out of our Toronto office. You must be willing to travel to our Toronto office two days/week.
What You'll Do:
• Kubernetes Operations: Manage and scale multi-cluster Kubernetes deployments, ensuring high availability, performance, and reliability. • Networking & Traffic: Design and operate the networking and traffic layer, from the edge to service-to-service communication. This includes safe rollout strategies (for example, canary releases, blue/green deployments, and gradual rollouts) with service mesh technologies such as Istio/Envoy. • Release Engineering & Developer Platform: Build and maintain CI/CD pipelines and self-service deployment tooling. Automate deployments and rollbacks, and improve the release experience for all engineers at Tubi. • Infrastructure as Code (IaC): Use Terraform and other IaC tools to provision and manage cloud infrastructure, ensuring consistency and auditability. • Observability & Incident Response: Build and operate the monitoring, logging, and tracing platform that all teams use. Troubleshoot and resolve production issues quickly to keep systems stable. • Documentation & Knowledge Sharing: Write and maintain clear technical documentation (system architecture, release processes, traffic policies, runbooks, best practices) to enable effective onboarding and collaboration. • Cross-Team Collaboration: Partner with developers, SREs, and platform teams to design scalable release and traffic strategies, and drive adoption of engineering best practices. Mentor junior engineers, including on effective AI-assisted workflows, and take part in interviewing and onboarding. • Technical Leadership: Lead major infrastructure initiatives end to end, from an ambiguous problem statement through design, rollout, and adoption. Help shape the team's technical roadmap and propose new investments.