AI Inference Engineer
About this role
About the Role
Premier Global Links LLC is seeking an experienced Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up.
This role is focused on LLM inference, model serving, distributed systems, and inference runtime performance . The ideal candidate has hands-on experience with production inference systems and strong systems engineering skills, with Rust experience highly valued.
Key Responsibilities
• Build and optimize production LLM inference and model-serving systems . • Develop inference runtime components using Rust and other systems-level technologies. • Design and implement batching, scheduling, request routing, and serving infrastructure. • Build and optimize KV cache and prefix caching systems. • Scale inference workloads across multi-GPU and multi-node environments. • Profile, benchmark, and optimize latency, throughput, reliability, and cost. • Work with inference engines such as vLLM, SGLang, or TensorRT-LLM . • Investigate performance bottlenecks across the inference stack. • Contribute to core architecture and technical decisions for the platform. • Collaborate with a small, hands-on engineering team in a fast-paced environment.
Required Qualifications
• 2–10 years of experience in backend, distributed systems, or systems engineering. • Hands-on experience building, operating, or optimizing LLM inference or serving systems . • Deep understanding of transformer inference internals, including attention, KV cache, batching, and scheduling. • Experience with a production inference engine such as vLLM, SGLang, or TensorRT-LLM . • Strong programming experience with Rust, C++, Go, or systems-level Python/PyTorch . • Experience building performance-critical systems where latency, throughput, and cost are important. • Strong distributed systems and production software engineering fundamentals. • Ability to work on-site 5 days per week in Palo Alto, CA.
Preferred Qualifications
• Production Rust experience. • CUDA or Triton kernel development experience. • Multi-GPU or multi-node serving experience. • Experience with NCCL, NVLink, or RDMA . • Experience with prefix caching, speculative decoding, or prefill/decode disaggregation. • Contributions to open-source inference projects such as vLLM, SGLang, or Dynamo . • Experience working on inference systems at an AI provider, accelerator company, research lab, or similar organization. • Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related technical field.