Staff ML Infrastructure Engineer

Date PostedApril 30, 2026LocationRemote / HybridCompanyGeneral MotorsSalary$189,300.00 to $290,700.00TypeSenior Level

Job Summary

Join the Embodied AI team at General Motors as a Staff ML Infrastructure Engineer to drive the development of core systems enabling rapid dataset generation, training, evaluation, and iteration of advanced Autonomous Driving models. You will deliver performant, easy-to-use, and reliable model training pipelines to dramatically accelerate the machine learning development cycle. Your success will be measured by the velocity and impact of the ML models relying on the scalable and high-performance training platforms you help create.

Responsibilities

  • Lead design and implementation of scalable ML training and evaluation platforms
  • Own complex technical projects end-to-end and drive technical prioritization
  • Collaborate with partner teams and mentor junior engineers and interns

Required Skills

  • 5+ years building large-scale distributed systems or advanced ML systems
  • Expertise in Python or C++
  • Experience with cloud infrastructure and MLOps practices
  • Strong understanding of machine learning algorithms and distributed training

Job Details

Role: Are you passionate about accelerating the future of autonomous driving? Join the Embodied AI team at General Motors. Our team is developing and deploying machine learning solutions that support safe and reliable autonomous vehicle behavior across real-world scenarios. As a Staff ML Infra Engineer, you will drive the development of core systems that enable rapid dataset generation, training, evaluation, and iteration of our most advanced Autonomous Driving models. From enabling large foundational driving models to distilling multi-stage production deployed models, your goal will be to dramatically accelerate the machine learning development cycle from one modeling hypothesis to next. You will deliver model training pipelines that are performant, easy to use, and exceptionally reliable. Your success will be measured by the velocity and impact of the ML models that rely on the scalable, intuitive, and high-performance training platforms you help create. What you'll do: - Lead the design, implementation, and deployment of scalable platforms and tools that drive machine learning model training and evaluation workflows across GM. - Own complex technical projects end-to-end, making key architectural decisions and technical trade-offs. You will be a core contributor to team planning, design reviews, and code quality. - Take a holistic view of projects, considering their impact across multiple teams, and across a longer timeline. Proactively drive technical prioritization. - Collaborate closely with partner teams to ensure maximum benefit from the systems we build. - Help shape our team through technical interviewing with high, well-calibrated standards, and play an essential role in recruiting. - Mentor and onboard junior engineers and interns, helping them grow their careers. What you'll bring: - 5+ years of experience building large-scale distributed systems, applications, or advanced ML systems-scale distributed systems, applications, or advanced ML systems - Proven track record of designing robust frameworks with high-quality, durable APIs. - Deep understanding of machine learning algorithms with hands-on application - Expertise in building reliable, high-performance, and cost-efficient systems on modern cloud infrastructure-performance - End-to-end experience across the ML development lifecycle, including MLOps practices - Strong cross functional collaboration skills across teams and organizations - Exceptional coding skills in Python or C++ - Strong interest in autonomous driving and its transformative potential - BS, MS, or PhD in Computer Science, Mathematics, or equivalent practical experience Nice to have: - Experience with distributed training methodologies - Experience scaling ML training across large GPU/CPU clusters or other accelerators - Familiarity with deep learning frameworks (e.g., PyTorch, TensorFlow) - Experience with performance profiling and state-of-the-art training optimization techniques, including their impact on model performance -of-the-art training optimization techniques, including their impact on convergence. - Experience with advanced build systems (e.g., Bazel, Buck, Blaze, CMake) - Proficiency with containerization and orchestration technologies (e.g., Docker, Kubernetes)