Software Engineer, ML Infrastructure Platform

Nuro
Mountain View, CA
On-site

Who this role is best for

Best suited to candidates with a strong background in distributed systems and machine learning infrastructure working in autonomous vehicle technology.

Best fit for

  • Candidates with a background in distributed systems and ML infrastructure who are committed to operational excellence
    — “We care as much about reliability and operational maturity as we do about raw scale.
  • Individuals with experience in Kubernetes and Python who can manage complex data pipelines
    — “Hands-on experience running production infrastructure on Kubernetes.
  • Candidates who have a track record of improving infrastructure reliability while reducing costs
    — “A track record of driving down infrastructure cost while improving reliability.

Things to consider

  • The role demands ownership of systems and their long-term operational stability
    — “A demonstrated ownership mindset: you drive systems to operational maturity.

How to stand out

  • Highlight experience with ML workflow automation and observability in your resume and interviews
    — “Design and develop agentic-first ML workflows - data-to-training-to-evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.
  • Showcase your ability to manage and optimize multi-cluster and GPU-based systems
    — “Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration.
  • Demonstrate your ability to define alerting and incident response practices for ML pipelines
    — “Own reliability for critical training and release pipelines: instrument them, define meaningful alerting, and build the on-call and incident-response practices.
Pace · Fast PacedCollaboration · HighAutonomy · MediumDecision Impact · TeamLevel · Mid

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • operational maturity
  • reliability
  • efficiency
Typical background
BS, MS, or PhD in Computer Science, Electrical Engineering, or a closely related field1+ years of relevant work experience

Skills & requirements

Required

Distributed SystemsKubernetesPythonProduction InfrastructureMonitoring

Preferred

GCPKubernetes-native OrchestrationGPU TrainingTraining Observability

Stack & domain

Machine LearningAutonomous DrivingDistributed GPU TrainingClosed-loop Reinforcement LearningWorkflowsOrchestrationObservabilityCost ManagementMulti-cluster SchedulingLarge-scale Data PipelinesBatch IngestionStreaming IngestionStorage LayoutHigh-throughput Data GenerationAgentic-first ML WorkflowsData-to-training-to-evaluation PipelinesMonitoringAlertingRunbooksKubernetesGCPKubernetes-native OrchestrationGPU / Distributed Training InternalsNCCLCollective CommunicationGPU And Training Observability ToolingOwnership MindsetOperational MaturityImplementation Deep-diveTechnical And Operational StandardsProficiency In PythonComfort With C++ Or GoSolid Distributed-systems FundamentalsPerformance ReasoningFailure Modes ReasoningReliability ReasoningComplex System ReasoningDistributed SystemsGPU Training

About the role

Original posting from Nuro

Who We Are 

Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides.

Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles.

With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected.

Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors

About the Role

Nuro takes a machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that possible. We own the systems that train the models at the core of the Nuro Driver™ - from distributed GPU training and closed-loop reinforcement learning, to the workflows, orchestration, observability, and cost management that keep the fleet running efficiently.

Our work sits directly on the critical path of autonomy development. When a training run stalls, when a pipeline silently regresses, or when GPU utilization slips, it shows up in how fast the rest of the company can ship. We care as much about reliability and operational maturity as we do about raw scale.

About the Work

Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration.

Design and operate large-scale data pipelines - batch and streaming ingestion, storage layout, and high-throughput data generation and storage.

Design and develop agentic-first ML workflows - data-to-training-to-evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.

Own reliability for critical training and release pipelines: instrument them, define meaningful alerting, and build the on-call and incident-response practices that let the team catch regressions.

About You

BS, MS, or PhD in Computer Science, Electrical Engineering, or a closely related field, plus 1+ years of relevant work experience.

Willingness to deep-dive into implementation and to raise the technical and operational standards of the broader engineering organization.

A demonstrated ownership mindset: you drive systems to operational maturity e.g. through monitoring, alerting, runbooks.

Strong proficiency in Python (and comfort with C++, Go or a similar systems language).

Hands-on experience running production infrastructure on Kubernetes.

Solid distributed-systems fundamentals and the ability to reason about performance, failure modes, and reliability across a complex system.

Bonus Points

Strong working knowledge of GCP.

Experience with building large scale data generation pipelines.

Experience with Kubernetes-native orchestration for ML workloads.

Depth in GPU / distributed training internals, including NCCL and collective communication.

Familiarity with GPU and training observability tooling and using it to diagnose real bottlenecks.

A track record of driving down infrastructure cost while improving reliability.

At Nuro, your base pay is one part of your total compensation package. For this position, the reasonably expected base pay range is between $160,360 and $240,540 for the level at which this job has been scoped. Your base pay will depend on several factors, including your experience, qualifications, education, location, and skills. In the event that you are considered for a different level, a higher or lower pay range would apply. This position is also eligible for an annual performance bonus, equity, and a competitive benefits package.

At Nuro, we celebrate differences and are committed to a diverse workplace that fosters inclusion and psychological safety for all employees. Nuro is proud to be an equal opportunity employer and expressly prohibits any form of workplace discrimination based on race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other legally protected characteristics.

Source: Nuro careers

Similar roles