Helix AI Engineer, Training Performance

Figure AI
San Jose, CA
On-site

Who this role is best for

A natural match if you have experience optimizing large-scale AI training systems and working with distributed GPU clusters.

Best fit for

  • Candidates with mid-level experience in GPU kernel optimization and large-scale distributed training frameworks
    — “Optimize training performance for a 100B+ parameter models across 100k+ GPUs
  • Individuals who have co-designed models with researchers for hardware efficiency
    — “Partner with researchers to co-design model architectures and training recipes that are performant at scale
  • Professionals with a background in heterogeneous hardware evaluation and accelerator porting
    — “Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads

Things to consider

  • This role requires a strong commitment to on-site work in San Jose, CA
    — “Headquartered in San Jose, CA
  • The role demands a high level of independence and leadership in performance improvement projects
    — “3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects

How to stand out

  • Highlight experience with model parallelism strategies and large-scale training optimization
    — “Explore different model/data parallelisms (FSDP, context parallel, expert parallel, etc.)
  • Demonstrate your ability to build performance monitoring tooling and analyze root causes
    — “Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis
  • Showcase your contributions to open-source ML systems or custom kernel development
    — “Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)
  • Emphasize your skills in profiling and optimizing GPU performance using industry tools
    — “Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar)
  • Demonstrate your ability to evaluate and adapt to new hardware architectures
    — “Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware)
Pace · Fast PacedCollaboration · HighAutonomy · MediumDecision Impact · TeamLevel · Senior

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • optimize training performance for large models
  • collaborate on hardware procurement
  • build tooling for performance monitoring
Typical background
AI engineeringperformance optimizationdistributed systems

Skills & requirements

Required

AI Performance EngineeringGPU OptimizationDistributed TrainingCustom Kernel DevelopmentPerformance Monitoring

Preferred

Heterogeneous Training SetupsCross-cluster OrchestrationNon-nvidia Accelerators

Stack & domain

AI Performance EngineeringGPU ArchitectureProfiling ToolsCollective CommunicationModern Networking ConceptsPythonCuda/c++Framework InternalsDebugging Performance RegressionsHardware-efficiency MetricsHardware ProcurementCustom KernelsKernel CompilersAgentic SystemsEmerging Accelerator ArchitecturesModel/data ParallelismsCollaborationProblem-solvingTeamworkFast-paced EnvironmentMultiple Priorities

About the role

Original posting from Figure AI via Greenhouse

Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA.

Figure's vision is to deploy autonomous humanoids at a global scale. Our Helix team is looking for an experienced AI Training Performance Engineer to take our model training to the next level. This role is focused on improving distributed training frameworks for large scale model training, optimizing GPU kernels, exploring the relative gains of different accelerator types and co-designing our models to maximize utilization of our hardware.

Responsibilities

Optimize training performance for a 100B+ parameter models across 100k+ GPUs.

Collaborate with the broader team on accelerator choice, cluster topology, scheduling, and hardware procurement decisions to inform future scaling.

Write and optimize custom kernels (Triton/CUDA)

Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis across training jobs

Optimize data loading and preprocessing pipelines so I/O never gates the accelerators

Improve checkpointing, fault tolerance, and elastic restart so large jobs recover quickly from node failures without losing significant wall-clock time

Partner with researchers to co-design model architectures and training recipes that are performant at scale (e.g., activation checkpointing strategies, mixed precision, sequence packing)

Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators

Build and extend agentic systems that automatically generate, benchmark, and iterate on custom kernels

Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads, and lead proof-of-concept ports/benchmarks

Explore different model/data parallelisms (FSDP, context parallel, expert parallel, etc.) to determine optimal configuration per model size.

Requirements

Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field

3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects

Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy)

Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations

Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement).

Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals

Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence)

Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities

Bonus Qualifications

Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration

Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)

Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management.

The US base salary range for this full-time position is between $200,000 - $400,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Source: Figure AI careers (Greenhouse)

Similar roles