Site Reliability Engineer (Manufacturing Infrastructure)

SpaceX
Bastrop, TX
On-site

Who this role is best for

Strong fit for infrastructure-focused software engineers who operate in production environments and manage mission-critical systems, with a preference for candidates experienced in Linux and capable of on-site work across multiple U.S. locations.

Best fit for

  • Infrastructure-focused software engineers with production experience and a track record of system reliability
    — “compute, storage, and networking that run our factories must be as reliable as the products we build
  • Candidates who have translated high-level requirements into technical implementations and prioritized proactive maintenance
    — “Ability to translate high-level requirements into implementations from first principles
  • Individuals with experience in infrastructure as code and observability practices
    — “Manage infrastructure as code and use observability to provide a complete picture of platform health

Things to consider

  • This role requires on-site presence and does not allow for remote or hybrid work arrangements
    — “This role requires you to be onsite. Remote and/or hybrid work will not be considered
  • Candidates must be prepared for significant travel between multiple U.S. locations
    — “Must be able to travel to different sites (Hawthorne, CA; Redmond, WA; Cape Canaveral, FL; Starbase, TX)
  • This role may require passing a U.S. Air Force background check for certain locations
    — “Ability to pass Air Force background check for Cape Canaveral

How to stand out

  • Emphasize experience in proactive capacity planning and lifecycle management in your resume and interview responses
    — “Practice proactive maintenance: capacity planning, lifecycle management
  • Highlight your ability to communicate clearly with stakeholders and reduce toil in system operations
    — “Communicate clearly with stakeholders and teammates, and take ownership of work
  • Showcase projects where you designed for reliability and scalability in production environments
    — “Design for reliability, stability, and scale; find and remove bottlenecks with measurement and engineering
  • Demonstrate your experience with infrastructure as code using tools like Terraform or Ansible
    — “Infrastructure as code (Terraform, Ansible, Puppet, or similar)
  • Include examples of how you've improved system lifecycle from design to deployment and refinement
    — “Improve the full lifecycle—from design through deployment, operation, and continuous refinement
Pace · Fast PacedCollaboration · HighAutonomy · MediumDecision Impact · Team

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • Deploy, upgrade, operate, maintain, and scale compute, storage, and networking for manufacturing systems
Typical background
Bachelor’s degree in computer science, information systems, or an engineering discipline

Skills & requirements

Required

LinuxInfrastructure As CodeContainersVirtualizationDatabasesSoftware DevelopmentCommunication

Preferred

Aerospace Experience

Stack & domain

Linux Operating SystemsInfrastructure As Code (terraform, Ansible, Puppet, Or Similar)Containers And Virtualization (docker, Kubernetes, Vsphere, Qemu, KVM, Etc.)Databases And Data Modeling (postgres, Clickhouse, Etc.)CommunicationProblem-solvingTeamworkManufacturingAerospace

About the role

Original posting from SpaceX via Greenhouse

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.

SITE RELIABILITY ENGINEER (MANUFACTURING INFRASTRUCTURE)

The application software team is the central nervous system of SpaceX. Manufacturing is how SpaceX turns designs into hardware. The compute, storage, and networking that run our factories must be as reliable as the products we build. This team owns infrastructure supporting Starship, Starlink, Starshield, and Terafab. This position will have a direct impact on factory uptime, throughput, and production scale across programs.

The ideal candidate has strong software engineering fundamentals and a passion for infrastructure: reliability, stability, proactive maintenance, and scalability. You understand the system before you change it, solve hard problems, communicate clearly with stakeholders and teammates, and take ownership of work that manufacturing depends on.

Aerospace experience is not required. We value smart, motivated, collaborative engineers who treat teammates with fairness, respect, and support, and who want to take full ownership of challenging problems to help make humanity multi-planetary.

RESPONSIBILITIES:

Deploy, upgrade, operate, maintain, and scale compute, storage, and networking for manufacturing systems across Starship, Starlink, Starshield, and Terafab

Manage infrastructure as code and use observability to provide a complete picture of platform health

Design for reliability, stability, and scale; find and remove bottlenecks with measurement and engineering

Practice proactive maintenance: capacity planning, lifecycle management, and reducing toil before it becomes an incident

Partner with software engineers, manufacturing stakeholders, and site teams to build operable, maintainable systems

Improve the full lifecycle—from design through deployment, operation, and continuous refinement

Practice sustainable incident response and blameless postmortems

Provide high-quality support to manufacturing and engineering users

Communicate clearly with stakeholders and teammates

Participate in on-call and travel to sites as needed for deployments, incidents, and cross-site reliability

BASIC QUALIFICATIONS:

Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 3+ years of professional experience in SRE or DevOps in lieu of a degree

1+ years of software development experience 

Experience with Linux operating systems

PREFERRED SKILLS AND EXPERIENCE:

Experience with compute, storage, and/or networking infrastructure in production

Infrastructure as code (Terraform, Ansible, Puppet, or similar)

Containers and virtualization (Docker, Kubernetes, vSphere, QEMU, KVM, etc.)

Databases and data modeling (Postgres, Clickhouse, etc.)

Ability to translate high-level requirements into implementations from first principles

Comfort with mission-critical systems and appropriate urgency and care

Skillful communication with customers, peers, and management

Comfort operating across multiple sites and manufacturing programs

ADDITIONAL REQUIREMENTS:

Must be able to work extended hours and weekends as needed

Must be able to travel to different sites (Hawthorne, CA; Redmond, WA; Cape Canaveral, FL; Starbase, TX)

Ability to pass Air Force background check for Cape Canaveral 

This role requires you to be onsite. Remote and/or hybrid work will not be considered 

ITAR REQUIREMENTS:

To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.  

SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.

Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com. 

Source: SpaceX careers (Greenhouse)

Similar roles