OT Systems Engineer (Supercomputer Infrastructure) - Memphis

Xai
Memphis +1 more
On-site

Who this role is best for

A natural match if you have experience with OT systems in high-reliability industrial environments and a background in integrating IT technologies with industrial control systems.

Best fit for

  • Candidates with OT systems experience in data centers or energy sectors and a track record of integrating IT technologies with industrial control systems
    — “OT environments in high-reliability settings
  • Individuals who have worked on mission-critical systems such as cooling plants, power generation, and BESS and have experience with BMS, SCADA, and PLC systems
    — “large cooling plants, medium-voltage distribution, and mission-critical facilities
  • Professionals comfortable with on-call rotations and working in plant or data hall environments, including at heights and with heavy lifting
    — “Ability to work at heights and in plant / data hall environments

Things to consider

  • On-call responsibilities may require working evenings, weekends, and after-hours to support 24/7 operations
    — “Willingness to participate in an after-hours on-call rotation
  • Travel is expected between Memphis, Southaven, and future campus locations as the company expands
    — “Willingness to travel (up to 20%) between Memphis, Southaven, and other sites
  • Physical demands such as lifting 30 lbs and working at heights are required for onsite tasks
    — “Ability to lift 30 lbs. Ability to work at heights

How to stand out

  • Highlight experience with industrial control networks and real-time systems in high-reliability environments
    — “Experience supporting real-time systems, industrial control networks
  • Emphasize your ability to write and maintain system documentation and operational procedures
    — “Write and maintain standards, architectures, best practices, and documentation
  • Demonstrate proficiency in automation frameworks like Terraform and Ansible for OT environments
    — “Proficiency in automation frameworks (Puppet, Terraform, Ansible, etc.)
  • Showcase experience with industrial protocols such as BACnet, Modbus, and Ethernet/IP
    — “Working knowledge of industrial protocols (BACnet, Modbus, OPC UA, MQTT, Ethernet/IP)
  • Demonstrate a history of balancing operational excellence with a sense of urgency in high-stakes environments
    — “thrives in high-stakes, 24/7 environments, brings a strong sense of urgency balanced with operational excellence
Pace · Fast PacedCollaboration · HighAutonomy · MediumDecision Impact · Team

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • Implement and support backend infrastructure for AI supercomputer campuses
  • Deploy and maintain development environments
  • Proactively monitor services and respond to incidents
  • Collaborate with cross-functional teams
Typical background
Experience in OT systems and controls software platformsBackground in virtualization and IT technologiesExperience in automation tools and infrastructure-as-code

Skills & requirements

Required

OT SystemsControls Software PlatformsVirtualizationVDIGitopsNetwork-level RedundancyEdge ComputeIT TechnologiesDevelopment EnvironmentsSystem UpgradesAutomation ToolsInfrastructure-as-codeDevOpsSimulation And Emulation EnvironmentsStandards And Documentation

Preferred

AI SystemsSupercomputer Infrastructure

Stack & domain

VirtualizationVDIGitopsNetwork-level RedundancyEdge ComputeBMSEPMSScadaPLCHMIAutomation ToolsInfrastructure-as-codeDevOpsICSOT EnvironmentsCommunicationInitiativeCuriosityUrgencyOperational ExcellencePrioritizationTeamworkSupercomputer InfrastructureControls And Industrial Software PlatformsAI Supercomputer CampusesPower GenerationCooling PlantsElectrical DistributionLiquid-cooling LoopsOn-site GenerationBESSData Hall Systems

About the role

Original posting from Xai via Greenhouse

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

SpaceXAI is looking for a highly skilled and versatile OT (Operational Technology) Systems Engineer to implement and support the backend infrastructure for next-generation controls and industrial software platforms critical to our hyperscale AI supercomputer campuses and co-located power generation. This role sits with the Supercomputer Physical Infrastructure team and supports dedicated controls, facilities, and power organizations while leveraging adjacent IT expertise, tooling, and technologies. You will balance sustainment of live cooling, power, and facility control systems with modernization and continuous improvement. The ideal candidate thrives in high-stakes, 24/7 environments, brings a strong sense of urgency balanced with operational excellence, and combines deep OT expertise with IT technologies such as virtualization, VDI, GitOps, network-level redundancy, and edge compute to create streamlined and secure ICS environments for ultra-dense AI compute.

RESPONSIBILITIES:

Design, deploy, and augment next-generation OT environments that support cooling plants, electrical distribution, liquid-cooling loops (CDUs / facility water), on-site generation, BESS, and data hall systems across the Memphis / Southaven campus and future sites.

Install, configure, maintain, and support industry-standard controls software platforms (BMS, EPMS, SCADA, PLC) as well as internally developed HMIs.

Integrate established and emergent IT technologies to simplify management, improve security, and create scalable, highly available plant and facility environments.

Deploy and maintain development, test, and staging environments to enable controlled, systematic change introduction on live critical systems.

Provide direct support during cluster bring-up, capacity expansions, commissioning, and production training campaigns.

Perform systems/software upgrades and maintenance between critical operations (including evenings and weekends as needed).

Proactively monitor services and respond rapidly to incidents to maintain high availability and performance of power, cooling, and environmental control.

Leverage automation tools and contribute to infrastructure-as-code and broader DevOps initiatives with a focus on ICS / OT environments.

Work with controls, mechanical, and electrical engineers to iterate on simulation and emulation environments to enhance test coverage and change control.

Write and maintain standards, architectures, best practices, and documentation (system overviews, design drawings, operational procedures), with emphasis on highly reliable and secure industrial environments, especially at IT/OT and data-hall boundaries.

Collaborate with cross-functional teams (IT, security, controls, facilities operations, construction, and power) as well as vendors and integrators to design robust OT architectures and resolve technical issues.

Ensure OT systems are configured and maintained in compliance with industry and cybersecurity standards (e.g., Purdue Model, IEC 62443).

BASIC QUALIFICATIONS:

3+ years of experience in OT systems engineering or industrial control systems administration.

Hands-on experience with multiple industry-standard controls software platforms and tools.

Significant experience designing, deploying, supporting, and troubleshooting OT environments in high-reliability settings.

PREFERRED SKILLS AND EXPERIENCE:

Experience supporting real-time systems, industrial control networks, or OT environments in data centers, power generation, semiconductor, energy, or similar high-reliability industries.

Direct experience with BMS, EPMS, SCADA, and PLC systems serving large cooling plants, medium-voltage distribution, and mission-critical facilities.

Experience designing architectures that incorporate hyperconverged, rugged industrial edge, and distributed compute technologies.

Working knowledge of industrial protocols (BACnet, Modbus, OPC UA, MQTT, Ethernet/IP, DNP3), controls networks, and OT cybersecurity best practices.

Proficiency in scripting (Bash / PowerShell / Python) and automation frameworks (Puppet, Terraform, Ansible, etc.).

Experience with configuration management, provisioning, infrastructure as code, and DevOps concepts/tools.

Familiarity with Active Directory, multi-platform authentication, and identity environments in OT contexts.

System administration experience managing Windows and Linux servers, rudimentary database administration, and storage/backup.

Network administration experience and understanding of the OSI model, especially Layer 1/2/3 considerations as they apply to industrial and facility networks (segmentation, VRFs, MDFs/IDFs).

Excellent communication skills with the ability to work with internal teams, vendors, and management in both formal and informal settings.

ADDITIONAL REQUIREMENTS:

Willingness to participate in an after-hours on-call rotation and work extended hours or weekends as necessary to support live campus operations.

Willingness to travel (up to 20%) between Memphis, Southaven, and other sites as the campus expands.

Ability to lift 30 lbs.

Ability to work at heights and in plant / data hall environments.

Ability to drive (active valid driver’s license).

Ability to work onsite in the Memphis, TN / Southaven, MS area.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Source: Xai careers (Greenhouse)

Similar roles