Engineer, Storage and Data Protection

Thinkahead
Gurugram +1 more
On-site

Who this role is best for

Strong fit for mid-level storage engineers with HPC/AI infrastructure experience who excel in cross-team collaboration and automation. Requires hands-on expertise with distributed filesystems and on-call availability in Gurugram.

Best fit for

  • Candidates with 5+ years in HPC/AI storage and Linux sysadmin expertise
    — “5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering
  • Individuals who have optimized storage for AI training and inference workloads
    — “Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference
  • Professionals adept at troubleshooting storage, Linux, and network bottlenecks
    — “Troubleshoot storage, Linux, and I/O bottlenecks across storage clusters and fabrics

Things to consider

  • On-call rotation implies potential for unscheduled work hours
    — “Participate in on-call rotation

How to stand out

  • Highlight automation projects involving storage provisioning and lifecycle management
    — “Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations
  • Demonstrate experience with GPU clusters and AI/ML workflows in HPC environments
    — “Experience supporting storage solutions for GPU clusters and AI/ML workflows
  • Emphasize cross-functional problem-solving with infrastructure and research teams
    — “Work across technical teams to troubleshoot complex infrastructure issues
  • Showcase scripting skills in Python/Bash for operational efficiency
    — “Scripting or programming experience with Python and Bash
Pace · SteadyCollaboration · HighAutonomy · MediumDecision Impact · TeamLevel · Mid

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • Optimizing storage performance for AI training
  • Building automation for storage operations
Typical background
5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering

Skills & requirements

Required

HPC StorageDistributed FilesystemsStorage Performance TuningAutomationData ProtectionVendor ManagementCustomer Service

Preferred

Object StorageTerraformAnsibleObservability PlatformsScripting

Stack & domain

Linux Systems AdministrationDistributed FilesystemsStorage Performance TuningHPC SchedulersHigh-speed InterconnectsData Protection MechanismsMachine Learning WorkflowsCustomer ServiceProblem-solvingCommunicationScriptingProgrammingTeamworkHPCAI InfrastructureLarge-scale Storage Engineering

About the role

Original posting from Thinkahead via Lever

The High-Performance Computing Storage Engineer is primarily responsible for the overall health and maintenance of storage technologies in our managed services customer's environments. Our Storage Engineers are a valued member of the Managed Services Infrastructure Practice responsible for Tier 3 incident management, service request management and change management infrastructure support for all Managed Services customers.  

Key Responsibilities:

Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities 

Administer parallel and distributed filesystems such as Lustre, GPFS, BeeGFS, Ceph, Weka, or Vast 

Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference 

Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations 

Plan and perform maintenance activities 

Assess customer environments for performance and design issues and propose resolutions 

Work across technical teams to troubleshoot complex infrastructure issues 

Create and maintain detailed documentation 

Serve as a subject matter expert and escalation point for storage technologies 

Work with vendors to resolve storage issues 

Communicate with customers and internal team with transparency 

Support data movement workflows including ingest, replication, caching, tiering, and archiving 

Troubleshoot storage, Linux, network, and I/O bottlenecks across storage clusters and fabrics 

Partner with infrastructure, platform, and research teams to support production AI/HPC workloads 

Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency 

Communicate with customers and internal team with transparency 

Participate in on-call rotation 

Required Qualifications:

5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering 

Bachelor’s degree or equivalent Information Systems or related field. Unique education, specialized experience, skills, knowledge, training, or certification may be substituted for education 

Strong experience with Linux systems administration 

Hands-on experience configuring, managing, and tuning distributed or parallel filesystems 

Experience tuning storage for performance-sensitive workloads 

Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes 

Familiarity with high-speed interconnects such as InfiniBand or RDMA 

Ability to troubleshoot complex issues across storage, compute, and networking layers 

Understanding of data protection mechanisms, including data replication, backup strategies, and disaster recovery in HPC environments 

Experience with machine learning or data science workflows in HPC environments 

Managed Services or consulting experience 

Strong background with customer service 

High level problem-solving and communication skills 

Strong oral and written communications skills 

Managed Services or consulting experience 

Preferred Qualifications:

Experience supporting storage solutions for GPU clusters and AI/ML workflows 

Familiarity with object storage such as S3, MinIO, or Ceph Object Gateway 

Experience with Terraform, Ansible, Helm, or GitOps workflows 

Knowledge of observability platforms such as Prometheus and Grafana 

Experience with multi-petabyte environments, caching architectures, and storage isolation in multi-tenant systems 

Experience with machine learning or data science workflows in HPC environments 

Scripting or programming experience with Python and Bash 

Related Storage certifications are a bonus

Source: Thinkahead careers (Lever)

Similar roles