Site Reliability Engineer

Xsolla
Russia
Hybrid

Who this role is best for

Candidates with infrastructure and software engineering experience in high-traffic commerce systems will find this role suits a technically proficient individual who thrives in fast-paced, collaborative environments.

Best fit for

  • Candidates with experience in building and maintaining production systems in cloud environments and a strong understanding of Kubernetes and observability tools.
    — “experience in operating production services in a cloud environment (GCP/GKE or comparable)
  • Individuals who can bridge development and infrastructure by contributing to design decisions with a reliability focus.
    — “the ability to hold a dual perspective - understanding both how developers ship features and what infrastructure needs to stay reliable
  • Candidates who have worked on high-traffic transactional systems in payments, fintech, or e-commerce domains.
    — “Experience in payments, fintech, e-commerce, or gaming — high-traffic transactional systems

Things to consider

  • The role requires active participation in on-call duty rotations and incident response activities.
    — “Participate in the SRE duty rotation, supporting developers across the company
  • Candidates must be comfortable with both software development and infrastructure management.
    — “equally comfortable writing code and running production systems

How to stand out

  • Highlight experience with application-level infrastructure ownership and reliability improvements from incidents.
    — “own the application-level infrastructure and reliability of a high-traffic commerce domain end to end
  • Showcase ability to define and enforce production-readiness standards for services.
    — “Run Production Readiness Reviews for new services and major changes; define and enforce what 'production-ready' means for the domain
  • Emphasize hands-on Kubernetes and observability tooling experience, especially with Datadog and OpenTelemetry.
    — “Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry
  • Demonstrate skills in building automation and tooling, preferably in Go, Python, or Bash.
    — “Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash)
  • Show collaboration with product development teams and involvement in architecture reviews.
    — “Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change
Pace · Fast PacedCollaboration · HighAutonomy · MediumDecision Impact · Team

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • owns the application-level infrastructure
  • designs and implements observability
  • sets up and evolves CI/CD pipelines
  • performs capacity planning
  • supports domain incident response
Typical background
pragmaticproduct-mindedcomfortable writing code and running production systems

Skills & requirements

Required

KubernetesObservabilitySoftware EngineeringCloud EnvironmentOperating Production ServicesCi/cd PipelinesCapacity PlanningPerformance TuningIncident ResponseAutomation

Preferred

Gitlab CIGithub ActionsDatadogOpentelemetrySlos/slisProduction Readiness Reviews

Stack & domain

KubernetesObservabilitySoftware EngineeringOperating Production ServicesCloud EnvironmentPartnering Closely With Product Development TeamsHolding A Dual PerspectiveUnderstanding Both How Developers Ship Features And What Infrastructure Needs To Stay ReliableDesigning And Implementing Slos/slisMonitorsAlertsDashboardsDatadogOpentelemetry-based ToolingSetting Up And Evolving Ci/cd PipelinesDeploy And Rollback AutomationCapacity PlanningPerformance TuningLoad TestingPerformance Regression InvestigationProduction Readiness ReviewsDefining And Enforcing What 'production-ready' MeansBuilding Domain-specific AutomationReducing Operational ToilMaintaining And Driving A Forward-looking Reliability RoadmapParticipating In Product Team PlanningRefinementsArchitecture ReviewsBringing The Reliability PerspectiveCommerceMonetizationVideo Game Industry

About the role

Original posting from Xsolla via Lever

ABOUT YOU 

We are looking for a Site Reliability Engineer who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our Infrastructure department's SRE team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to own the application-level infrastructure and reliability of a high-traffic commerce domain end to end - from deploy pipelines and Kubernetes manifests to SLOs, capacity planning, and production readiness.

Strong Kubernetes, observability, and software engineering skills are essential, along with experience in operating production services in a cloud environment (GCP/GKE or comparable) and partnering closely with product development teams. The ability to hold a dual perspective - understanding both how developers ship features and what infrastructure needs to stay reliable - and to bring the reliability lens into design decisions early will be key to your success in this role.

If you're passionate about making complex distributed systems boringly reliable and love building the commerce and monetization backbone that lets game developers around the world get paid, we would love to hear from you!

ABOUT US

Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.

For more information, visit xsolla.com.

Responsibilities:

Own the application-level infrastructure: Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service-level networking and integrations

Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards for critical services on Datadog and OpenTelemetry-based tooling

Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation

Perform capacity planning and performance tuning ahead of expected load - product launches, sales events, and regional rollouts - including load testing and performance regression investigation

Run Production Readiness Reviews for new services and major changes; define and enforce what "production-ready" means for the domain

Support domain incident response: assist with deep investigation of complex incidents, contribute to post-mortems, drive follow-up reliability improvements, and maintain runbooks

Build domain-specific automation that reduces operational toil: runbook automation, deploy helpers, recurring operational scripts

Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads

Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change

Co-author company-wide SLO/SLI, capacity, and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems

Participate in the SRE duty rotation, supporting developers across the company

Qualifications & Skills:

3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response duty, SLO/monitoring ownership, deploy pipeline and infrastructure work for production services

Software development background: you have built and shipped backend services, not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g., Go, PHP)

Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging application-level performance and networking issues (GKE or another managed Kubernetes)

Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry

  • Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams
  • GCP experience (IAM, networking, managed services)

Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions)

Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash)

Practical experience with incident response, post-mortems, and driving reliability improvements from incidents

Strong collaboration and communication skills — this role works embedded with product development teams daily

  • Experience in payments, fintech, e-commerce, or gaming — high-traffic transactional systems

Nice to Have:

Kubernetes certifications

Google Cloud Platform certifications

HashiCorp certifications

Benefits:

We are passionate about fostering a supportive environment for our team, so we prioritize the physical, mental, and emotional well-being of our employees and their families through a comprehensive Benefits Program. This includes 100% company-paid medical, dental, and vision plans, unlimited Flexible Time Off, and a personalized career roadmap for each employee. By investing in professional development through training and educational opportunities, we ensure that our team thrives both personally and professionally. Together, we’re not just building a business; we’re cultivating a community that values creativity, collaboration, and the transformative power of play.

By submitting the following job application form, you consent to Xsolla processing your data for career-related inquiries and potential employment opportunities. We process your data in accordance with this Xsolla Privacy Notice for Job Applicants. Please direct any inquiries regarding your data privacy to careers@xsolla.com.

Source: Xsolla careers (Lever)

Similar roles