AI QA & Evaluation Engineer

Elastic
Bangalore, India
On-site

Who this role is best for

Best suited to candidates with experience in AI/ML testing and evaluation, working in a collaborative, cross-functional environment focused on governance and ethical AI practices.

Best fit for

  • Candidates with experience in AI/ML testing and structured evaluation frameworks
    — “Design and implement comprehensive test strategies for AI/ML systems
  • Individuals who can navigate cross-functional collaboration across IT, engineering, and compliance teams
    — “Collaborate with teams from different areas. These areas include IT Engineering, IT Operations, Data & Integrations, PMO, CRM, Risk & Compliance, and business technology.
  • Professionals familiar with LLM evaluation tools and cloud-based AI infrastructure
    — “Direct experience with LLM evaluation frameworks and benchmarking tools such as LangSmith, Confident AI, etc.

Things to consider

  • The role requires maintaining high attention to detail, especially in detecting subtle inconsistencies in AI outputs
    — “High attention to detail and ability to notice subtle patterns or inconsistencies
  • Candidates must be prepared to work with a variety of cloud platforms and DevOps tools
    — “Experience with Cloud platforms (Azure, GCP, AWS)

How to stand out

  • Highlight experience with LLM evaluation and bias testing in your resume and interviews
    — “Design and implement comprehensive test strategies for AI/ML systems, including accuracy, bias, robustness, and regression testing.
  • Emphasize your ability to document nuanced observations and feedback in your application materials
    — “Advanced written communication skills, especially for documenting nuanced observations and feedback.
  • Showcase your understanding of AI ethics and data governance in your profile or interview responses
    — “Comprehension with AI ethics, risk management, and data governance.
  • Demonstrate hands-on experience with automation tools like GitHub and Terraform
    — “Thorough understanding of DevOps/automation/CI/CD tools: GitHub, Terraform.
  • Include specific examples of rubric-based evaluation or prompt engineering work in your application
    — “Design and implement self-contained evaluation tasks, including prompts, supporting files, and detailed grading rubrics.
Pace · Fast PacedCollaboration · HighAutonomy · MediumDecision Impact · Team

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • Designed and implemented comprehensive test strategies for AI/ML systems
  • Validated prompt engineering outputs
Typical background
Proficiency in Python, TypeScript, or other programming languages used in AI and test automation

Skills & requirements

Required

AI TestingML Systems TestingLLM Evaluation FrameworksAgentic/multi-agent SystemsCi/cd PipelinesData ValidationPrompt Engineering

Preferred

Llms: Azure Openai, Vertex AI, Chatgpt Enterprise

Stack & domain

PythonTypeScriptLLM Evaluation FrameworksLangsmithConfident AIRetrieval Augmented Generation (rag)Attention To DetailCollaborationTechnical RecommendationsAIQAEvaluation

About the role

Original posting from Elastic

Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale — unleashing the potential of businesses and people. The Elastic Search AI Platform, used by more than 50% of the Fortune 500, brings together the precision of search and the intelligence of AI to enable everyone to accelerate the results that matter. By taking advantage of all structured and unstructured data — securing and protecting private information more effectively — Elastic’s complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI.

What is The Role

We are looking for a skilled QA & Evaluation Engineer to join our team. The role blends strategic QA leadership with hands-on technical validation and structured evaluation to safeguard the accuracy, reliability, compliance, and ethical use of AI models. You will partner across IT and Engineering teams to identify, design, implement and run robust testing frameworks and evaluation rubrics for a portfolio of GenAI solutions that will be used across our organization.

What You Will Be Doing

Be a primary contributor to our AI strategy, helping validate and test AI infrastructure, custom solutions and third-party SaaS offerings.

Test Strategy & Execution: Design and implement comprehensive test strategies for AI/ML systems, including accuracy, bias, robustness, and regression testing.

Rubric-Based Evaluation: Design and implement self-contained evaluation tasks, including prompts, supporting files, and detailed grading rubrics to assess AI performance on functional workflows.

Automation & CI/CD: Automate validation suites for agentic/multi-agent systems, integration testing, and CI/CD pipelines for ML models.

Data Validation: Validate that AI/ML models are consuming accurate, authorized, and properly structured data sources; ensuring data quality across training and inference.

Observation & Reporting: Meticulously observe and document AI agent behaviors, producing crisp, precise summaries and reports on model performance and hallucinations.

Output Grounding: Validate prompt engineering outputs from a data accuracy standpoint, ensuring responses are grounded in verified data sources.

Refinement & Iteration: Iterate and refine evaluation tasks and rubrics based on feedback and team collaboration to ensure robust benchmarking methodologies.

Security & Governance: Ensure all AI data sources and structures meet governance, regulatory, and compliance standards, while implementing best practices for security and data privacy.

Collaborate with teams from different areas. These areas include IT Engineering, IT Operations, Data & Integrations, PMO, CRM, Risk & Compliance, and business technology.

Stay current on the latest work in AI and make technical recommendations to the organization.

What You Bring

Proficiency in Python, TypeScript, or other programming languages used in AI and test automation.

Proven skill in designing or applying rubric-based evaluation, grading against set criteria, or building structured scoring frameworks.

Direct experience with LLM evaluation frameworks and benchmarking tools such as LangSmith, Confident AI, etc.

Knowledge of the GenAI stack and solutions including Retrieval Augmented Generation (RAG).

LLMs: Azure OpenAI, Vertex AI, ChatGPT Enterprise or similar.

High attention to detail and ability to notice subtle patterns or inconsistencies (such as data hallucinations or logic errors) that others might miss.

Advanced written communication skills, especially for documenting nuanced observations and feedback.

Experience with Cloud platforms (Azure, GCP, AWS).

Thorough understanding of DevOps/automation/CI/CD tools: GitHub, Terraform.

Comprehension with AI ethics, risk management, and data governance.

Additional Information - We Take Care of Our People

As a distributed company, diversity drives our identity. Whether you’re looking to launch a new career or grow an existing one, Elastic is the type of company where you can balance great work with great life. Your age is only a number. It doesn’t matter if you’re just out of college or your children are; we need you for what you can do.

We strive to have parity of benefits across regions and while regulations differ from place to place, we believe taking care of our people is the right thing to do.

Competitive pay based on the work you do here and not your previous salary

Health coverage for you and your family in many locations

Ability to craft your calendar with flexible locations and schedules for many roles

Generous number of vacation days each year

Increase your impact - We match up to $2000 (or local currency equivalent) for financial donations and service

Up to 40 hours each year to use toward volunteer projects you love

Embracing parenthood with minimum of 16 weeks of parental leave

Different people approach problems differently. We need that. Elastic is an equal opportunity employer and is committed to creating an inclusive culture that celebrates different perspectives, experiences, and backgrounds. Qualified applicants will receive consideration for employment without regard to race, ethnicity, color, religion, sex, pregnancy, sexual orientation, gender perception or identity, national origin, age, marital status, protected veteran status, disability status, or any other basis protected by federal, state or local law, ordinance or regulation.

We welcome individuals with disabilities and strive to create an accessible and inclusive experience for all individuals. To request an accommodation during the application or the recruiting process, please email candidate_accessibility@elastic.co. We will reply to your request within 24 business hours of submission.

Applicants have rights under Federal Employment Laws, view posters linked below: Family and Medical Leave Act (FMLA) Poster; Pay Transparency Nondiscrimination Provision Poster; Employee Polygraph Protection Act (EPPA) Poster and Know Your Rights (Poster)

Elasticsearch develops and distributes technology and information that is subject to U.S. and other countries’ export controls and licensing requirements for individuals who are located in or are nationals of the following sanctioned countries and regions: Belarus, Cuba, Iran, North Korea, Syria, or Russia, including the Ukrainian territories annexed by Russia (The Crimea region of Ukraine, The Donetsk People's Republic (DNR), The Luhansk People's Republic (LNR), Kherson or Zaporizhzhia). If you are located in or are a national of one of the listed countries or regions, an export license may be required as a condition of your employment in this role. Please note that national origin and/or nationality do not affect eligibility for employment with Elastic.

Please see here for our Privacy Statement.

Source: Elastic careers

Similar roles