AI Evaluation Engineer — Evals & Systems Verification
AI Evaluation Engineer — Evals & Systems Verification
Location: HSR, Bengaluru (On-site)
Experience: 3–6 years
About the Role
Building with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.
Responsibilities
Build and maintain evaluation harnesses for AI-facing features to measure and tune system quality (e.g., capture quality, retrieval quality, guidance quality)
Own end-to-end and API test infrastructure (Playwright-class), supporting a continuous, daily-release cycle
Design and execute host-behavior probes using scripted user sessions across diverse AI assistants, ensuring product behavior aligns with expectations
Gate production releases through thorough user acceptance testing (UAT), regression analysis, and quality reporting
Requirements
Background in SDET/QA automation or ML evaluation, with proven ownership of test or evaluation infrastructure—not just executing tests
Strong programming skills in Python or TypeScript, with experience in API-level testing and tools like Playwright or Cypress
Familiarity with LLM applications or a demonstrated interest in evaluating non-deterministic systems
Highly autonomous and able to define your own workflows and processes for evaluation and verification