Skip to content
ENGINEERING

AI Quality & Evaluation Engineer

Own the question every client asks: “how do we know the AI is good?” You will build evaluation harnesses, red-team suites and regression gates that let enterprises ship LLM systems with evidence instead of vibes.

  • Remote
  • Engineering
  • Full-time

What you’ll do

  • Design evaluation strategies per engagement: metrics, datasets, pass thresholds and review cadence
  • Build automated eval suites that run in CI and block releases on behavioural regressions
  • Red-team client systems for jailbreaks, injection and data-leakage paths before launch
  • Stand up human-in-the-loop review workflows and calibrate them against automated judges
  • Report model quality to client stakeholders in terms executives can act on

What we’re looking for

  • 2+ years in QA, test automation or ML evaluation
  • Experience building eval pipelines for LLM systems — golden sets, LLM-as-judge, human review loops
  • Familiarity with eval and observability tooling such as LangSmith, Braintrust, promptfoo or OpenAI Evals
  • Strong Python or TypeScript scripting for test harnesses and data wrangling
  • Understanding of failure modes: hallucination, prompt injection, regression across model upgrades

What we offer

  • Competitive salary
  • Health, dental, and vision coverage
  • Flexible remote work policy
  • Professional development stipend
  • Paid time off and holidays
Back to job search