ENGINEERING
AI Quality & Evaluation Engineer
Own the question every client asks: “how do we know the AI is good?” You will build evaluation harnesses, red-team suites and regression gates that let enterprises ship LLM systems with evidence instead of vibes.
- Remote
- Engineering
- Full-time
What you’ll do
- Design evaluation strategies per engagement: metrics, datasets, pass thresholds and review cadence
- Build automated eval suites that run in CI and block releases on behavioural regressions
- Red-team client systems for jailbreaks, injection and data-leakage paths before launch
- Stand up human-in-the-loop review workflows and calibrate them against automated judges
- Report model quality to client stakeholders in terms executives can act on
What we’re looking for
- 2+ years in QA, test automation or ML evaluation
- Experience building eval pipelines for LLM systems — golden sets, LLM-as-judge, human review loops
- Familiarity with eval and observability tooling such as LangSmith, Braintrust, promptfoo or OpenAI Evals
- Strong Python or TypeScript scripting for test harnesses and data wrangling
- Understanding of failure modes: hallucination, prompt injection, regression across model upgrades
What we offer
- Competitive salary
- Health, dental, and vision coverage
- Flexible remote work policy
- Professional development stipend
- Paid time off and holidays