Skip to content
ENGINEERING

Site Reliability Engineer, Model Serving

Keep frontier-model systems up when they become business-critical. You will define SLOs for AI services, build the observability to defend them and run incident response when a provider, a prompt or a pipeline misbehaves.

  • New York
  • Engineering
  • Full-time

What you’ll do

  • Define and defend SLOs for latency, availability and answer quality on AI services
  • Build multi-provider failover and degradation modes for frontier-model dependencies
  • Automate capacity and quota management across model providers and GPU pools
  • Lead incident response and blameless postmortems for AI system failures
  • Advise client operations teams on runbooks for the AI systems we hand over

What we’re looking for

  • 4+ years in SRE, DevOps or production engineering
  • Experience operating LLM or ML serving infrastructure (vLLM, Triton, managed model endpoints)
  • Strong observability skills — OpenTelemetry, Prometheus, Grafana — extended to tokens, quality and cost
  • Incident management experience, ideally including third-party API dependency failures
  • Solid programming ability in Python or Go for automation and tooling

What we offer

  • Competitive salary
  • Health, dental, and vision insurance
  • Retirement savings plan
  • On-call compensation and recovery time
  • Flexible work arrangements
Back to job search