ENGINEERING
Site Reliability Engineer, Model Serving
Keep frontier-model systems up when they become business-critical. You will define SLOs for AI services, build the observability to defend them and run incident response when a provider, a prompt or a pipeline misbehaves.
- New York
- Engineering
- Full-time
What you’ll do
- Define and defend SLOs for latency, availability and answer quality on AI services
- Build multi-provider failover and degradation modes for frontier-model dependencies
- Automate capacity and quota management across model providers and GPU pools
- Lead incident response and blameless postmortems for AI system failures
- Advise client operations teams on runbooks for the AI systems we hand over
What we’re looking for
- 4+ years in SRE, DevOps or production engineering
- Experience operating LLM or ML serving infrastructure (vLLM, Triton, managed model endpoints)
- Strong observability skills — OpenTelemetry, Prometheus, Grafana — extended to tokens, quality and cost
- Incident management experience, ideally including third-party API dependency failures
- Solid programming ability in Python or Go for automation and tooling
What we offer
- Competitive salary
- Health, dental, and vision insurance
- Retirement savings plan
- On-call compensation and recovery time
- Flexible work arrangements