ENGINEERING
Backend Engineer, LLM Platforms
Build the server-side machinery behind our clients’ AI products: orchestration services, retrieval layers, token-metered APIs and the queues and caches that keep frontier-model calls fast and affordable at enterprise scale.
- New York
- Engineering
- Full-time
What you’ll do
- Design and implement orchestration and retrieval services that ground model answers in client data
- Build token accounting, caching and semantic deduplication to keep inference costs predictable
- Implement guardrail layers: input validation, output schemas, moderation and audit logging
- Integrate client systems of record into tool-use and RAG pipelines securely
- Load-test and tune model-backed endpoints against latency and cost budgets
What we’re looking for
- 3+ years of backend experience in Node.js, Python or Go
- Experience operating LLM API integrations in production — rate limits, retries, streaming, structured outputs
- Solid database skills across PostgreSQL and a vector store (pgvector, Pinecone or similar)
- Understanding of caching, queuing and batching strategies for high-volume inference traffic
- Familiarity with multi-provider routing across Anthropic, OpenAI and open-weight models served via vLLM
What we offer
- Competitive salary package
- Comprehensive health coverage
- Retirement savings plan
- Frontier-model API budget for experimentation
- Flexible work hours