Skip to content
ENGINEERING

Backend Engineer, LLM Platforms

Build the server-side machinery behind our clients’ AI products: orchestration services, retrieval layers, token-metered APIs and the queues and caches that keep frontier-model calls fast and affordable at enterprise scale.

  • New York
  • Engineering
  • Full-time

What you’ll do

  • Design and implement orchestration and retrieval services that ground model answers in client data
  • Build token accounting, caching and semantic deduplication to keep inference costs predictable
  • Implement guardrail layers: input validation, output schemas, moderation and audit logging
  • Integrate client systems of record into tool-use and RAG pipelines securely
  • Load-test and tune model-backed endpoints against latency and cost budgets

What we’re looking for

  • 3+ years of backend experience in Node.js, Python or Go
  • Experience operating LLM API integrations in production — rate limits, retries, streaming, structured outputs
  • Solid database skills across PostgreSQL and a vector store (pgvector, Pinecone or similar)
  • Understanding of caching, queuing and batching strategies for high-volume inference traffic
  • Familiarity with multi-provider routing across Anthropic, OpenAI and open-weight models served via vLLM

What we offer

  • Competitive salary package
  • Comprehensive health coverage
  • Retirement savings plan
  • Frontier-model API budget for experimentation
  • Flexible work hours
Back to job search