Skip to content
DATA SCIENCE

Computer Vision Engineer, Multimodal AI

Build vision systems for enterprises that now expect them to talk: defect detection that explains itself, document intelligence that reads like an analyst, and multimodal assistants grounded in what the camera sees.

  • San Francisco
  • Data Science
  • Full-time

What you’ll do

  • Build VLM-powered document, inspection and monitoring systems for client operations
  • Fine-tune and evaluate vision models against client-specific defect and document classes
  • Combine classic CV, VLMs and OCR into pipelines with the right cost profile
  • Deploy and monitor vision services in cloud and edge environments
  • Stay current with multimodal research and fold advances into delivery patterns

What we’re looking for

  • MS or PhD in Computer Science, Computer Vision, or equivalent experience
  • Strong PyTorch and experience with modern vision stacks (YOLO, SAM, DINOv2)
  • Hands-on with vision-language models (GPT-4o class, Claude, Gemini, LLaVA, Qwen-VL)
  • Experience deploying vision inference at scale — batching, GPU optimisation, edge constraints
  • Grounding in dataset curation and annotation quality for visual tasks

What we offer

  • Competitive compensation package
  • Comprehensive health benefits
  • GPU compute budget
  • Flexible work arrangements
  • Creative work environment
Back to job search