DATA SCIENCE
Computer Vision Engineer, Multimodal AI
Build vision systems for enterprises that now expect them to talk: defect detection that explains itself, document intelligence that reads like an analyst, and multimodal assistants grounded in what the camera sees.
- San Francisco
- Data Science
- Full-time
What you’ll do
- Build VLM-powered document, inspection and monitoring systems for client operations
- Fine-tune and evaluate vision models against client-specific defect and document classes
- Combine classic CV, VLMs and OCR into pipelines with the right cost profile
- Deploy and monitor vision services in cloud and edge environments
- Stay current with multimodal research and fold advances into delivery patterns
What we’re looking for
- MS or PhD in Computer Science, Computer Vision, or equivalent experience
- Strong PyTorch and experience with modern vision stacks (YOLO, SAM, DINOv2)
- Hands-on with vision-language models (GPT-4o class, Claude, Gemini, LLaVA, Qwen-VL)
- Experience deploying vision inference at scale — batching, GPU optimisation, edge constraints
- Grounding in dataset curation and annotation quality for visual tasks
What we offer
- Competitive compensation package
- Comprehensive health benefits
- GPU compute budget
- Flexible work arrangements
- Creative work environment