DATA SCIENCE
Data Engineer, AI Pipelines
Build the data foundations every AI engagement stands on: ingestion from messy enterprise systems, document processing for retrieval, and pipelines that keep embeddings, features and evals fresh.
- Dallas
- Data Science
- Full-time
What you’ll do
- Design ingestion and document-processing pipelines that feed client RAG and fine-tuning workloads
- Keep embedding indexes, feature stores and eval datasets versioned and reproducible
- Build data-quality gates that stop bad inputs before they become bad model behaviour
- Optimise pipelines for the volume and freshness each engagement actually needs
- Work with client data teams to hand over pipelines they can run themselves
What we’re looking for
- 3+ years of data engineering with Python and SQL
- Experience with modern data stacks — dbt, Airflow or Dagster, Spark, and warehouses like Snowflake or BigQuery
- Hands-on with unstructured-data pipelines: parsing, chunking, embedding and indexing documents at scale
- Familiarity with vector databases (pgvector, Pinecone, Weaviate) and hybrid search
- Understanding of data quality and lineage practices in regulated environments
What we offer
- Competitive salary
- Comprehensive health benefits
- Professional development opportunities
- Flexible work arrangements
- Collaborative work environment