Please Accept our Privacy Policy
Summary :
"We're looking for a Data Engineer to build and maintain the data infrastructure powering our AI/ML initiatives including data pipelines for model training, retrieval-augmented generation (RAG), and production ML systems. This role sits at the intersection of data engineering and applied AI, ensuring high-quality, well-governed data is available for both traditional analytics and AI consumption."
Key Responsibilities:
Design and build pipelines to ingest, clean, and transform structured and unstructured data (documents, PDFs, transcripts, logs) for AI/ML use cases
Build and maintain vector embedding pipelines and support vector database integration
Develop feature stores/pipelines for ML model training and inference
Implement data pipelines supporting RAG architectures chunking strategies, embedding refresh cycles, metadata tagging
Partner with ML engineers/data scientists to understand data requirements for model training and evaluation
Ensure data lineage, quality, and governance across pipelines feeding AI systems
Optimize pipelines for cost and performance (e.g., Dynamic Tables, streaming ingestion)
Support MLOps/LLMOps workflows versioning datasets, tracking data used in model training for reproducibility and compliance
Implement entitlement/access control logic so AI systems only access data they're authorized to use
Required Skills/Experience:
Strong SQL + one of Python/Scala for pipeline development
Experience with a cloud data platform
Experience with orchestration tools (Airflow, dbt)
Understanding of embedding models, vector search, and RAG patterns
Experience with unstructured data processing
Familiarity with ML lifecycle concepts
Nice to have:
Experience with Snowflake Cortex, Cortex Search, or similar in-warehouse AI tooling
Experience in a regulated industry (financial services) with data entitlement/compliance requirements
Exposure to LLM application frameworks (LangChain, LlamaIndex)