
Hey! We're Scale Up, and our client is looking for a Data Engineer – RAG & Generative AI to help build the data infrastructure powering next-generation AI applications.
Type of Employment: Contractor
Work Modality: 100% Remote
Work Schedule: Full-time
Location: LATAM
Start Date: August 3, 2026
End Date: December 31, 2026 (with the possibility of extension)
Our client is a fast-growing technology company building AI-powered products that help organizations unlock the full potential of their data. They foster a collaborative, engineering-driven culture where Data, AI, and Product teams work closely together to build scalable, production-ready solutions using modern cloud technologies.
We're looking for a Data Engineer to design and build the data pipelines that power Retrieval-Augmented Generation (RAG) and Generative AI applications.
In this role, you'll transform raw, unstructured data into AI-ready knowledge by developing scalable ingestion, processing, embedding, and vector indexing pipelines. You'll work closely with AI Engineers, Machine Learning teams, and Product stakeholders to ensure our AI systems have fast, reliable access to high-quality data that drives accurate and relevant responses.
Design and build scalable data pipelines for RAG and Generative AI applications.
Process and transform data from multiple formats, including PDFs, Word documents, HTML, JSON, and XML.
Implement document parsing, chunking, embedding generation, and vector indexing workflows.
Manage, optimize, and maintain vector databases for semantic search and retrieval.
Build data validation, metadata enrichment, and deduplication processes.
Monitor, troubleshoot, and optimize production data pipelines for reliability and performance.
Collaborate with AI Engineers and Product teams to continuously improve retrieval quality and system performance.
3+ years of experience in Data Engineering or building production data pipelines.
Strong proficiency in Python and SQL.
Experience working with cloud platforms such as AWS, Azure, or GCP.
Hands-on experience with data engineering tools like Spark, Airflow, Kafka, or similar technologies.
Experience with vector databases such as Pinecone, Weaviate, Milvus, Chroma, or Qdrant.
Familiarity with document processing libraries and RAG frameworks such as LangChain or LlamaIndex.
Experience with Docker and CI/CD pipelines.
Experience building Generative AI or Retrieval-Augmented Generation (RAG) solutions.
Knowledge of NLP concepts, embedding models, and semantic search.
Experience with Kubernetes and Infrastructure as Code (Terraform or CloudFormation).
Exposure to MLOps practices or enterprise-scale AI deployments.

We help start up founders to scale up their teams by taking care of the hassle and stress of recruiting and hiring remote team members.