Senior RAG - Document AI Engineer
About this role
Senior RAG - Document AI Engineer Design, build and operate the Python/FastAPI microservices that turn heterogeneous enterprise documents into accurate, cited answers — from document parsing and chunking through enrichment, hybrid retrieval and evaluation. You own retrieval quality and the evidence for it. Responsibilities ● Build document processing services: format-aware parsing across native text, OCR and vision-language models; table and figure extraction. ● Build a configurable chunking service with strategies selectable per document class, maintained as versioned pipeline profiles. ● Implement LLM-based metadata extraction into strict JSON schemas, with confidence scoring and human-in-the-loop review. ● Build the retrieval service: query rewriting, hybrid dense + BM25 with rank fusion, metadata pre-filtering, cross-encoder reranking, context assembly with citations. ● Own the evaluation harness — golden Q&A sets, recall@k, nDCG/MRR, groundedness — and gate releases on measured quality. ● Instrument, monitor and support the services in production. Qualifications ● 6–10 years software engineering, with 2+ years building retrieval or document-AI systems used by real users in production. ● Production RAG at scale. 100k+ documents and millions of chunks; p95 query latency under 2s, sustained under concurrent load. ● Operational maturity. Incremental and delta ingest; has re-indexed a live corpus without downtime after a chunking or embedding-model change. ● Chunking as a measured design decision. Hierarchical parent–child and section- aware strategies — not a single global token size. ● Retrieval depth. Hybrid dense + sparse (BM25) retrieval and rank fusion; cross-encoder reranking; embedding model selection and evaluation. ● Retrieval evaluation in practice. Recall@k, nDCG/MRR and a groundedness measure, used to justify changes. ● Document processing. Messy real-world PDF, PPTX and DOCX; OCR pipelines and their failure modes; vision-language models for charts and infographics. ● Python and FastAPI in production. Python 3.11+, async, Pydantic, streaming/SSE, OpenAPI contracts, API versioning. ● Service engineering. Queue and worker patterns, idempotency, retries, dead-letter handling; Docker, pytest, Git and CI; structured logging and tracing. ● AWS as a consumer. S3, ECS/EKS, Lambda, Bedrock, Textract, OpenSearch; a vector database in production with collection design and metadata filtering. Preferred Qualifications ● Life sciences, pharma or other regulated content; PII/PHI handling. ● Multimodal retrieval; MCP or agent tool surfaces. ● Judgement on frameworks — has used LangChain or LlamaIndex and can say when not to. Key Notes: NOT A FIT FOR THIS ROLE Infrastructure or DevOps engineers · data platform and ETL engineers · data scientists without service-building experience · candidates whose retrieval exposure is limited to calling a managed RAG API.
Key Responsibilities
- Manage evaluation harness
- Implement metadata extraction
Requirements
Must have
- 6–10 years of experience