Senior LLM / Generative AI Engineer
We are looking for a Senior LLM / Generative AI Engineer to own the end-to-end engineering of production-grade Large Language Model (LLM) solutions. The role will cover model selection, prompt engineering, RAG, fine-tuning, AI agents, deployment, LLMOps, evaluation, security, performance optimization, and ongoing production operations.
The ideal candidate should have strong hands-on experience building and deploying LLM-powered enterprise applications and should be comfortable working across AI engineering, cloud/GPU infrastructure, DevOps, security, and system architecture.
Key Responsibilities
- Evaluate and select commercial and open-source LLMs based on accuracy, latency, context length, cost, licensing, data residency, and governance requirements.
- Design and implement prompt engineering strategies, few-shot prompting, structured outputs, function/tool calling, and context management.
- Build production-grade RAG pipelines, semantic search, document understanding, summarization, embeddings, chunking, metadata extraction, and vector search.
- Implement LLM customization using LoRA, QLoRA, PEFT, and fine-tuning where required.
- Build agentic and multi-agent workflows with tool/API integration, memory, state management, and human approval mechanisms.
- Deploy and operate LLM inference platforms across cloud, hybrid, and GPU environments.
- Work with technologies such as vLLM, TensorRT-LLM, TGI, or equivalent inference/serving frameworks.
- Optimize inference using quantization, batching, KV-cache management, speculative decoding, and prompt caching.
- Implement LLMOps practices including model/prompt versioning, deployment pipelines, rollback, canary releases, and A/B testing.
- Develop automated LLM evaluation frameworks, regression testing, hallucination/groundedness checks, and LLM-as-a-judge workflows.
- Implement observability covering latency, TTFT, tokens/second, token usage, failures, refusals, quality regression, and model drift.
- Optimize LLM infrastructure and operational costs through model routing, caching, right-sizing, and capacity planning.
- Implement AI security controls covering prompt injection, data leakage, PII protection, insecure tool use, and adversarial inputs.
- Apply responsible AI, privacy, governance, auditability, and human-in-the-loop principles.
- Mentor engineers and contribute to architecture reviews, technical standards, documentation, and internal training.
- Collaborate with product, data, security, infrastructure, and business teams to deliver enterprise AI solutions.
Required Skills & Experience
- 7–12+ years of experience in software engineering, AI/ML, or related fields.
- Strong hands-on experience with LLMs and Generative AI.
- Strong Python programming and API/microservices development experience.
- Experience with RAG, embeddings, vector databases, semantic search, and LLM applications.
- Experience with LangChain, LangGraph, LlamaIndex, Semantic Kernel, or similar frameworks.
- Experience with LLM fine-tuning and LoRA/QLoRA/PEFT.
- Experience with LLM inference and serving technologies such as vLLM, TensorRT-LLM, TGI, or equivalent.
- Strong understanding of Docker, Kubernetes, CI/CD, Git, and cloud environments.
- Experience with AWS, Azure, or GCP and GPU-based infrastructure.
- Experience with vector databases such as Pinecone, Milvus, Qdrant, Weaviate, pgvector, Elasticsearch/OpenSearch, or equivalent.
- Understanding of LLM evaluation, monitoring, observability, and LLMOps/MLOps.
- Strong understanding of AI security risks, including prompt injection, data leakage, and insecure tool calling.
- Excellent problem-solving, architecture, communication, and technical leadership skills.
Preferred Experience
- Experience with open-source models such as Llama, Mistral, Qwen, or equivalent.
- Experience with MLflow, Weights & Biases, LangSmith, RAGAS, DeepEval, promptfoo, or similar tools.
- Experience with GPU optimization and inference performance tuning.
- Knowledge of OWASP Top 10 for LLM Applications.
- Experience working with enterprise AI governance, privacy, and responsible AI frameworks.
What We Are Looking For
We are looking for someone who can take an LLM solution from:
Model Selection → Prompt/RAG → Fine-Tuning → Evaluation → Agent/API Integration → Deployment → Monitoring → Optimization → Production Operations
Candidates with strong hands-on production LLM engineering experience will be preferred over candidates whose experience is limited to basic ChatGPT/API integrations or proof-of-concept chatbot development.
Interested candidates can apply with their updated CV highlighting relevant LLM/Generative AI projects and production experience.
Pay: AED12,000.00 - AED15,000.00 per month
Application Question(s):
- Do you have hands-on production experience with LLM/Generative AI systems?
- Do you have hands-on experience building and deploying RAG (Retrieval-Augmented Generation) solutions?
Experience:
- hands-on experience do you have with LLM/Generative AI: 4 years (Required)
Location:
- Dubai (Preferred)
Willingness to travel:
- 100% (Preferred)
Work Location: In person
Tags & Focus Areas
About INFRA ASSURE
Ready to Join the Team?
Apply once with DevFound — we route your profile to INFRA ASSURE and keep you posted on matching AI roles.