Senior LLM / Generative AI Engineer in Dubai at INFRA ASSURE
INFRA ASSURE

Senior LLM / Generative AI Engineer

INFRA ASSURE Dubai, DU, AE
Full-time $12k Posted 12 days ago

We are looking for a Senior LLM / Generative AI Engineer to own the end-to-end engineering of production-grade Large Language Model (LLM) solutions. The role will cover model selection, prompt engineering, RAG, fine-tuning, AI agents, deployment, LLMOps, evaluation, security, performance optimization, and ongoing production operations.

The ideal candidate should have strong hands-on experience building and deploying LLM-powered enterprise applications and should be comfortable working across AI engineering, cloud/GPU infrastructure, DevOps, security, and system architecture.

Key Responsibilities

  • Evaluate and select commercial and open-source LLMs based on accuracy, latency, context length, cost, licensing, data residency, and governance requirements.
  • Design and implement prompt engineering strategies, few-shot prompting, structured outputs, function/tool calling, and context management.
  • Build production-grade RAG pipelines, semantic search, document understanding, summarization, embeddings, chunking, metadata extraction, and vector search.
  • Implement LLM customization using LoRA, QLoRA, PEFT, and fine-tuning where required.
  • Build agentic and multi-agent workflows with tool/API integration, memory, state management, and human approval mechanisms.
  • Deploy and operate LLM inference platforms across cloud, hybrid, and GPU environments.
  • Work with technologies such as vLLM, TensorRT-LLM, TGI, or equivalent inference/serving frameworks.
  • Optimize inference using quantization, batching, KV-cache management, speculative decoding, and prompt caching.
  • Implement LLMOps practices including model/prompt versioning, deployment pipelines, rollback, canary releases, and A/B testing.
  • Develop automated LLM evaluation frameworks, regression testing, hallucination/groundedness checks, and LLM-as-a-judge workflows.
  • Implement observability covering latency, TTFT, tokens/second, token usage, failures, refusals, quality regression, and model drift.
  • Optimize LLM infrastructure and operational costs through model routing, caching, right-sizing, and capacity planning.
  • Implement AI security controls covering prompt injection, data leakage, PII protection, insecure tool use, and adversarial inputs.
  • Apply responsible AI, privacy, governance, auditability, and human-in-the-loop principles.
  • Mentor engineers and contribute to architecture reviews, technical standards, documentation, and internal training.
  • Collaborate with product, data, security, infrastructure, and business teams to deliver enterprise AI solutions.

Required Skills & Experience

  • 7–12+ years of experience in software engineering, AI/ML, or related fields.
  • Strong hands-on experience with LLMs and Generative AI.
  • Strong Python programming and API/microservices development experience.
  • Experience with RAG, embeddings, vector databases, semantic search, and LLM applications.
  • Experience with LangChain, LangGraph, LlamaIndex, Semantic Kernel, or similar frameworks.
  • Experience with LLM fine-tuning and LoRA/QLoRA/PEFT.
  • Experience with LLM inference and serving technologies such as vLLM, TensorRT-LLM, TGI, or equivalent.
  • Strong understanding of Docker, Kubernetes, CI/CD, Git, and cloud environments.
  • Experience with AWS, Azure, or GCP and GPU-based infrastructure.
  • Experience with vector databases such as Pinecone, Milvus, Qdrant, Weaviate, pgvector, Elasticsearch/OpenSearch, or equivalent.
  • Understanding of LLM evaluation, monitoring, observability, and LLMOps/MLOps.
  • Strong understanding of AI security risks, including prompt injection, data leakage, and insecure tool calling.
  • Excellent problem-solving, architecture, communication, and technical leadership skills.

Preferred Experience

  • Experience with open-source models such as Llama, Mistral, Qwen, or equivalent.
  • Experience with MLflow, Weights & Biases, LangSmith, RAGAS, DeepEval, promptfoo, or similar tools.
  • Experience with GPU optimization and inference performance tuning.
  • Knowledge of OWASP Top 10 for LLM Applications.
  • Experience working with enterprise AI governance, privacy, and responsible AI frameworks.

What We Are Looking For

We are looking for someone who can take an LLM solution from:

Model Selection → Prompt/RAG → Fine-Tuning → Evaluation → Agent/API Integration → Deployment → Monitoring → Optimization → Production Operations

Candidates with strong hands-on production LLM engineering experience will be preferred over candidates whose experience is limited to basic ChatGPT/API integrations or proof-of-concept chatbot development.

Interested candidates can apply with their updated CV highlighting relevant LLM/Generative AI projects and production experience.

Pay: AED12,000.00 - AED15,000.00 per month

Application Question(s):

  • Do you have hands-on production experience with LLM/Generative AI systems?
  • Do you have hands-on experience building and deploying RAG (Retrieval-Augmented Generation) solutions?

Experience:

  • hands-on experience do you have with LLM/Generative AI: 4 years (Required)

Location:

  • Dubai (Preferred)

Willingness to travel:

  • 100% (Preferred)

Work Location: In person

Tags & Focus Areas

Fulltime Ai Ai Engineer Generative Ai

Ready to Apply?

Join INFRA ASSURE and help shape the future of AI.

Save for later

About INFRA ASSURE

Ready to Join the Team?

Apply once with DevFound — we route your profile to INFRA ASSURE and keep you posted on matching AI roles.