Senior Data Scientist (UAE)
Actively Hiring
Full-time Posted 2 days ago
Responsibilities
- check_circle Query and analyse large domain- or topic-specific data sets from both structured and unstructured sources, identify patterns and features. Ensure data meets quality standards and requirements before model development.
- check_circle Design and fine-tune Large Language Models (LLMs) to parse complex regulatory texts (e.g., building codes, military standards) and extract structured rules for automated compliance checking.
- check_circle Convert interpreted regulations into computer-processable formats (e.g., object-property-condition-value tuples) that can be executed by downstream compliance engines.
- check_circle Architect methods for LLMs to map natural language requirements directly to specific metadata entities within various schemas (e.g., mapping 'systems design' to specified attributes).
- check_circle Implement Retrieval-Augmented Generation (RAG) pipelines that allow systems to query vast repositories of technical documentation and historical project data with high accuracy and low hallucination rates.
- check_circle Develop time-series forecasting models to predict spend categories and material demand by correlating internal ERP data with external macroeconomic signals.
- check_circle Build machine learning classifiers to categorize supplier risks and operational anomalies, integrating data from diverse sources to create dynamic risk scores.
- check_circle Design robust pipelines to extract and transform raw data (from Data Lakehouse, external web sources, or SAP and other databases) into features required for predictive modeling and automated rule checking.
- check_circle Work with Back End Engineers to integrate AI models into a cohesive 'compliance engine' or 'risk engine' that can be invoked programmatically via robust APIs.
- check_circle Streamline model performance to ensure complex checks (e.g., analyzing large datasets or processing thousands of supplier records) can be executed within reasonable timeframes, potentially using batching or asynchronous processing.
- check_circle Validate model outputs against known test cases and historical data, debugging false positives/negatives to refine algorithms and ensure 'defense-grade' reliability.
Basic qualifications
- Expert proficiency in Python and standard ML libraries (TensorFlow/PyTorch, Scikit-learn, Pandas, NumPy). Strong grasp of both supervised and unsupervised learning techniques.
- Deep experience with transformer-based models (GPT, BERT, Llama) and prompt engineering techniques (few-shot learning, fine-tuning) for domain-specific tasks.
- Proficiency in handling complex data structures (JSON, XML) and familiarity with database querying (SQL/NoSQL) or graph data structures. Experience with data extraction from specialized formats is a significant plus.
- Understanding of how to expose models via RESTful APIs (Flask/FastAPI) and integrate them into larger software architectures.
- Solid understanding of statistics, probability distribution, A/B testing. Adept at identifying and mitigating biases in datasets
Tags & Focus Areas
Remote Ai Data Science Nlp Generative Ai
About CloudPSO
Ready to Join the Team?
Apply once with DevFound — we route your profile to CloudPSO and keep you posted on matching AI roles.