Find The RightJob.
The AI Infrastructure Engineer is a platform specialist responsible for architecting, building, and operating high-performance AI infrastructure to support advanced AI workloads, including LLMs, GenAI, Computer Vision, and MLOps. This role will focus on managing GPU clusters (NVIDIA A100/H100), deploying and maintaining Red Hat OpenShift AI (RHODS), and ensuring secure, scalable, and cost-efficient AI platforms across SDD’s Sovereign Cloud and hybrid/multi-cloud environments. The engineer will enable enterprise-grade AI adoption for 200+ government entities.
Key Responsibilities & Deliverables
GPU & AI Platform Architecture
Design and implement GPU-based compute clusters. Define reference architectures for LLM hosting, Vector Databases, MLOps, and high-performance storage/networking.
Fully operational GPU-based AI infrastructure. GPU Cluster Uptime and Performance Utilization. Reduction in Cost per Training/Inference Workload.
GPU Cluster Operations
Install, configure, and optimize core components: CUDA, cuDNN, NCCL, NVIDIA Drivers, and GPU Operators. Implement GPU partitioning, scheduling, and performance tuning for high-end GPUs (e.g., A100/H100).
High-availability architecture for all AI workloads. Complete documentation and runbooks.
OpenShift AI (RHODS) Management
Deploy, configure, and maintain the Red Hat OpenShift AI (RHODS) platform for multi-tenant use. Manage the integration of NVIDIA GPU Operator for efficient GPU scheduling and support Data Scientists with Notebooks, Training, and Inference Endpoints.
Production-ready OpenShift AI (RHODS) platform. AI Project Onboarding Speed.
LLM & Model Serving
Build and manage infrastructure for hosting and serving open-source LLM frameworks (Llama, Falcon, Mistral) and supporting RAG pipelines, LoRA adapters, and Vector Databases (Milvus, pgvector).
Multi-model LLM serving environment for entities. MLOps Pipeline Success Rate and Deployment Frequency.
MLOps & Automation
Implement IaC (Terraform, Ansible) and GitOps for the automated lifecycle management of the AI platform (node onboarding, scaling, model rollout/rollback). Build robust MLOps pipelines for data prep, training, evaluation, and monitoring (using tools like MLflow/Kubeflow).
Infrastructure automation via Terraform & Ansible. Automation Coverage for AI Infrastructure.
Required Qualifications & Experience
Essential Skills & Competencies
Preferred Certifications
Similar jobs
VentureOne
Abu Dhabi, United Arab Emirates
29 days ago
Klanik
Abu Dhabi, United Arab Emirates
29 days ago
Deeplight AI
Dubai, United Arab Emirates
29 days ago
Deeplight
Abu Dhabi, United Arab Emirates
29 days ago
TAT IT Technolgies
Abu Dhabi, United Arab Emirates
29 days ago
Ceenex Global LLC
Sharjah, United Arab Emirates
29 days ago
Jaheziya
Abu Dhabi, United Arab Emirates
about 1 month ago
© 2026 Qureos. All rights reserved.