About the Role

We build and operate a large-scale, air-gapped, on-premises intelligence platform deployed into enterprise and government data centers. Our infrastructure spans Kubernetes clusters on bare-metal Dell PowerEdge servers, NVIDIA GPU nodes (H100/H200 SXM), distributed object storage, and a multi-zone network architecture with strict air-gap requirements. You will own the full infrastructure lifecycle — from initial rack-and-stack to observability, security hardening, and GitOps-driven continuous deployment.

This is not a cloud-native role. We run on-prem first, with Azure DR as a secondary layer. You need to be comfortable at the hardware level and in the Kubernetes control plane simultaneously.

Key Responsibilities

  • Deploy, manage, and scale RKE2 Kubernetes clusters on bare-metal Dell R750xs/R740xd servers, including control plane hardening and CIS benchmark compliance
  • Manage distributed storage infrastructure: MinIO erasure-coded object storage (300TB+), OpenSearch/Elasticsearch hot/warm/cold tiers, and NVMe-backed stateful workloads
  • Build and maintain GitOps pipelines using ArgoCD, GitHub Actions, and Harbor private registry — with full air-gap support for offline image lifecycle management
  • Configure and administer Keycloak for enterprise SSO, RBAC, OAuth2/OIDC, and LDAP federation across all platform services
  • Implement and operate the full observability stack: OpenTelemetry Collector → Loki (logs) + Prometheus (metrics) + Tempo (traces) + Grafana (dashboards)
  • Manage HashiCorp Vault in offline Raft mode for secrets, certificates, and encryption key lifecycle
  • Configure and troubleshoot Istio service mesh for inter-service mTLS, traffic policies, and distributed tracing propagation
  • Support GPU cluster operations: NVIDIA driver management, CUDA toolkit versioning, MIG partitioning on H100/H200 SXM nodes
  • Manage network infrastructure: Cisco Nexus spine-leaf topology, 100GbE Intel E810 NICs, VLAN segmentation, BGP, and OOB management via iDRAC/IPMI
  • Run disaster recovery operations across dual on-prem sites (Abu Dhabi primary / Dubai secondary) and Azure DR cloud fallback
  • Participate in on-call; lead incident response, post-mortems, and RCA documentation

Tech Stack You'll Own

  • Kubernetes RKE2 — air-gapped, CIS-hardened. Helm, ArgoCD (GitOps). Istio service mesh.
  • Storage MinIO (distributed object storage, 300TB+). OpenSearch (hot/warm/cold). Neo4j (Graph DB). PGVector. PostgreSQL HA.
  • Observability OpenTelemetry Collector. Loki. Prometheus. Tempo. Grafana. OPIK. Arize AI (ML observability).
  • Security / Auth Keycloak (SSO, RBAC, OAuth2/OIDC). HashiCorp Vault (offline Raft). CIS benchmarks. Network policies.
  • Registry / CI Harbor (private, air-gap). GitHub Actions. GitLab CI. ArgoCD. SonarQube (code quality gates).
  • Messaging RabbitMQ. Confluent Kafka. Redis (dedup, caching).
  • Hardware Dell R750xs / R740xd. NVIDIA H100/H200 SXM. Intel E810 100GbE. Cisco Nexus N9K. iDRAC/IPMI.
  • DR / Cloud Azure (secondary DR only). Dual on-prem sites. Offline-first architecture.

Required Qualifications

  • 4–7 years of hands-on infrastructure, systems, or DevOps engineering experience
  • Kubernetes — deep operational experience (CKA/CKAD preferred); RKE2 or k3s experience a strong plus
  • Linux systems administration at depth: systemd, networking stack, storage, kernel tuning
  • Experience operating MinIO or comparable distributed object storage at scale
  • GitOps workflows: ArgoCD or Flux, Helm, private registry management
  • Strong networking: Cisco switching, BGP, VLAN, firewall rules, service mesh (Istio/Linkerd)
  • Scripting proficiency: Python and Bash minimum; Go is a bonus
  • Familiarity with Keycloak or comparable IAM/SSO platforms

Strong Advantage

  • Air-gapped or on-prem-first deployment experience — this is not optional in our client environments
  • NVIDIA GPU infrastructure: driver management, CUDA, MIG partitioning, InfiniBand/NVLink
  • HashiCorp Vault in offline/Raft mode
  • Istio distributed tracing and mTLS configuration
  • OpenTelemetry instrumentation and full observability stack ownership
  • Dell PowerEdge hardware (R750xs, R740xd) hands-on experience

What We Offer

  • Competitive remote compensation benchmarked globally
  • Work on one of the most technically complex on-prem AI deployments in the region
  • GPU infrastructure exposure (H100/H200 SXM) you won't find at most companies
  • Async-first culture with strong engineering discipline

How to Apply

Submit your resume and a cover letter outlining your experience