About the Role
We build and operate a large-scale, air-gapped, on-premises intelligence platform deployed into enterprise and government data centers. Our infrastructure spans Kubernetes clusters on bare-metal Dell PowerEdge servers, NVIDIA GPU nodes (H100/H200 SXM), distributed object storage, and a multi-zone network architecture with strict air-gap requirements. You will own the full infrastructure lifecycle — from initial rack-and-stack to observability, security hardening, and GitOps-driven continuous deployment.
This is not a cloud-native role. We run on-prem first, with Azure DR as a secondary layer. You need to be comfortable at the hardware level and in the Kubernetes control plane simultaneously.
Key Responsibilities
- Deploy, manage, and scale RKE2 Kubernetes clusters on bare-metal Dell R750xs/R740xd servers, including control plane hardening and CIS benchmark compliance
- Manage distributed storage infrastructure: MinIO erasure-coded object storage (300TB+), OpenSearch/Elasticsearch hot/warm/cold tiers, and NVMe-backed stateful workloads
- Build and maintain GitOps pipelines using ArgoCD, GitHub Actions, and Harbor private registry — with full air-gap support for offline image lifecycle management
- Configure and administer Keycloak for enterprise SSO, RBAC, OAuth2/OIDC, and LDAP federation across all platform services
- Implement and operate the full observability stack: OpenTelemetry Collector → Loki (logs) + Prometheus (metrics) + Tempo (traces) + Grafana (dashboards)
- Manage HashiCorp Vault in offline Raft mode for secrets, certificates, and encryption key lifecycle
- Configure and troubleshoot Istio service mesh for inter-service mTLS, traffic policies, and distributed tracing propagation
- Support GPU cluster operations: NVIDIA driver management, CUDA toolkit versioning, MIG partitioning on H100/H200 SXM nodes
- Manage network infrastructure: Cisco Nexus spine-leaf topology, 100GbE Intel E810 NICs, VLAN segmentation, BGP, and OOB management via iDRAC/IPMI
- Run disaster recovery operations across dual on-prem sites (Abu Dhabi primary / Dubai secondary) and Azure DR cloud fallback
- Participate in on-call; lead incident response, post-mortems, and RCA documentation
Tech Stack You'll Own
- Kubernetes RKE2 — air-gapped, CIS-hardened. Helm, ArgoCD (GitOps). Istio service mesh.
- Storage MinIO (distributed object storage, 300TB+). OpenSearch (hot/warm/cold). Neo4j (Graph DB). PGVector. PostgreSQL HA.
- Observability OpenTelemetry Collector. Loki. Prometheus. Tempo. Grafana. OPIK. Arize AI (ML observability).
- Security / Auth Keycloak (SSO, RBAC, OAuth2/OIDC). HashiCorp Vault (offline Raft). CIS benchmarks. Network policies.
- Registry / CI Harbor (private, air-gap). GitHub Actions. GitLab CI. ArgoCD. SonarQube (code quality gates).
- Messaging RabbitMQ. Confluent Kafka. Redis (dedup, caching).
- Hardware Dell R750xs / R740xd. NVIDIA H100/H200 SXM. Intel E810 100GbE. Cisco Nexus N9K. iDRAC/IPMI.
- DR / Cloud Azure (secondary DR only). Dual on-prem sites. Offline-first architecture.
Required Qualifications
- 4–7 years of hands-on infrastructure, systems, or DevOps engineering experience
- Kubernetes — deep operational experience (CKA/CKAD preferred); RKE2 or k3s experience a strong plus
- Linux systems administration at depth: systemd, networking stack, storage, kernel tuning
- Experience operating MinIO or comparable distributed object storage at scale
- GitOps workflows: ArgoCD or Flux, Helm, private registry management
- Strong networking: Cisco switching, BGP, VLAN, firewall rules, service mesh (Istio/Linkerd)
- Scripting proficiency: Python and Bash minimum; Go is a bonus
- Familiarity with Keycloak or comparable IAM/SSO platforms
Strong Advantage
- Air-gapped or on-prem-first deployment experience — this is not optional in our client environments
- NVIDIA GPU infrastructure: driver management, CUDA, MIG partitioning, InfiniBand/NVLink
- HashiCorp Vault in offline/Raft mode
- Istio distributed tracing and mTLS configuration
- OpenTelemetry instrumentation and full observability stack ownership
- Dell PowerEdge hardware (R750xs, R740xd) hands-on experience
What We Offer
- Competitive remote compensation benchmarked globally
- Work on one of the most technically complex on-prem AI deployments in the region
- GPU infrastructure exposure (H100/H200 SXM) you won't find at most companies
- Async-first culture with strong engineering discipline
How to Apply
Submit your resume and a cover letter outlining your experience