workfromanywhereworkfromanywhere
All jobs
nameDevOps

AI Infrastructure Engineer

Remote (Global)$150K - $200KPosted 2 days ago

vCluster is seeking an AI Infrastructure Specialist to work directly with customers on deploying GPU infrastructure and Kubernetes environments, from initial setup to production. The role involves technical deployment, optimization, validation, knowledge transfer, and pre-sales collaboration in a fast-growing AI cloud platform.

Location: Remote (Global)

Salary: $150K - $200K

Responsibilities

  • Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment.
  • Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand.
  • Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.
  • Knowledge Transfer: Work alongside customer teams to build self-sufficiency, ensuring they can operate and grow the platform independently.
  • Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer's head start.
  • Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap.
  • Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value.

Requirements

  • 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments.
  • Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes.
  • Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments.
  • Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn.
  • Comfort operating in ambiguous, fast-moving environments where you are often writing the playbook in real time.
  • You thrive in environments that reject legacy tech and prefer a modern stack where you can solve a variety of problems from pipelines to internal services.

Benefits

  • Competitive Salary: We offer a competitive compensation package, including equity.
  • Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).
  • Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.
  • Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

Additional Information

  • This role is remote and involves working closely with customers and internal teams to deploy and optimize GPU infrastructure and Kubernetes environments. The company is a venture-backed startup with a focus on Kubernetes virtualization for AI workloads, offering a competitive salary, benefits, and a flexible work culture.

Location

Remote (Global)

Salary

$150K - $200K

Category

DevOps

Company

name

Source

himalayas

Posted

2 days ago

Share this job

XLinkedIn

Similar remote jobs

ToparoNewDevOps

DevOps Director

Remote (Canada)
today
ModivcareNewDevOps

Senior Manager, DevOps

Remote$140,000-$185,000
today
MiratechNewDevOps

DevOps Engineer With Splunk

Remote (India)
today
FirstupNewDevOps

Director Of Cloud Operations

Remote - US$200,000–$228,000
today
LeidosNewDevOps

OCI DevOps Engineer

Remote (US)$107,900.00 - $195,050.00
today