Browse DevOps & SRE Jobs
Search 95 curated tech job listings scraped in real-time from LinkedIn, Glassdoor, RemoteOK, and more. Filter by role, location, seniority, and source to find your next opportunity.
95 jobs found for "GPU platforms"
Page 1 of 5
Infrastructure Engineer
…company's first dedicated infrastructure hire. You'll own the infrastructure supporting its cloud platform, GPU workloads, and distributed edge-device fleet as deployments scale across hundreds of customer locations. What … pipelines and deployment automation Implement logging, metrics, alerts, and incident response Improve platform reliability, scalability, and cost efficiency Support GPU workload observability and edge-device fleet automation Establish infrastructure best practices…
DevOps Engineer
…help improve reliability, scalability, security, and observability across Kubernetes clusters, AWS infrastructure, GPU-enabled workloads, SageMaker environments, and data/ML platform components. They will support production-ready ML workflows by helping improve experimentation … Product and Data teams, helping reduce operational overhead through automation, documentation, alerting, GPU resource management, and repeatable platform standards. What you bring: 2–5 years of hands-on DevOps or cloud engineering…
Pre-Sales Solution Architect – AI Infrastructure
…open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge … AI/ML deployment pipelines model serving and inference platforms Kubeflow, MLflow, Ray, or similar frameworks GPU-accelerated workloads Platform Engineering GitOps workflows infrastructure automation Cloud & Hybrid Infrastructure AWS, Azure, GCP hybrid cloud…
Senior DevOps Engineer, Platform Engineering
…computer graphics, through our invention of the GPU. The GPU has proven extremely effective in solving complex computer science challenges. Today, NVIDIA's GPU powers deep learning algorithms, simulating human intelligence … exciting time to join us! NVIDIA invites applications for a Senior DevOps Platform Engineer skilled in Platform and Release Engineering to join the Metropolis team. The role involves developing, building, and maintaining…
Software Engineer, Machine Learning Infrastructure
…engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives. Shape the future of the centralized GenAI platform — including emerging directions such … evaluation GPU performance work — multi-node/distributed inference, KV-cache/memory optimization, quantization (FP8/INT8/AWQ/GPTQ), or cold-start/throughput tuning Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (e.g., Modal…
DevOps Engineer
…description Design, deploy, automate, and operate scalable infrastructure and cloud-native platform services Contribute to Kubernetes-based AI/ML and HPC platforms, including CI/CD, GitOps, observability, security, and operational tooling Collaborate with researchers … systems concepts, APIs, scalability, observability, identity and access management, and security AI/ML platforms and supporting infrastructure services HPC systems, GPU clusters, and large-scale infrastructure environments Platform engineering and developer productivity tooling…
DevOps Engineer
…Description Design, deploy, automate, and operate scalable infrastructure and cloud-native platform services Contribute to Kubernetes-based AI/ML and HPC platforms, including CI/CD, GitOps, observability, security, and operational tooling Collaborate with researchers … systems concepts, APIs, scalability, observability, identity and access management, and security AI/ML platforms and supporting infrastructure services HPC systems, GPU clusters, and large-scale infrastructure environments Platform engineering and developer productivity tooling…
Platform Engineer
…related cloud-native technologies. Develop automation solutions for provisioning, configuration management, monitoring, and platform lifecycle management. Ensure platform reliability through proactive monitoring, capacity planning, performance tuning, and incident management. Support production environments … technologies such as Python, Bash/Shell, or Ansible. Solid knowledge of container orchestration platforms, particularly Kubernetes, Terraform or OpenShift and GPU Hands-on experience in Prometheus and Grafana…
DevOps Engineer - AWS
…Cloud Engineer to design, provision, optimize, and support the AWS infrastructure powering our AMD GPU AI/HPC platform. This is a hands-on execution role — you'll work closely with Rust backend engineers … collaborate across engineering teams Preferred Qualifications Experience with AI/ML, GPU, or HPC workloads Kubernetes on AWS (EKS or self-managed) Observability platforms: Prometheus, Grafana, Loki, OpenTelemetry, Datadog AWS cost optimization: right-sizing…
Senior Platform Engineer – Cloud & ML Platform
…Develop infrastructure-as-code, automation, and GitOps workflows to ensure reproducible, auditable, and efficient platform operations. Manage GPU-enabled workloads, scheduling, storage, networking, secrets, access control, and cost-aware resource utilization. Improve … English is a matter of course for you. Additional plus Experience building internal developer platforms or ML platform products. Experience with distributed storage systems, object storage, data lake architectures, or high-throughput…
DevOps Engineer New
…Kubernetes-based AI infrastructure. This hands-on role will focus on cloud-native platform engineering, GPU-accelerated workloads, bare-metal servers, automation and systems reliability, helping to deliver a secure, scalable … high-performing AI platform for ML engineers, data scientists and enterprise customers. You will help build and operate our AI infrastructure, supporting Linux systems, Kubernetes platforms and automation initiatives. Working alongside experienced…
DevOps Engineer
…looking for a DevOps / Platform Engineer to own and evolve our cloud and Kubernetes infrastructure. You will work closely with Product, Engineering, and ML to keep our platform scalable, secure, and production … scripting skills (Bash, Python, etc.) Interest or experience supporting AI/ML systems Nice to have GPU scheduling / ML platform experience Helm or Kustomize Exposure to security and compliance environments…
Staff Platform Engineer New
…that spans hybrid infrastructure in Europe - including large on-prem GPU clusters powering DeepL's research and hundreds of on-prem GPU nodes serving production inference - as well as AWS regions across … reliability, capacity and cost efficiency of DeepL's compute infrastructure, CPU and GPU: the hybrid platform serving our products across AWS and on-prem, and the on-prem clusters powering our research…
DevOps Engineer
…grade environments Improve scalability, reliability and resilience across the platform Design infrastructure that grows with the business Platform Engineering Build internal developer platforms that accelerate engineering productivity Improve deployment pipelines and developer … posture Ensure production environments remain secure as the company scales Experience supporting AI infrastructure, GPU workloads, ML platforms or large-scale distributed systems would be a significant advantage. What We're Looking…
SRE - Seattle
…observability and operational maturity Minimum Qualifications 5+ years in SRE, Infra, or Platform Engineering roles Strong experience with cloud platforms (AWS/GCP/Azure/OCI) Hands-on with Kubernetes and distributed systems Experience in on-call … programming language (Go, Python, Java) Preferred Qualifications Experience supporting AI/ML or data-intensive platforms Experience of GPU cluster management Knowledge of SLO/SLA frameworks Experience in fast-growing or global products Exposure…
DevOps Engineer
…production support, infrastructure automation, platform reliability, and continuous improvement of deployment and operational standards. Key Responsibilities Platform Engineering & DevOps Design, build, automate, and maintain DevOps platforms supporting AI, ML, Agentic, and traditional … environments. Support LLM-based, Agentic AI, Retrieval-Augmented Generation (RAG), and AI workflow platforms. Manage and optimize GPU-based infrastructure for AI training and inference workloads. Collaborate with Data Science…
Site Reliability Engineer
…Site Reliability / Platform Engineer We are recruiting for a growing technology infrastructure business that is building a new operational capability to support large-scale, high-performance computing environments. This is an excellent … Halo, Jira Service Management, OpenTelemetry, distributed tracing, Slack/Teams automation, datacentre or colocation environments, GPU infrastructure, DCIM, IPAM, virtualisation platforms or LLM-assisted operational automation. This is not an AI/ML development position…
Senior DevOps Engineer
…Raft builds mission-critical data, AI, and operational platforms that process high-volume sensor, operational, and mission data across multiple classification levels. These platforms must run reliably in cloud-hosted, classified … shared Helm libraries, environment templates, or self-service deployment workflows Experience supporting GPU-enabled Kubernetes workloads, model-serving platforms, data pipelines, or mission AI/ML workflows Prior experience supporting Pacific Command, PACAF…
DevOps Engineer
…years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or a closely related role. · Hands-on experience with cloud platforms such as AWS and strong familiarity with containerization and orchestration technologies … digital products, ideally in banking, fintech, or other regulated industries. · Exposure to AI/ML platforms, model-serving environments, GPU-enabled infrastructure, or MLOps practices. · Familiarity with security, compliance, and governance expectations in enterprise…
Platform Engineer II/III
…Design and implement auto-scaling compute infrastructure for simulation workloads using cloud platforms • Build and maintain on-premises GPU and CPU clusters for simulation and machine learning training • Architect hybrid cloud solutions … experience in platform engineering, DevOps, SRE, or cloud infrastructure roles • Strong hands-on experience with Kubernetes for container orchestration and workload management • Experience with cloud computing platforms and services (compute, storage, networking…