Browse DevOps & SRE Jobs
Search 862 curated tech job listings scraped in real-time from LinkedIn, Glassdoor, RemoteOK, and more. Filter by role, location, seniority, and source to find your next opportunity.
862 jobs found for "Operations and Resource"
Page 1 of 44
Site Reliability Engineer II New
…Code & Cloud Operations Provision and manage Azure infrastructure using Terraform or OpenTofu, maintaining drift-free state aligned with GitOps principles. Own Kubernetes cluster operations including workload scheduling, resource optimization, RBAC, network policy … equivalent services. AWS/GCP backgrounds considered with clear willingness to operate in Azure. Kubernetes: Deep operational experience with Kubernetes in production: resource management, network policies, RBAC, HPA/VPA, persistent volumes, and debugging live workload…
Site Reliability Engineer II
…Code & Cloud Operations Provision and manage Azure infrastructure using Terraform or OpenTofu, maintaining drift-free state aligned with GitOps principles. Own Kubernetes cluster operations including workload scheduling, resource optimization, RBAC, network policy … equivalent services. AWS/GCP backgrounds considered with clear willingness to operate in Azure. Kubernetes: Deep operational experience with Kubernetes in production: resource management, network policies, RBAC, HPA/VPA, persistent volumes, and debugging live workload…
Site Reliability Engineer II
…Code & Cloud Operations Provision and manage Azure infrastructure using Terraform or OpenTofu, maintaining drift-free state aligned with GitOps principles. Own Kubernetes cluster operations including workload scheduling, resource optimization, RBAC, network policy … equivalent services. AWS/GCP backgrounds considered with clear willingness to operate in Azure. Kubernetes: Deep operational experience with Kubernetes in production: resource management, network policies, RBAC, HPA/VPA, persistent volumes, and debugging live workload…
Site Reliability Engineer II
…Code & Cloud Operations Provision and manage Azure infrastructure using Terraform or OpenTofu, maintaining drift-free state aligned with GitOps principles. Own Kubernetes cluster operations including workload scheduling, resource optimization, RBAC, network policy … equivalent services. AWS/GCP backgrounds considered with clear willingness to operate in Azure. Kubernetes: Deep operational experience with Kubernetes in production: resource management, network policies, RBAC, HPA/VPA, persistent volumes, and debugging live workload…
Site Reliability Engineer II
…Code & Cloud Operations Provision and manage Azure infrastructure using Terraform or OpenTofu, maintaining drift-free state aligned with GitOps principles. Own Kubernetes cluster operations including workload scheduling, resource optimization, RBAC, network policy … equivalent services. AWS/GCP backgrounds considered with clear willingness to operate in Azure. Kubernetes: Deep operational experience with Kubernetes in production: resource management, network policies, RBAC, HPA/VPA, persistent volumes, and debugging live workload…
Site Reliability Engineer II
…Code & Cloud Operations Provision and manage Azure infrastructure using Terraform or OpenTofu, maintaining drift-free state aligned with GitOps principles. Own Kubernetes cluster operations including workload scheduling, resource optimization, RBAC, network policy … equivalent services. AWS/GCP backgrounds considered with clear willingness to operate in Azure. Kubernetes: Deep operational experience with Kubernetes in production: resource management, network policies, RBAC, HPA/VPA, persistent volumes, and debugging live workload…
Site Reliability Engineer II
…Code & Cloud Operations Provision and manage Azure infrastructure using Terraform or OpenTofu, maintaining drift-free state aligned with GitOps principles. Own Kubernetes cluster operations including workload scheduling, resource optimization, RBAC, network policy … equivalent services. AWS/GCP backgrounds considered with clear willingness to operate in Azure. Kubernetes: Deep operational experience with Kubernetes in production: resource management, network policies, RBAC, HPA/VPA, persistent volumes, and debugging live workload…
Platform Engineer - US
…infrastructure-as-code (e.g., Terraform, Helm) to provision and configure cloud and on-prem resources. Deploy and operate workloads across cloud platforms (AWS/Azure) and Kubernetes using GitOps (Argo CD/Flux). Follow tagging, cost … configuration standards when provisioning resources. Help deploy and operate AI workload infrastructure, including model gateways, retrieval services, orchestration components, and supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services…
Platform Engineer - London
…infrastructure-as-code (e.g., Terraform, Helm) to provision and configure cloud and on-prem resources. Deploy and operate workloads across cloud platforms (AWS/Azure) and Kubernetes using GitOps (Argo CD/Flux). Follow tagging, cost … configuration standards when provisioning resources. Help deploy and operate AI workload infrastructure, including model gateways, retrieval services, orchestration components, and supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services…
Platform Engineer
…infrastructure-as-code (e.g., Terraform, Helm) to provision and configure cloud and on-prem resources. Deploy and operate workloads across cloud platforms (AWS/Azure) and Kubernetes using GitOps (Argo CD/Flux). Follow tagging, cost … configuration standards when provisioning resources. Help deploy and operate AI workload infrastructure, including model gateways, retrieval services, orchestration components, and supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services…
Platform Engineer
…infrastructure-as-code (e.g., Terraform, Helm) to provision and configure cloud and on-prem resources. Deploy and operate workloads across cloud platforms (AWS/Azure) and Kubernetes using GitOps (Argo CD/Flux). Follow tagging, cost … configuration standards when provisioning resources. Help deploy and operate AI workload infrastructure, including model gateways, retrieval services, orchestration components, and supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services…
Infrastructure Engineer New
…join our Technology Solutions Group (TSG) , the team responsible for designing, building, and operating the technology platforms and services that enable Bain's global business. Working alongside infrastructure engineers, cloud specialists, architects … optimize performance, reliability, security, and operational efficiency across cloud environments. Resolve infrastructure issues independently and contribute to continuous operational improvements. Improve technical documentation and knowledge resources to support operational excellence. Partner with…
Infrastructure Engineer New
…join our Technology Solutions Group (TSG) , the team responsible for designing, building, and operating the technology platforms and services that enable Bain's global business. Working alongside infrastructure engineers, cloud specialists, architects … optimize performance, reliability, security, and operational efficiency across cloud environments. Resolve infrastructure issues independently and contribute to continuous operational improvements. Improve technical documentation and knowledge resources to support operational excellence. Partner with…
Infrastructure Engineer
…join our Technology Solutions Group (TSG) , the team responsible for designing, building, and operating the technology platforms and services that enable Bain's global business. Working alongside infrastructure engineers, cloud specialists, architects … optimize performance, reliability, security, and operational efficiency across cloud environments. Resolve infrastructure issues independently and contribute to continuous operational improvements. Improve technical documentation and knowledge resources to support operational excellence. Partner with…
Principal DevOps Engineer
…infrastructure, automation, Linux operations, Kubernetes platforms, security, compliance, and operational excellence. This role is distinct from SRE and is centered on infrastructure and platform operations rather than Java application reliability. What … Large-scale AWS cloud infrastructure operations Terraform and configuration management using Ansible/Puppet Linux fleet management and automation Kubernetes and cloud-native platform operations Operational tooling, observability implementation and platform supportability Security, compliance…
DevOps and DataOps Specialist
…well as pipeline automation including planning, development, continuous delivery, and operations phases (CI/CD, DataOps). During the production launch and operational maintenance of hybrid platforms, you will be in charge of requirement gathering … infrastructure for AI workloads. FinOps Optimization: Analyzing and optimizing cloud costs to ensure efficient resource utilization. Daily Operations: Monitoring system performance, access management, and security updates. YOUR JOURNEY INCLUDES University or college…
Cloud Engineer New
…operations. This is an excellent opportunity for a technically curious individual eager to build a career in cloud engineering. How You'll Create Impact Provision, configure, and maintain Azure cloud resources across … cloud operations including incident response, change management, access reviews, and cloud resource tagging. Deploy, operate, and support containerized applications on Azure Kubernetes Service (AKS) with an emphasis on secure-by-default…
DevOps Engineer
…experience with git version control, git branching, and CI/CD practices Establishing visibility into cloud operations through: Leveraging resource tagging to allocate costs and optimize resource planning Assisting in preparing cost analysis based … which may include: Migrating applications using microservices architectures Confirming the migration of resources into AWS and decommissioning on-premises resources Supporting rigorous project governance and execution achieved through: Meeting with team members…
Head of Physical Infrastructure
…physical compute infrastructure, from the first assessment of a potential deployment through to the reliable operation of the resulting cluster. You will scope sites, coordinate the design cluster infrastructure and deployment across … team responsible for deploying new clusters and maintaining the reliability and availability of the resources we operate What Sets You Apart Experience designing, deploying, or operating high-density compute infrastructure, HPC clusters…
DevOps Specialist
…Infrastructure as Code, CI/CD pipelines, observability, security/compliance and high-availability operations. The ideal candidate partners closely with Development, QA, Security and Operations teams to embed DevOps practices and continuously improve delivery reliability … availability, disaster recovery, fault tolerance and operational standards. Build and operate container platforms using OpenShift/Kubernetes, Docker or compatible runtimes, registries, ingress/routes, namespaces/projects, RBAC, secrets/config maps, resource controls, operators and deployment manifests. Implement…