Browse DevOps & SRE Jobs
Search 555 curated tech job listings scraped in real-time from LinkedIn, Glassdoor, RemoteOK, and more. Filter by role, location, seniority, and source to find your next opportunity.
555 jobs found for "Datadog"
Page 1 of 28
Site Reliability Engineer
…Reliability Engineering (SRE), DevOps, or Systems Architecture, with at least 3+ years specializing deeply in Datadog administration and configuration. Cloud & Container Expertise: Deep professional experience working with AWS, Azure, or GCP, paired … scripting languages (e.g., Python, Go, Bash, or JavaScript) and experience instrumenting applications for APM. Datadog Mastery: Deep understanding of Datadog’s core pillars—Infrastructure, APM, Logs, Metrics, Synthetics, and Security Monitoring. Datadog…
Site Reliability Engineer
…observability of our infrastructure across GCP and Vercel Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards Manage IAM, service accounts, and security best practices … Solid understanding of IAM, security, and cloud best practices Experience with observability tools like Datadog Familiarity with Node.js environments AI is part of your daily engineering workflow you use AI tools thoughtfully…
Site Reliability Engineer
…OpenShift, Kubernetes). Expertise in automation tools experience (preferably Ansible). Expertise in observability tools including APM (Datadog), synthetic monitoring and log aggregation (Elk). Experience in dashboarding tools such as Grafana and Kibana. Understating … Cloud and on Prem deployments. Good experience in Systems Observability and APM tools, preferably Datadog. Strong ability to track and contribute to technical discussions around application integration and high-availability, resilience…
DevOps Specialist
…manage scalable CI/CD pipelines, manage infrastructure through GitOps principles, and ensure observability using tools like Datadog. The Opportunity: Key Roles & Responsibilities: Design, implement, and optimize GitLab CI/CD pipelines for build, test … production releases Manage pipeline orchestration, triggers, and scheduling mechanisms Implement logging, monitoring, and ing using Datadog dashboards and pipeline logs Ensure high availability, reliability, and scalability of CI/CD platforms Collaborate with development…
DevOps Engineer
…strong focus will be placed on Azure DevOps, AKS, CI/CD automation, Infrastructure-as-Code, Datadog observability, cloud-native operations, and AI-assisted automation . Responsibilities: Design, implement, maintain, and continuously improve CI/CD pipelines … release management, deployment governance, and production delivery processes; Implement and maintain observability capabilities using Datadog, including monitors, dashboards, logs, metrics, alerts, and integrations; Collaborate with SRE and engineering teams to improve system…
Site Reliability Engineer
…standards. Implement policy-as-code and automation frameworks. Reduce manual infrastructure management through automation. Observability & Datadog Strategy Define observability standards and monitoring frameworks. Manage Datadog implementation including dashboards, alerts, APM, logs … standards. Partner with DevOps teams on release governance. Preferred Skill And Experience Harness Helm Charts Datadog Incident Management Preferred Certifications AWS DevOps Engineer Professional AWS Solutions Architect Professional Terraform Associate FinOps Practitioner…
DevOps Engineer New
…DevOps Engineer – Financial Services, CI/CD, AWS, Jenkins, DataDog I am looking for a DevOps Engineer to join the Equity Technology team of a leading financial services organisation based in Galway. This … maintain performance testing scripts, frameworks, and automation solutions. Build dashboards and monitoring solutions using DataDog to improve operational visibility. Collaborate with developers, QA teams, and business stakeholders within an Agile environment Required…
Site Reliability Engineer
…ensure reliable, scalable cloud environments. Implement and enhance observability solutions using tools like New Relic, DataDog, Sumologic and Splunk for monitoring, logging, and alerting. Perform code deployments and manage CI/CD pipelines using … Helm, GitHub. Experience with monitoring tools and incident management processes like Prometheus, Grafana, New Relic, DataDog, Splunk, Cloudwatch, Sumologic, etc. Extensive understanding of networking and security concepts. Bonus Points For Specialized…
Site Reliability Engineer
…observability of our infrastructure across GCP and Vercel Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards Manage IAM, service accounts, and security best practices … Solid understanding of IAM, security, and cloud best practices Experience with observability tools like Datadog Familiarity with Node.js environments AI is part of your daily engineering workflow you use AI tools thoughtfully…
Platform Engineer
…maintain dashboards, alerts, and telemetry pipelines using Grafana, Prometheus/Mimir, Loki, Grafana Alloy, Tempo, OpenTelemetry, Datadog, and CloudWatch. Create actionable metrics, log views, and traces that help engineering and operations teams see what … Auto Scaling, Kubernetes, Helm Observability: Grafana, Prometheus/Mimir, Loki, Alloy, Tempo, OpenTelemetry, Datadog, CloudWatch IaC & Automation: Terraform, Terraform Cloud, Ansible, Bash, PowerShell, Python CI/CD & Source Control: Azure DevOps Pipelines, GitHub, GitLab Linux & Edge…
Site Reliability Engineer
…reliable and scalable cloud environments. Implement and enhance observability solutions using tools like New Relic, DataDog, Sumologic and Splunk for monitoring, logging, and alerting. Perform code deployments and manage CI/CD pipelines using … Helm, GitHub. Experience with monitoring tools and incident management processes like Prometheus, Grafana, New Relic, DataDog, Splunk, Cloudwatch, Sumologic etc. Extensive understanding of networking and security concepts. Bonus Points For Specialized…
Platform Engineer
…maintain dashboards, alerts, and telemetry pipelines using Grafana, Prometheus/Mimir, Loki, Grafana Alloy, Tempo, OpenTelemetry, Datadog, and CloudWatch. Create actionable metrics, log views, and traces that help engineering and operations teams see what … Auto Scaling, Kubernetes, Helm Observability: Grafana, Prometheus/Mimir, Loki, Alloy, Tempo, OpenTelemetry, Datadog, CloudWatch IaC & Automation: Terraform, Terraform Cloud, Ansible, Bash, PowerShell, Python CI/CD & Source Control: Azure DevOps Pipelines, GitHub, GitLab Linux & Edge…
DevOps Engineer
…DevOps Engineer – Financial Services, CI/CD, AWS, Jenkins, DataDog I am looking for a DevOps Engineer to join the Equity Technology team of a leading financial services organisation based in Galway. This … maintain performance testing scripts, frameworks, and automation solutions. Build dashboards and monitoring solutions using DataDog to improve operational visibility. Collaborate with developers, QA teams, and business stakeholders within an Agile environment Required…
DevOps Engineer
…Cloud Infrastructure, Containers & K8s, CI/CD Pipelines, Observability, Security & DevSecOps, Automation, Python , Shell Scripting, GitLab / GitLab, Datadog, New Relic Experience: 5-7+ years Location: Englewood Cliffs, NJ 5 Days Onsite We at Coforge … automated workflows using GitLab CI. Observability: Set up deep monitoring, logging, and proactive alerting via Datadog and New Relic. Security & DevSecOps: Integrate security scans, IAM policies, and secrets management; support AI/ML infrastructure…
Sr. DevOps Engineer
…supporting Aurora PostgreSQL, ElastiCache Redis, and Amazon MQ for RabbitMQ in production environments. Experience with Datadog or similar observability platforms. Strong understanding of cloud networking, including VPCs, subnets, routing, security groups … services, ConfigMaps, Secrets, autoscaling. Argo CD, GitOps, CI/CD, release automation, Python, Bash or shell scripting, Datadog, metrics, logs, tracing, dashboards, alerting. Experience supporting high-availability SaaS or customer-facing platforms. Experience with…
Sr. DevOps Engineer
…supporting Aurora PostgreSQL, ElastiCache Redis, and Amazon MQ for RabbitMQ in production environments. Experience with Datadog or similar observability platforms. Strong understanding of cloud networking, including VPCs, subnets, routing, security groups … services, ConfigMaps, Secrets, autoscaling. Argo CD, GitOps, CI/CD, release automation, Python, Bash or shell scripting, Datadog, metrics, logs, tracing, dashboards, alerting. Experience supporting high-availability SaaS or customer-facing platforms. Experience with…
Sr. DevOps Engineer
…supporting Aurora PostgreSQL, ElastiCache Redis, and Amazon MQ for RabbitMQ in production environments. Experience with Datadog or similar observability platforms. Strong understanding of cloud networking, including VPCs, subnets, routing, security groups … services, ConfigMaps, Secrets, autoscaling. Argo CD, GitOps, CI/CD, release automation, Python, Bash or shell scripting, Datadog, metrics, logs, tracing, dashboards, alerting. Experience supporting high-availability SaaS or customer-facing platforms. Experience with…
Sr. DevOps Engineer
…supporting Aurora PostgreSQL, ElastiCache Redis, and Amazon MQ for RabbitMQ in production environments. Experience with Datadog or similar observability platforms. Strong understanding of cloud networking, including VPCs, subnets, routing, security groups … services, ConfigMaps, Secrets, autoscaling. Argo CD, GitOps, CI/CD, release automation, Python, Bash or shell scripting, Datadog, metrics, logs, tracing, dashboards, alerting. Experience supporting high-availability SaaS or customer-facing platforms. Experience with…
DevOps Support Engineer
…Terraform and Ansible. Monitor platform health, reliability, availability and performance using tools such as Datadog and BigPanda, and implement improvements based on observed trends. Participate actively in incident management, including root cause … solutions, including Terraform. Strong familiarity with source control and deployment automation best practices. experience with Datadog for monitoring, observability and alerting. experience with BigPanda or similar event correlation and incident management platforms…
Site Reliability Engineer
…maintain full observability stacks (logging, metrics, tracing) using tools like Prometheus, Grafana, Datadog, OpenTelemetry, or ELK. Analyze telemetry and logs to identify trends, anomalies, and opportunities for improvement. Conduct post-incident reviews … based frameworks. Proficiency with scripting and automation (Bash, PowerShell, Python). Experience with observability tools (Phobos,Datadog, Prometheus, Grafana, OpenTelemetry, ELK). Hands-on experience with cloud platforms (AWS, Azure, or GCP). Strong PostgreSQL…