Site Reliability Engineer

Open Systems Technologies · Montreal, Quebec, Canada
LinkedIn

Posted

Aug 14, 2026 (Aug 14)

Seniority

Not Specified

Work Model

Not Specified

Type

Not Specified

Category

DevOps & SRE

Salary

Not specified

Skills

AWS Azure CI/CD Java Kubernetes Linux Observability

Description

Responsibilities: Operate and maintain the Kubernetes based platform across public and private cloud environments. Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation. Manage incidents from end-to-end, including incident response, root cause analysis, and post incident reviews. Onboard and support new clients onto the platform. Build automation and diagnostic tooling that cuts manual effort and evaluates performance. Identify and deliver process improvements. Support software and hardware upgrades and keep components up to date. Proactively manage capacity so the platform scales with demand. Required Skills: Hands on experience operating Kubernetes in production, ideally with Service Mesh. Strong Linux and command line fundamentals. Confident debugging & troubleshooting complex systems, from the application layer through to lower-level infrastructure. Experience working with a public cloud provider, preferable Azure or AWS. Working knowledge of Grafana, Prometheus, Loki and Tempo is a plus. Scripting or coding in Python or Java is a strong plus. CI/CD, infrastructure as code such as Helm or Terraform is a plus A financial services background is not required.