Site Reliability Engineer

Aisle · New York, NY
LinkedIn

Posted

Aug 21, 2026 (8d ago)

Seniority

Not Specified

Work Model

Not Specified

Type

Contract

Category

DevOps & SRE

Salary

Not specified

Skills

CI/CD Datadog GCP Google Cloud Kubernetes LLM Next.js Node.js Observability PostgreSQL Pub/Sub React Redis Serverless SRE TypeScript Vercel

Description

What is Aisle? Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened. Aisle builds a real-time control layer on top of physical retail. We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact. This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time. Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar. Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment. We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis. You'll win here if… You enjoy communicating across technical and non-technical teams looking to build shared understanding that can turn into collaborative actio Can solve problems in stages with appropriate level of urgency when needed: immediate triage, temporary patch, and long term fix that increases durability of the system You’re comfortable moving quickly, shipping improvements, and iterating in production AI is a core part of how you work - you’ve gone well beyond basic copilots and actively use AI tools to design, debug, and ship systems faster You treat AI agents as chaotic, untrusted components - your instinct is isolation, rate limiting, and permissions scoping, etc. (gist of harness engineering) You have experience building or orchestrating AI/agent workflows - in production or through serious side projects About your role Reliability & Infrastructure Own the reliability, scalability, and observability of our infrastructure across GCP and Vercel Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards Manage IAM, service accounts, and security best practices across our cloud environment Participate in on-call rotation, incident response, and post-mortems — and turn learnings into systemic improvements Core Infrastructure & State Stabilize core infrastructure and stateful systems under heavy concurrent load from both consumer traffic and agent-driven workflows Build and maintain our event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deployments Investigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolution Stabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows) Automate infrastructure provisioning, deployments, and operational workflows AI & Next-Gen Tooling: Agent Ops Build agent operations infrastructure that enables AI agents to run safely and reliably in production Develop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetry Help define CI/CD for an agent-driven engineering environment - including how code gets shipped, reviewed, validated, and deployed when agents are in the loop Own visibility into AI usage, reliability, and spend as our agent footprint scales Cross-functional Impact Partner closely with engineering and product teams to maintain reliability without slowing development velocity Act as a force multiplier across the team — helping engineers ship faster and more safely About your skillsMust haves 4+ years in SRE, DevOps, or infrastructure/platform engineering Strong, hands-on experience with a major cloud platform (preferable GCP) Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.) Solid understanding of IAM, security, and cloud best practices Experience with observability tools like Datadog Familiarity with Node.js environments AI is part of your daily engineering workflow you use AI tools thoughtfully to improve engineering velocity, reliability, and operational excellence beyond basic code generation. Experience building or orchestrating AI/agent workflows (work or serious personal projects) High ownership, strong curiosity, and a bias toward action Nice to haves Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controls Familiarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns is a plus Hands-on experience with GCP Workflows for orchestration Familiarity with Prisma , pgbouncer , and PostgreSQL connection pooling Experience with Vercel deployment and edge computing Familiarity with the k8s ecosystem Familiarity with Redis and BullMQ Understanding of SOC 2 compliance requirements and implementation Previous experience in a high-growth startup environment Previous backend engineering experience to bridge the gap between infrastructure and code About the stack Cloud: Google Cloud Platform (GCP) Infrastructure: Cloud Functions, Pub/Sub, Workflows, Kubernetes Database: PostgreSQL with pgbouncer, Prisma Observability: Datadog Runtime: Node.js, TypeScript Deployment: Vercel, GCP Frontend: React, Next.js, TypeScript Backend: TypeScript, PostgreSQL, Next.js