Senior Site Reliability Engineer

Status
Open
Headcount
2/2 open
Level
Senior Individual Contributor
Location
USA (EST time)
Language
English C1
Duration
12 months
Candidate Location
EU (remote)
Project start
9/4/2026

Description

Client type: Corporate/Enterprise/Radiology

Client location: USA (EST time)

Project technology stack: AWS (ECS, ALB, IAM, VPC, CloudWatch, S3), Terraform, SNS/SQS, Cloudflare, GitHub Actions, PostgreSQL, Redis/BullMQ, Node.js/NestJS, Datadog, TypeScript, React, SQL

Position: Senior Site Reliability Engineer

Workload: 100% (Full-time)

Project start date: ASAP

Engagement period: Ongoing

Candidate location: EU / remote

Language: English C1 

Interview timeline: ASAP

Interview process: 1) CV review 2) Interview with our CTO 3) Interview with end client 

Number of interviews: 2

About Our Client

Our client is a leading veterinary teleradiology provider in the US, offering expert interpretation of X-rays, CT, MRI, and ultrasound studies. The company delivers fast, reliable diagnostic reports from board-certified radiologists to support veterinarians in clinical decision-making.

Their mission is to provide veterinary practices with fast, expert, and affordable diagnostic imaging interpretations, ensuring high-quality radiology support whenever needed. The company's platform leverages computer vision and machine learning to help radiologists make fast, informed decisions

Role Overview

We are looking for a Senior Engineer to join the infrastructure and platform team supporting this teleradiology platform. This is a hands-on, ongoing engagement combining cloud infrastructure ownership (AWS, Terraform, messaging, observability) with the ability to read the codebase and patch small defects across the stack. The role sits at the intersection of platform reliability and product engineering support.

TECH STACK

Infrastructure & Platform

  • AWS: ECS (task definitions, services, autoscaling, rolling deploys), ALB, IAM, VPC, CloudWatch, S3, Secrets Manager / Parameter Store
  • Terraform (Infrastructure as Code)
  • SNS and SQS: fan-out topics, queue policies, visibility timeouts, redrive policies and DLQs, message replay
  • Cloudflare: DNS, TLS, WAF, caching, rate limiting
  • CI/CD pipelines, GitHub Actions
  • Secrets management, least-privilege IAM, patching discipline

Data & Runtime

  • PostgreSQL operations: replication, PITR backups, connection pooling, query and lock diagnosis
  • Redis and BullMQ: backlog, retries, DLQ handling
  • Node.js runtime operations and debugging (NestJS apps, memory and event-loop issues)

Observability & Reliability

  • Datadog: APM, log pipelines, monitors, dashboards, synthetics, ECS container metrics
  • SLOs and error budgets
  • Incident management and on-call practice
  • Debugging in production: reading traces and logs back to a specific line

Development Capability (fixes small bugs)

  • TypeScript and NestJS — well enough to read the codebase, trace a request end to end, and patch small defects
  • React basics for frontend fixes (state, data fetching, browser devtools)
  • SQL and the ORM/query layer, plus writing a safe migration (expand / contract)
  • Git workflow, code review, and the team's testing setup so fixes ship with coverage

Valuable, Teachable on the Job

  • DICOM/PACS and HL7 fundamentals
  • Auth0
  • Healthcare and veterinary data retention and audit expectations

Scope Boundaries

  • Bug fixes are bounded to a single service; no schema or contract changes without product engineering.
  • Same PR review standard as product engineers.
  • The team's ownership of integration support (per-clinic PACS/PIMS issues) versus a separate ops function will be decided up front.

Please note: Background and reference checks will be conducted as part of the selection process.  

Want to know more about Devspace?

We build software teams for companies across Europe and the US. Have a look at what we do.

Visit devspace.no