Description
Client type: Corporate/Enterprise/Radiology
Client location: USA (EST time)
Project technology stack: AWS (ECS, ALB, IAM, VPC, CloudWatch, S3), Terraform, SNS/SQS, Cloudflare, GitHub Actions, PostgreSQL, Redis/BullMQ, Node.js/NestJS, Datadog, TypeScript, React, SQL
Position: Senior Site Reliability Engineer
Workload: 100% (Full-time)
Project start date: ASAP
Engagement period: Ongoing
Candidate location: EU / remote
Language: English C1Â
Interview timeline: ASAP
Interview process: 1) CV review 2) Interview with our CTO 3) Interview with end clientÂ
Number of interviews: 2
About Our Client
Our client is a leading veterinary teleradiology provider in the US, offering expert interpretation of X-rays, CT, MRI, and ultrasound studies. The company delivers fast, reliable diagnostic reports from board-certified radiologists to support veterinarians in clinical decision-making.
Their mission is to provide veterinary practices with fast, expert, and affordable diagnostic imaging interpretations, ensuring high-quality radiology support whenever needed. The company's platform leverages computer vision and machine learning to help radiologists make fast, informed decisions
Role Overview
We are looking for a Senior Engineer to join the infrastructure and platform team supporting this teleradiology platform. This is a hands-on, ongoing engagement combining cloud infrastructure ownership (AWS, Terraform, messaging, observability) with the ability to read the codebase and patch small defects across the stack. The role sits at the intersection of platform reliability and product engineering support.
TECH STACK
Infrastructure & Platform
- AWS: ECS (task definitions, services, autoscaling, rolling deploys), ALB, IAM, VPC, CloudWatch, S3, Secrets Manager / Parameter Store
- Terraform (Infrastructure as Code)
- SNS and SQS: fan-out topics, queue policies, visibility timeouts, redrive policies and DLQs, message replay
- Cloudflare: DNS, TLS, WAF, caching, rate limiting
- CI/CD pipelines, GitHub Actions
- Secrets management, least-privilege IAM, patching discipline
Data & Runtime
- PostgreSQL operations: replication, PITR backups, connection pooling, query and lock diagnosis
- Redis and BullMQ: backlog, retries, DLQ handling
- Node.js runtime operations and debugging (NestJS apps, memory and event-loop issues)
Observability & Reliability
- Datadog: APM, log pipelines, monitors, dashboards, synthetics, ECS container metrics
- SLOs and error budgets
- Incident management and on-call practice
- Debugging in production: reading traces and logs back to a specific line
Development Capability (fixes small bugs)
- TypeScript and NestJS — well enough to read the codebase, trace a request end to end, and patch small defects
- React basics for frontend fixes (state, data fetching, browser devtools)
- SQL and the ORM/query layer, plus writing a safe migration (expand / contract)
- Git workflow, code review, and the team's testing setup so fixes ship with coverage
Valuable, Teachable on the Job
- DICOM/PACS and HL7 fundamentals
- Auth0
- Healthcare and veterinary data retention and audit expectations
Scope Boundaries
- Bug fixes are bounded to a single service; no schema or contract changes without product engineering.
- Same PR review standard as product engineers.
- The team's ownership of integration support (per-clinic PACS/PIMS issues) versus a separate ops function will be decided up front.
Please note: Background and reference checks will be conducted as part of the selection process. Â
Want to know more about Devspace?
We build software teams for companies across Europe and the US. Have a look at what we do.