Description
AWS Senior Data Engineer
Client: Leading Life Sciences & Technology Company (Corporate/Enterprise, USA – EST) Employment Type: Full-time (100%) Location: Remote (EU-based candidates) Start Date: ASAP Engagement: Ongoing Language Requirement: English C1
About the Client
Our client is a leading life sciences and technology company supplying analytical instruments, laboratory equipment, reagents, software, and services used in scientific research, healthcare, biotechnology, and pharmaceutical manufacturing worldwide. Their mission is to help customers make the world healthier, cleaner, and safer — supporting breakthroughs from life-sciences research and clinical diagnostics to complex analytical challenges and drug development.
Their software portfolio includes enterprise lab informatics platforms such as Laboratory Information Management Systems (LIMS), Electronic Lab Notebooks (ELN), and data management tools that help laboratories manage, analyze, and automate data and workflows.
Role Overview
We're looking for a Senior AWS Data Engineer to design, build, and maintain robust data pipelines and cloud-based data platforms supporting enterprise lab informatics products. This role involves end-to-end ownership of ETL/data ingestion pipelines, cloud infrastructure automation, and increasingly, support for AI/ML and RAG-based initiatives leveraging proprietary enterprise data.
Core AWS Stack: Glue, Lambda, S3, Athena, RDS, IAM, VPC, EC2, CloudWatch, CloudTrail, Step Functions
Programming: Python, PySpark
Data Architecture: Data Lakes, S3-based data lake patterns, Apache Iceberg, Apache Parquet, PostgreSQL
DevOps/IaC: GitHub Actions or Jenkins (CI/CD), Terraform or CloudFormation
Emerging Focus: RAG (Retrieval-Augmented Generation) & Agentic Workflows for LLMs
Governance: Data validation, PII detection, privacy & compliance controls
Experience Level: 8–10 years in Data Engineering / Data Lake development
Key Responsibilities
- Implement robust ETL pipelines using AWS Glue, defining extraction methods, transformation logic, and load procedures across diverse data sources.
- Assess application data requirements and build AWS Lambda-based solutions for efficient data integration, processing, and application support.
- Orchestrate jobs using AWS Step Functions and Lambda.
- Implement event-driven pipelines for various input formats.
- Support CI/CD using GitHub Actions or Jenkins.
- Automate infrastructure using Terraform or CloudFormation.
- Monitor pipeline execution using CloudWatch and CloudTrail.
Must-Have Skills
Critical / Non-Negotiable
- 8–10 years of experience developing Data Lakes with ingestion from disparate sources (relational databases, flat files, APIs, streaming data)
- Strong AWS data engineering experience: Glue, Lambda, S3, Athena, RDS, IAM, VPC, EC2, CloudWatch, CloudTrail
- Proficiency in Python and PySpark for efficient data processing
- Hands-on experience implementing ETL pipelines with AWS Glue (extraction, transformation, load logic)
- Experience with S3-based data lake patterns
Data Architecture & Modeling
- Design and development of Data Platforms and ingestion pipelines from disparate sources into the cloud
- Working knowledge of PostgreSQL and Apache Iceberg
- Experience with Apache Parquet and Apache Iceberg as table formats for analytical datasets
- Data modeling experience for relational databases
Infrastructure & Operations
- Ability to collaborate with the Infrastructure team for AWS service provisioning (databases, IAM roles)
- Ability to work with AWS support for issue resolution
- CI/CD experience using GitHub Actions
- Terraform or CloudFormation experience
AI/ML & Emerging Tech
- Solid understanding of RAG (Retrieval-Augmented Generation) and Agentic Workflows for providing enterprise data context to LLMs
- Strong analytical and problem-solving skills for data and AI/ML engineering use cases
- Familiarity with data observability tools to ensure pipeline and ML model quality/accuracy
Governance & Compliance
- Good understanding of data validation, error handling, and audit logging
- Knowledge of data classification, PII detection, privacy, and compliance controls
- Experience with data privacy and compliance requirements, especially related to PII data
Interview process: 1) CV review 2) Interview with Devspace manager 3) Interview with end client
Number of interviews: 2
Please note: Background and reference checks will be conducted as part of the selection process.
Want to know more about Devspace?
We build software teams for companies across Europe and the US. Have a look at what we do.