AWS SRE
Job Description:
Role: AWS Site Reliability Engineer (SRE) - Data Platform
Location: Remote (UK-based)
Contract Length: Initial 3 months (with potential for extension)
IR35 Status: Inside IR35
Clearance: SC Clearance preferred
About the Role
We are looking for an experienced AWS Site Reliability Engineer (SRE) to join our team on an initial 3-month contract. You will be Embedded within a high-impact team dedicated to ensuring the reliability, scalability, and performance of our AWS-hosted Data Platform.
If you live and breathe observability, love tearing down operational toil through automation, and know how to keep complex cloud ecosystems running smoothly, we want to hear from you.
Key Responsibilities
- Define and Operationalise Reliability: Establish, refine, and operationalise SLIs, SLOs, and error budgets for critical data services, mapping them to the four golden signals (latency, errors, traffic, and saturation).
- Observability Frameworks: Build and maintain comprehensive SLO dashboards and end-to-end monitoring (metrics, logs, and traces) utilizing Dynatrace and Prometheus.
- Cloud and Container Management: Navigate the AWS ecosystem confidently, managing and optimizing containerized workloads deployed on Amazon EKS (Kubernetes).
- Toil Reduction and Automation: Drive aggressive automation initiatives to eliminate repetitive operational tasks and streamline system efficiency.
- Collaboration and Resilience: Partner closely with developers and architects to improve architecture reliability, contribute to continuous improvement backlogs, and lead root cause analysis (RCA) via blameless post-mortems.
Technical Skills and Experience
Required Expertise
- Strong background as an SRE or DevOps Engineer within an AWS environment.
- Hands-on experience managing and scaling workloads on Amazon EKS (Kubernetes).
- Proven track record with observability stacks, specifically Dynatrace and Prometheus.
- Deep understanding of SRE principles, including error budgets, alerting thresholds, and full-stack tracing.
- Excellent Scripting/automation skills (eg, Python, Bash, or Go).
- Data Platform Experience: Prior exposure to data platforms, batch/streaming data pipelines (eg, Kafka, Spark), and the unique challenges of data observability and workload reliability.
- Active or recent SC Clearance.
£300.00 - £375.00/day
Talent International UK and it's subsidiaries, Digital Gurus, Infinite Talent and Rethink act as an employment agency for permanent recruitment and employment business for the supply of temporary workers. By applying for this opportunity, you accept the TandC's, Privacy Policy and Disclaimers which can be found on our website
Reference: 3142444043
AWS SRE
Posted on Jul 21, 2026 by Talent International
Job Description:
Role: AWS Site Reliability Engineer (SRE) - Data Platform
Location: Remote (UK-based)
Contract Length: Initial 3 months (with potential for extension)
IR35 Status: Inside IR35
Clearance: SC Clearance preferred
About the Role
We are looking for an experienced AWS Site Reliability Engineer (SRE) to join our team on an initial 3-month contract. You will be Embedded within a high-impact team dedicated to ensuring the reliability, scalability, and performance of our AWS-hosted Data Platform.
If you live and breathe observability, love tearing down operational toil through automation, and know how to keep complex cloud ecosystems running smoothly, we want to hear from you.
Key Responsibilities
- Define and Operationalise Reliability: Establish, refine, and operationalise SLIs, SLOs, and error budgets for critical data services, mapping them to the four golden signals (latency, errors, traffic, and saturation).
- Observability Frameworks: Build and maintain comprehensive SLO dashboards and end-to-end monitoring (metrics, logs, and traces) utilizing Dynatrace and Prometheus.
- Cloud and Container Management: Navigate the AWS ecosystem confidently, managing and optimizing containerized workloads deployed on Amazon EKS (Kubernetes).
- Toil Reduction and Automation: Drive aggressive automation initiatives to eliminate repetitive operational tasks and streamline system efficiency.
- Collaboration and Resilience: Partner closely with developers and architects to improve architecture reliability, contribute to continuous improvement backlogs, and lead root cause analysis (RCA) via blameless post-mortems.
Technical Skills and Experience
Required Expertise
- Strong background as an SRE or DevOps Engineer within an AWS environment.
- Hands-on experience managing and scaling workloads on Amazon EKS (Kubernetes).
- Proven track record with observability stacks, specifically Dynatrace and Prometheus.
- Deep understanding of SRE principles, including error budgets, alerting thresholds, and full-stack tracing.
- Excellent Scripting/automation skills (eg, Python, Bash, or Go).
- Data Platform Experience: Prior exposure to data platforms, batch/streaming data pipelines (eg, Kafka, Spark), and the unique challenges of data observability and workload reliability.
- Active or recent SC Clearance.
£300.00 - £375.00/day
Talent International UK and it's subsidiaries, Digital Gurus, Infinite Talent and Rethink act as an employment agency for permanent recruitment and employment business for the supply of temporary workers. By applying for this opportunity, you accept the TandC's, Privacy Policy and Disclaimers which can be found on our website
Reference: 3142444043
Alert me to jobs like this:
Amplify your job search:
Expert career advice
Increase interview chances with our downloads and specialist services.
Visit Blog