Site Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering to build and run large-scale, data & product, fault-tolerant systems. SRE ensures that dunnhumbys services—both our internally critical and our externally-visible systems—have reliability and uptime appropriate to users’ needs and a fast rate of improvement while keeping an ever-watchful eye on capacity and performance.
SRE is also a mindset and a set of engineering approaches to running better production systems—we build our own creative engineering solutions to operations problems. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to solve a broad spectrum of problems. Practices such as limiting time spent on operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting and dynamic day-to-day work.
SRE’s culture of diversity, intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.
- Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.
- Support services before they go live through activities such as system design consulting, developing analytics platforms and frameworks, capacity planning and launch reviews.
- Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
- Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity.
- Practice sustainable incident response and blameless postmortems.
- Interest in designing, analyzing and troubleshooting large-scale data & product platforms.
- Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive.
- Ability to debug and optimize solutions, automate routine tasks.
- BS degree or Technical related practical experience with excellent problem solving skills
- Work as a team with others
- Build strong relationships with colleagues and Technical Strategic Partners
- Drive and identify opportunities
- Identify personal and business priorities
- Demonstrate enthusiasm
- Deliver work to high standard
- Solve challenging problems
General Functional Skills:
- Business and Commercial Acumen: E
- Project Management: E
- Engineering & Design: E
- Stakeholder Management: E
- Financial Analysis: E
- dunnhumby Capabilities and Solutions: E
Sub Family Skills:
- Hosted platforms (Data Centres & Cloud): I
- Storage & Virtualozation: I
- Linux, Storage, Containers & orchestration: I
- Database, Hadoop: I
- Network & Security: I
What we offer:
Duration: 18 months
Working week: 37.5 hours Monday – Friday
Role location: 184 Shepherds Bush Rd, Hammersmith, London W6 7NL