Join a team at the heart of the global economy! The Department for Business and Trade (“DBT”) and Inspire People are partnering together to bring you an exciting opportunity for a Senior Site Reliability Engineer (with either strong Python skills or
About DBT
The Department for Business and Trade (DBT) has a clear mission – to grow the economy. Our role is to help businesses invest, grow and export to create jobs and opportunities right across the country. We do this in three ways.
Firstly, we help to build a strong, competitive business environment, where consumers are protected and companies rewarded for treating their employees properly.
Secondly, we open international markets and ensure resilient supply chains. This can be through Free Trade Agreements, trade facilitation and multilateral agreements.
Finally, we work in partnership with businesses every day, providing advance, finance and deal-making support to those looking to start up, invest, export and grow.
The Digital, Data and Technology (DDaT) directorate develops and operates tools and services to support us in this mission.
As a Senior Site Reliability Engineer (DevOps), you will play a key role in designing, building and running reliable, scalable and secure platform services that underpin critical DBT digital products. Working in multidisciplinary, agile teams, you’ll help ensure development teams have the tools and support they need, from observability and monitoring through to CI/CD pipelines so services are resilient, performant and centred around user needs. You’ll champion good engineering practices, helping teams use data and insight to continuously improve how services are delivered and operated.
This is a hands-on role where you’ll spend much of your time building and improving platform capabilities directly. You’ll work closely with product managers, architects and engineers to improve reliability and reduce operational burden, while also supporting and mentoring others across the engineering community. You’ll contribute to building and scaling our global platform, support live services through an on-call rota, and play an active role in initiatives such as enhancing observability and streamlining deployment processes to improve service quality and delivery outcomes.
Main responsibilities
You will:
- Build and maintain shared service products enabling developers to be more efficient.
- Write clean, maintainable and well-tested code to support platform and tooling development.
- Work closely with development teams to provide and improve platform tooling, including monitoring, logging, metrics, dashboards and CI/CD pipelines.
- Build, maintain and improve reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches.
- Support teams to adopt SRE practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets.
- Contribute to observability across services, helping teams better understand performance, reliability and user impact.
- Develop and improve CI/CD pipelines to enable safe, frequent and low-risk delivery of changes.
- Collaborate with product, delivery and architecture colleagues to ensure platform services meet user and business needs.
- Support live service operations, including incident response, troubleshooting and problem management, with a focus on learning and continuous improvement.
- Share knowledge and provide coaching and mentoring to colleagues, contributing to a supportive and inclusive engineering community.
- Contribute to improving security, resilience and compliance practices within the platform.
What tech will you be using?
- Python and Django framework
- PostgreSQL as a service (Amazon RDS)
- Redis/Elasticache
- AWS and Azure
- GitHub Actions and AWS CodePipelines/CodeBuild
- Terraform
- Docker, Elastic Container Service (ECS) and Elastic Container Registry (ECR)
- ElasticSearch/OpenSearch
- Datadog and Logstash
Essential Skills:
- Experience writing clean, maintainable code in at least one programming language.
- Experience designing, operating and improving distributed systems, with a focus on reliability, performance and user impact.
- Experience of building applications hosted on cloud platforms such as AWS, Azure or Google Cloud.
- Ability to build code-defined, reliable and well-tested infrastructure using tools such as Terraform, CloudFormation or similar.
- Knowledge of Linux/Unix fundamentals and TCP/IP networking.
- Strong communication skills, with the ability to build effective working relationships with both technical and non-technical stakeholders.
How we interview
At the interview stage for this role, you will be asked to demonstrate relevant Technical Skills and Behaviours from the Success Profiles framework. A role-specific list of these can be found below.
The technical element within the interview, where you will be asked a series of questions to demonstrate your specific professional skills and knowledge related directly to the job role and context, will assess against capabilities which are outlined under DevOps engineer within the DDaT framework which can be found here. As part of this process, you will be asked to complete a technical problem-solving exercise during your interview. Further details will be provided following sift.
Technical Skills
- Availability and capacity management
- Development process optimisations
- Information security
- Modern standards approach
- Programming and build (software engineering)
- Prototyping (with MVP and POC)
- Systems integration
- User focus
Behaviours
- Communicating and Influencing
- Developing Self and Others
How we Offer
Offers will be made in merit order based on location preferences. If you pass the bar at interview but are not the highest scoring you will be held on a 12-month reserve list in case a role becomes available. If you are judged a near miss at interview, you may be offered a post at the grade below the one you applied for.
This role requires SC clearance. DBT’s requirement for SC clearance is to have been present in the UK for at least 3 of the last 5 years. Failure to meet this requirement will result in your application being rejected and your offer will be withdrawn.
Checks will also be made against:
- departmental or company records (personnel files, staff reports, sick leave reports and security records)
- UK criminal records covering both spent and unspent criminal records
- your credit and financial history with a credit reference agency
- security services record
- location details
Benefits
If you join DBT, you will get:
- learning and development tailored to your role
- a flexible, hybrid working environment with options like condensed hours
- a culture encouraging inclusion and diversity
- a Civil Service pension with an average employer contribution of 28.97%
- annual leave starting at 25 days rising to 30 days with service
- three paid volunteering days a year
- an employee benefits programme including cycle to work
More about DBT
This role can only be worked from within the UK, not overseas. If you are based in London, you will receive London weighting. DBT employees work in a hybrid pattern, spending 2-3 days a week (pro rata) in the office on average. Travel to your primary office location will not be paid for by DBT, but costs for travel to an office which is not your main location will be covered.
You can find out more about our office locations, how we calculate salaries, our diversity statement and reasonable adjustments, the Recruitment Principles, the Civil Service code and our complaints procedure on our website.
Find out more about life at DBT, our benefits and meet the team by watching our video or reading our blog!
If you are a DevOps Engineer, Site Reliability Engineer or Linux/Unix Systems Administrator looking to enhance your career and make a difference across an expanding function, then apply today or contact Keesha Paulsen at Inspire People in complete confidence for further information.
.

