Logo

Site Reliability Engineer (SRE)

Job Description

Reliability Engineering & Operations

  • Drive the adoption of Site Reliability Engineering (SRE) principles, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), Error Budgets, monitoring, and alerting.
  • Design and implement automation solutions that improve operational efficiency, reduce manual effort, and increase service reliability.
  • Perform capacity planning, performance analysis, and system design to ensure platforms can scale effectively with business growth.
  • Lead troubleshooting efforts for complex production issues and perform root cause analyses (RCA) to prevent recurrence.
  • Participate in incident management processes, including major incidents, service outages, and on-call support when required.
  • Conduct post-incident reviews and drive corrective and preventive actions to continuously improve platform stability.
  • Leverage operational metrics and data-driven insights to proactively identify and mitigate potential risks.
  • Support engineering teams in adopting tools, processes, and best practices that improve reliability, maintainability, scalability, and extensibility.
  • Collaborate across Enterprise Platform teams to ensure standardized platforms meet enterprise-grade reliability and performance standards.
  • Champion the implementation of Non-Functional Requirements (NFRs), including availability, performance, security, and resilience, throughout the software development lifecycle.
  • Lead deep-dive investigations into problematic applications and services, identifying opportunities for architectural and operational improvements.
  • Oversee vulnerability management, platform lifecycle governance, and End-of-Life (EOL) remediation activities across supported products and services.

Leadership & Strategic Responsibilities

As a senior member of the team, you will:

  • Lead and coordinate strategic initiatives and key deliverables within the Platform SRE function.
  • Serve as an advocate for SRE principles and operational excellence across the Enterprise Platform organization and broader business community.
  • Review engineering practices, operational processes, and team outputs to drive continuous improvement.
  • Foster strong collaboration with development, infrastructure, security, and business teams.
  • Contribute to the long-term strategy, roadmap, and maturity of the Platform SRE function.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Software Engineering, or a related discipline.
  • Practical experience defining, implementing, and managing SLOs, SLIs, and Error Budgets.
  • Minimum of 5 years of experience supporting and operating production environments in SRE, DevOps, Platform Engineering, or related roles.
  • Experience or familiarity with monitoring and observability tools (e.g., Prometheus, preferably Grafana)
  • Professional fluency in English.
  • Proven ability to collaborate effectively within cross-functional teams.
  • Strong commitment to continuous learning, operational excellence, and continuous improvement.

Preferred Qualifications

  • Experience with Infrastructure as Code (IaC) and GitOps practices.
  • Knowledge of CI/CD platforms such as GitHub Actions, Jenkins, Azure DevOps, or GitLab.
  • Experience implementing Auto-Remediation and AIOps capabilities.
  • Understanding of security best practices, vulnerability management, and compliance frameworks.
  • AWS, Kubernetes, Terraform, or SRE-related certifications are advantageous.
  • Hands-on experience with observability, monitoring, logging, and alerting platforms such as Dynatrace, Datadog, Splunk, or similar tools.
  • Experience building automation solutions using technologies such as Python, Shell Scripting, PowerShell, Terraform, Ansible, Chef, Puppet, SQL, or equivalent.
  • Experience supporting middleware and platform technologies, including databases, web servers, messaging systems (MQ), and Kafka.
  • Familiarity with containerization and orchestration technologies such as Docker and Kubernetes.
Site Reliability Engineer (SRE)

business dinner 1

Exchanging ideas over a good meal always gets the best results.

Site Reliability Engineer (SRE)

contract renewal 1

At Blackfort, we value every client as an essential part of our growing family and trusted network. Each contract renewal is not just a milestone - it is a celebration of partnership, collaboration, and shared success. We also take this opportunity to recognize our team’s dedication and hard work, which make every renewed partnership possible.

Site Reliability Engineer (SRE)

contract renewal 2

Their smiles reflect the dedication, passion, and hard work we pour into every project. We are proud to share that our clients continue to place their trust in us — a testament to their satisfaction with the results we deliver and the value we create together.

Site Reliability Engineer (SRE)

team lunch

At Blackfort, we believe in making time for a little rest and recharge - even if it’s just over a shared meal.

Site Reliability Engineer (SRE)

business dinner 2

Another successful business planning session, made even more memorable by sharing a meaningful dinner together as one united team.

Site Reliability Engineer (SRE)

year end party

Here we are, celebrating the year's milestones and sending off good vibes as we welcome the new year together!