15 Sep 2026

Site Reliability Engineer

Work from home

Agilus by Synergie is recruiting for a Site Reliability Engineer in the Financial Services.
Seeking an experienced Application Manager (Site Reliability Engineer) that is responsible for the day-to-day reliability, health, and operational sustainability of assigned applications, data pipelines, platforms, and services within the applicable Business Technology – specifically aligned to support our Capital Markets and Total Fund Market Investment Analytics Team.

Responsibilities:
  • Key Responsibilities Operate and continuously improve the reliability, supportability, and operational efficiency of assigned Product Team including applications, data pipelines, platforms and services, including availability, latency, performance, and resilience. Implement and operationalize Service Level Availability (SLAs) and related reliability metrics in accordance with standards and recommend changes and improvements when applicable. Develop and maintain application, data pipeline, platform, and service health plans focused on reducing operational risk and improving sustainability and contribute identified priorities to the portfolio reliability roadmap Lead incident response for assigned services and ensure timely stakeholder communication escalating high-impact and cross-service events as required. Conduct Root Cause Analysis for key incidents, identify corrective and preventative actions, and track them through completion. Improve Mean Time to Detect (MTTD) and Mean Time to Restore (MTTR) through monitoring, diagnostics, runbooks, and operational learning. Perform capacity planning, load testing, and performance analysis to prevent incidents before they occur. Contribute to Disaster Recovery (DR) and resilience plans and participate in recovery testing. Conduct production readiness reviews and identify supportability gaps prior to go-live. Partner with Product Engineering, Data Solutions, and Data Platform, Architecture, Security, and Technology Services teams, as appropriate, during design reviews to identify reliability, resilience, and supportability risks. Validate that monitoring, resilience testing, runbooks, operational documentation and support arrangements are in place before release. Monitor vendor services against defined SLAs and escalate performance issues through the appropriate vendor and portfolio governance channels. Apply established source-control, testing, change-management, release, rollback, documentation, and production-validation practices to changes implemented by the SRE function

Successful candidates will have:
  • 6+ years of experience in Site Reliability Engineering, production engineering, platform engineering, application support, or technology service delivery roles. Strong understanding of application, data, and platform architectures, cloud platforms, and enterprise systems.? Strong collaboration, communication, and influencing skills, with a demonstrated ability to coordinate successful outcomes across multiple teams.? Hands-on experience operating and improving highly available production systems. Experience leading incident response and root cause analysis. Practical experience using approved AI-assisted engineering tools, including IDE-integrated assistants and Model Context Protocol (MCP)-enabled integrations, to troubleshoot applications, databases, code, and infrastructure; analyze logs and telemetry; develop and validate queries, scripts, and remediation options; and improve the speed, consistency, and quality of incident diagnosis and service restoration. Experience building or maintaining observability and automation solutions; scripting experience with Python or PowerShell is preferred. Experience supporting message queue systems, APIs, and distributed service architectures is an asset. Expertise in delivery methodologies (Agile, Waterfall, DevOps) and IT service management frameworks (ITIL, COBIT).? Experience with CI/CD pipelines, infrastructure-as-code, and DevOps/DataOps practices. Experience with JIRA, Confluence, Git Strong collaboration skills with architects, business analysts, DBAs, and QA/test peers.

Total rewards:
  • $70.00 - $95.00 hourly
  • 9 month contract to start

Soft skills:
  • Strong written and verbal communication skills required

This posting is for an active opening.

Interested?
Please apply directly online

Agilus by Synergie would like to thank all candidates for their interest in this opportunity. Due to the volume of resumes we receive; we may only be able to respond directly to those candidates being selected for an interview.
We encourage you to visit agilus.ca regularly or subscribe to our email alerts at agilus.ca/Account/Register as new exciting employment opportunities become available daily.

Job Code 157737

We value a diverse workplace.

Field Data Capture (PM Suite / PVR)

  • Calgary
  • Contract
  • May 01, 2026
View job posting

Software Engineer

  • Toronto
  • Permanent
  • June 16, 2026
View job posting

Senior SDET / QE Platform Engineer

  • Toronto
  • Contract
  • June 16, 2026
View job posting

Senior Data Scientist

  • Toronto
  • Permanent
  • June 19, 2026
View job posting

Senior Data Engineer

  • Toronto
  • Permanent
  • June 19, 2026
View job posting

Senior Full Stack Developer

  • Toronto
  • Permanent
  • June 19, 2026
View job posting