<p>We are looking for an experienced Site Reliability Engineer (SRE) to strengthen observability and operational resilience across a Microsoft Azure environment. This long-term Contract role will work closely with DevOps and engineering teams to establish monitoring standards, expand telemetry coverage, and improve service reliability across cloud-based platforms. The ideal candidate brings deep expertise in Azure operations, modern observability tooling, and production support, with the ability to turn data into actionable insight for faster troubleshooting and stronger system performance.</p><p><br></p><p>Responsibilities:</p><p>• Create and advance an observability framework for Azure-hosted systems and integrated third-party platforms, ensuring scalable monitoring coverage.</p><p>• Develop meaningful dashboards, alerting rules, log analysis views, and distributed tracing to provide actionable insight into application and infrastructure behavior.</p><p>• Utilize Azure services such as Azure Monitor, Log Analytics, Application Insights, Managed Prometheus, and Azure Managed Grafana to expand end-to-end visibility.</p><p>• Work alongside DevOps and software engineering teams to strengthen platform stability, incident response readiness, and service performance.</p><p>• Assess existing monitoring practices to uncover blind spots, reduce unnecessary alert volume, and support quicker root-cause identification.</p><p>• Improve insight into the health of applications, infrastructure components, and dependent services across production environments.</p><p>• Support reliability-focused engineering efforts by applying SRE principles such as service measurement, alert strategy refinement, and operational readiness improvements.</p>