Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert
RaceTrac Company OverviewJob Description:We are seeking a highly experienced Site Reliability Engineer (SRE) with deep expertise in Dynatrace, observability engineering, and Azure cloud technologies. This role will be exclusively focused on building, enhancing, and managing enterprise observability, telemetry, monitoring, and proactive reliability engineering practices across critical digital platforms.The ideal candidate must possess advanced hands-on expertise in Dynatrace, especially Dynatrace Query Language (DQL), along with strong knowledge of Azure Monitor, Azure KQL, Application Insights, Azure Functions, APIM, and distributed telemetry concepts. The candidate should have a strong understanding of .NET application architecture and the ability to read and analyze .NET code to support troubleshooting, root cause analysis, and observability implementation within Azure environments. Experience enabling observability for mobile platforms such as iOS and Android is also required.This is a highly technical, hands-on role requiring a proactive engineering mindset, strong analytical capabilities, and the ability to collaborate across engineering, cloud, mobile, and business teams.What You'll DoDynatrace & Observability EngineeringServe as the primary Dynatrace SME across the organization.Design, develop, and optimize enterprise observability solutions using Dynatrace.Develop advanced Dynatrace DQL queries, dashboards, workflows, alerts, and analytics.Implement intelligent monitoring strategies for applications, APIs, integrations, Azure services, mobile platforms, and distributed systems.Continuously improve observability maturity through telemetry standardization, proactive monitoring, and automation.Configure and tune alerting mechanisms to improve signal-to-noise ratio and reduce alert fatigue.Leverage Dynatrace Davis AI, anomaly detection, and AI-driven root cause analysis capabilities.Enable and enhance observability for mobile applications across iOS and Android platforms.Azure Monitoring & Cloud OperationsBuild and maintain monitoring solutions using:Azure MonitorApplication InsightsAzure Log AnalyticsAzure KQLMonitor and troubleshoot Azure Function Apps, App Services, APIs, integrations, and backend services.Analyze telemetry, traces, logs, metrics, and distributed transactions to identify root causes and performance bottlenecks.Troubleshoot cloud-native applications and Azure infrastructure issues.Develop proactive monitoring for cloud services, integrations, APIs, and backend processing systems.API & Integration MonitoringMonitor and troubleshoot Azure API Management (APIM), API Gateways, API endpoints, and integrations.Understand end-to-end API transaction flows and dependency mapping.Build observability solutions for APIs, middleware platforms, and integration services.Diagnose latency issues, transaction failures, authentication issues, and backend service degradation.Mobile Application ObservabilityEnable telemetry, mon...