Platform Reliability Engineer
Apify is the largest marketplace of tools for AI. 40,000+ Actors helping people and agents get real-time web data, track competitors, generate leads, or integrate their apps. Actors are built by a global creator community that now earns more than $1.2 million every month.
Join us to help people put the web to work. Apify can find missing children https://blog.apify.com/fighting-child-traffickers-with-technology/, protect consumers from fake discounts across the EU https://blog.apify.com/how-web-scraping-ai-and-the-eu-have-come-together-to-sweep-away-fake-discounts-in-europe/, and feed data to AI chatbots https://blog.apify.com/intercom-customer-support-ai-chatbot-web-scraping/.
To support our mission, we're looking for a Platform Reliability Engineer with a developer's mindset. You've shipped code and you care what happens when it runs in production (speed, failures, recovery). You'll help us strengthen how Apify monitors systems, handle incidents, and route alerts so engineering teams can ship with confidence. You won't be on-call. This role is focused on sustainable improvement, not after-hours emergency response.
WHAT YOU'LL BE WORKING ON:
- Monitoring & signals: Operate and improve our monitoring stack (Prometheus, Grafana, OpenTelemetry) - instrument services to expose the right metrics, define what we watch in production, and shape alerting so teams get actionable signals without the noise.
- When things go wrong: Help define how we run incidents - clear communication, structured learning afterward, and supporting artifacts (status page, runbooks).
- With the team: Work with platform and product engineers to make reliability standards practical - help teams adopt better tooling or practices when things change, and write documentation people actually use.
WHO WE'RE LOOKING FOR:
- You have hands-on experience choosing what to measure in production - not just reading dashboards, but picking signals that reflect the customer experien...