Buildkite

Staff ML Engineer

ANZ RegionFull timeStaffPosted 20 days ago
Apply on Buildkite โ†’

Sign into see who you know at Buildkite.

  Push a one-line fix. Then watch CI grind through forty minutes of tests, ninety-five percent of which never had a chance of touching what you changed. You already know the handful that mattered. The test suite doesn't โ€” so it runs everything, every time, just in case. That "just in case" is the most expensive habit in software delivery. Every engineering team pays it, because the alternative โ€” knowing which tests actually matter for a given change โ€” has been too hard to get right. We're building the team that gets it right. This role sits at the centre of it. ๐Ÿ”ง The problem worth solving Test Engine already ingests billions of test runs. We can see the tests, the code underneath them, and how the two move together โ€” at a scale very few people ever get to work with. The raw material for the answer is already here. Nobody's turned it into predictions yet. That's the step to take: for a given change, work out the slice of tests most likely to fail, and run only those. Get it right and teams stop re-running what hasn't changed, and spend that time where it counts โ€” like fixing the two percent of tests most likely to break. It's a genuinely difficult ML problem โ€” sparse signal, cold-start on new repos, generalising across languages and frameworks, and latency tight enough to sit in the critical path. It's also close to a blank page. There's no ML org above you setting the direction โ€” you'd set it. And not alone: we've just hired another ML engineer, so there's someone to think out loud with from day one. ๐Ÿš€ What you'll own Machine learning in Test Engine, end-to-end โ€” the strategy, the architecture, and the models running in production. That means shaping the whole path: pulling features out of code changes and test history, training and evaluating models, building the serving layer that keeps predictions fast, and closing the loop so the system keeps improving. You'd make the trade-offs that matter โ€” accuracy versus latency, what happens when confidence is low โ€” and build the platform underneath so the next model into production is quick and repeatable, not a one-off. โœจ The person we're picturing You've taken ML models the whole way โ€” from rough idea to something running reliably in production, monitored and retrained, owned rather than handed off. Two things matter more than any specific tool: You've built ML that generalised. Not one clever model โ€” a repeatable approach that worked across more than one use case. You're comfortable where the signal is noisy. Classification, ranking, prediction โ€” problems where the data doesn't hand you the answer. Day to day you'll live in Python and SQL, on AWS, with containerised workloads and data-at-scale tooling (Spark, Flink, or similar). Experience with code analysis, CI/CD systems, or ranking problems is a real head start โ€” a bonus, not a bar. The one thing we won't budge on: you've shipped and owned ML in production. Prototyped and handed off doesn't count here. ๐Ÿค” Is this you? You're likely a strong...