BEES

Senior Reliability Engineer

Campinas, São Paulo, BrazilFull timeSeniorPosted 24 days ago
Apply on BEES →

Sign into see who you know at BEES.

About BEES  Join us to build the future of B2B commerce!  BEES is AB InBev’s B2B platform. Through our ecosystem, merchants and retailers across 29 countries can stock their businesses quickly, easily, and securely. At BEES, we dream big, lead with purpose, and develop technology that transforms the way retailers and sellers grow.  Every line of code and every partnership is built in service of a single mission: to make commerce better for retailers and sellers around the world. Here, your work is not just important. It makes a difference!  What you will do: Architect and design highly available, fault-tolerant, and scalable infrastructure and systems. Collaborate with software development teams to influence design decisions and ensure reliability and scalability from the ground up. Develop and implement best practices, guidelines, and standards for SRE processes, infrastructure, and automation. Define and enforce service-level objectives (SLOs) and error budgets to drive system reliability and availability. Evaluate and recommend suitable technologies, tools, and platforms for infrastructure, monitoring, and observability. Drive automation efforts through the development and maintenance of infrastructure-as-code (IaC) solutions. Lead incident response and post-incident analysis efforts, identifying root causes and implementing preventive measures. Implement robust monitoring, logging, and alerting solutions to proactively detect and resolve issues. Collaborate with cross-functional teams, including development, operations, and quality assurance, to ensure alignment and successful project delivery. Stay updated with emerging technologies, industry trends, and best practices in SRE and cloud infrastructure. We are looking for people with: Bachelor's or Master's degree in computer science, engineering, or a related field. 3 years of experience in SRE or systems engineering/architecture roles, with a focus on designing and architecting highly reliable and scalable systems. Solid understanding of cloud platforms (e.g., AWS, Azure, GCP) and experience with infrastructure-as-code (IaC) tools like Terraform. Proficiency in at least one programming language (e.g., Python, Go, Java) and experience with scripting for automation. Strong knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes). Experience with monitoring and observability tools (e.g., New Relic, Dynatrace, Prometheus, Grafana, ELK stack). Proven track record of architecting and implementing highly available and scalable systems in a production environment. Strong problem-solving and analytical thinking abilities, with a focus on driving continuous improvement. Excellent communication and collaboration skills, with the ability to effectively work with cross-functional teams and influence technical decisions. Knowledge of advanced networking concepts, including load balancing, CDN, and DNS management is a plus. Knowledge...