The role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Brazil.
Join a globally distributed engineering team focused on operating and automating mission-critical infrastructure at scale. In this role, you will apply strong software engineering principles to site reliability, cloud operations, and infrastructure automation. You will work across the full technology stack, from bare-metal systems and Linux networking to containers, Kubernetes, and applications. Your work will help improve reliability, scalability, security, and operational efficiency for complex customer environments. The role combines hands-on engineering with continuous learning as infrastructure technologies and open source capabilities evolve. You will collaborate with experienced engineers in a high-accountability environment where automation, metrics, and code drive operational decisions.
Accountabilities:
- Design, develop, and maintain reliable infrastructure automation using Python and software engineering best practices.
- Operate and improve large-scale infrastructure spanning bare metal, private and public cloud environments, containers, and applications.
- Apply DevSecOps principles across the infrastructure and application lifecycle, with a strong focus on security, reliability, and automation.
- Architect, deploy, operate, and continuously improve technologies such as OpenStack, Kubernetes, and software-defined storage.
- Develop automation and operational tooling that reduces manual intervention and improves consistency, scalability, and resilience.
- Diagnose and resolve complex infrastructure and production issues across Linux systems, networking, storage, and application layers.
- Use metrics, monitoring, code, and data-driven analysis to understand system behaviour and improve operational performance.
- Contribute to infrastructure upgrades and modernization initiatives, helping keep environments aligned with current technology capabilities.
- Work across the full infrastructure stack, developing a broad understanding of systems ranging from operating-system and kernel components to containers and application platforms.
- Support mission-critical services and maintain high standards of reliability in demanding operational environments.
- Collaborate with globally distributed engineering teams to improve processes, tooling, documentation, and operational practices.
- Continuously evaluate emerging open source infrastructure technologies and identify opportunities to improve reliability and efficiency.
Requirements:
- Bachelor’s degree in Computer Science, Software Engineering, or a related technical discipline.
- Strong professional experience in site reliability engineering, cloud operations, infrastructure engineering, or a closely related field.
- Advanced Python software development skills and the ability to approach infrastructure automation as a software engineering discipline.
- Strong Linux expertise, including practical knowledge of Linux networking and storage.
- Hands-on operational experience managing production infrastructure and mission-critical services.
- Experience troubleshooting complex technical issues across multiple layers of the infrastructure stack.
- Familiarity with modern infrastructure automation, DevOps, and DevSecOps practices.
- Experience with OpenStack or Kubernetes deployment and operations is highly desirable.
- Familiarity with public or private cloud infrastructure and management platforms.
- Strong understanding of infrastructure reliability, scalability, monitoring, automation, and performance.
- Excellent interpersonal and communication skills, with the ability to collaborate effectively across distributed teams.
- Curiosity, flexibility, accountability, and a strong commitment to continuous learning.
- Ability to work effectively in a high-pressure operational environment while maintaining engineering rigor and attention to detail.
- Willingness and ability to travel internationally twice a year for company events lasting up to two weeks.
Benefits:
- Compensation based on geographical location, experience, and performance.
- Performance-driven annual bonus or commission in addition to base compensation.
- Distributed and remote-first working environment.
- Twice-yearly in-person team sprints and opportunities to collaborate with colleagues internationally.
- USD 2,000 annual personal learning and development budget.
- Annual compensation review.
- Recognition rewards.
- Annual holiday leave.
- Maternity and paternity leave.
- Employee Assistance Programme.
- Opportunities to travel to new locations and meet colleagues from around the world.
- Priority Pass and travel upgrades for long-haul company events.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1