
Lead Site Reliability Engineer
1 week ago
At Xero, we're here to help you supercharge your business. We do this by automating routine tasks, surfacing actionable insights and connecting businesses with the right data, advisors and apps. When that happens, we're not only making life better for small business, we'll be building a stronger economy that can change the world.
About the team
Xero's Incident and Problem Management team is part of the Site Reliability Engineering (SRE) organization and is responsible for the build, delivery and ongoing maintenance of robust processes and tooling around Incident management.
The team is responsible for driving enduring reliability at Xero through robust, consistent and fast response to high severity incidents. They are responsible for building a world-class process and ensuring that that process matures as the demands of the business grow.
About the roles
We\'re looking for a Lead Engineer to join Xero's Incident and Problem Management team. This position requires an experienced SRE professional with a strong technical background, deep experience in SRE, a passion for building and delivering robust processes, and extensive experience of leading technical response to high severity cloud issues.
You will drive best practice across the business and contribute to the ongoing transformation of the Xero SRE culture. As an expert communicator, you will lead technical discussions to identify and track actions associated with and identified during incident situations.
Across our SRE function, we\'re looking for those who are keen to deep dive into causes of incidents and proactively examine the potential causes of future incidents; working with engineering teams to remove the risk of that failure scenario. Ultimately building playbooks and automation to ensure quick and effective responses. In addition, provide ongoing training across the business to ensure the process is well understood and adhered to.
This role will form the backbone of a new team, providing a Technical Duty Officer (TDO) function within the business. TDOs are incident commanders who use SRE skillsets to drive fast mitigation and enduring resolution of impactful events.
What you\'ll do
- Own the incident management process, ensuring it drives enduring reliability across all products and services within Xero.
- Provide expert leadership during critical outages, coordinating multiple teams to ensure streamlined decision-making and quick resolution.
- Lead and advocate for the transformation to a world-leading SRE organization, promoting SRE principles within the Engineering Department.
- Promote a customer-focused approach by addressing and mitigating global customer environment issues, and fostering a culture of continuous learning and technical excellence within the SRE team.
- Develop and implement scalable process frameworks and observability strategies to ensure rapid problem diagnosis, response, and service reliability.
- Collaborate with product teams to thoroughly analyze failures and integrate insights to improve service reliability, scalability, and operational efficiency.
What you\'ll bring
- Previous career experience as a Site Reliability Engineer, in an Operations or Engineering environment
- Strong hands-on coding experience (preferably Python) and knowledge of software engineering best practice
- Networking knowledge and able to troubleshoot TCP/IP, SSL/TLS, DNSSEC, IPsec, and BGP issues
- Strong communication (oral & written) skills including the ability to translate technical issues/concepts into agreed actions
Why Xero?
Offering very generous paid leave to use however you'd like (plus statutory holidays), dedicated paid leave to care for your physical and mental wellbeing as well as an Employee Assistance Program to access mental health care for you and your family. Health insurance, life insurance, and income protection.
We offer wellbeing and sports programmes, employee resource groups, 26 weeks of paid parental leave for primary caregivers, an Employee Share Plan, beautiful offices, flexible working, career development, and many other benefits that reflect our human value.
You'll do the best work of your life at Xero
Seniority level
Not Applicable
Employment type
Full-time
Job function
Engineering and Information Technology
Industries
Software Development
#J-18808-Ljbffr-
Lead Site Reliability Engineer
1 week ago
Melbourne, Victoria, Australia Xero Full timeLead Site Reliability Engineer (Technical Duty Officer)At Xero, we're here to help you supercharge your business. We do this by automating routine tasks, surfacing actionable insights and connecting businesses with the right data, advisors and apps. When that happens, we're not only making life better for small business, we'll be building a stronger economy...
-
Site Reliability Engineer
3 days ago
Melbourne, Victoria, Australia Salient Group Full time $120,000 - $180,000 per yearSite Reliability Engineer | Scale a Next-Gen SaaS PlatformLocation:Melbourne (Hybrid)AboutSalient is proud to be partnering with a fast-growing fintech scale-up that's tackling one of the world's most pressing challenges: financial crime. Their AI-powered SaaS platform is already trusted by leading banks and financial institutions across Australia, New...
-
Site Reliability Engineer
2 days ago
Melbourne, Victoria, Australia Salient Group Full timeGet AI-powered advice on this job and more exclusive features.Direct message the job poster from Salient GroupSite Reliability Engineer | Scale a Next-Gen SaaS PlatformAboutSalient is proud to be partnering with a fast-growing fintech scale-up that's tackling one of the world's most pressing challenges: financial crime. Their AI-powered SaaS platform is...
-
Site Reliability Engineer
2 days ago
Melbourne, Victoria, Australia Salient Group Full timeGet AI-powered advice on this job and more exclusive features.Direct message the job poster from Salient GroupSite Reliability Engineer | Scale a Next-Gen SaaS PlatformAboutSalient is proud to be partnering with a fast-growing fintech scale-up that's tackling one of the world's most pressing challenges: financial crime. Their AI-powered SaaS platform is...
-
Site Reliability Engineers
4 weeks ago
Melbourne, Victoria, Australia Xero Full timeSite Reliability Engineers (Observability) Join to apply for the Site Reliability Engineers (Observability) role at Xero Site Reliability Engineers (Observability) Join to apply for the Site Reliability Engineers (Observability) role at Xero Get AI-powered advice on this job and more exclusive features.At Xero, we're here to help you supercharge your...
-
Site Reliability Engineer
3 days ago
Melbourne, Victoria, Australia BURGEON IT SERVICES Full time $125,000 - $175,000 per yearPosition: Site Reliability Engineer Lead Engineer Location: Melbourne, VIC Duration: 6 months Relevant Exp: 10 years Primary Focus: Ensuring system reliability, scalability, and performance. Key Responsibilities: - Defining SLOs, SLIs, and SLAs for reliability. - Monitoring system performance and reducing toil. - Incident response and root cause...
-
Site Reliability Engineers
3 weeks ago
Melbourne, Victoria, Australia Xero Full timeSite Reliability Engineers (Observability)Join to apply for the Site Reliability Engineers (Observability) role at XeroSite Reliability Engineers (Observability)Join to apply for the Site Reliability Engineers (Observability) role at XeroGet AI-powered advice on this job and more exclusive features.At Xero, we're here to help you supercharge your business....
-
Site Reliability Engineer
2 days ago
Melbourne, Victoria, Australia beBeeHighPerformance Full time $120,000 - $180,000Job Title: High Performance Infrastructure SpecialistAbout Our Organization:We partner with a leading fintech scale-up tackling financial crime through an AI-powered SaaS platform trusted by major banks and institutions.This opportunity allows you to join an engineering-led culture focused on reliability, scalability, and security. As we scale globally, we...
-
Site Reliability Engineer
5 days ago
Melbourne, Victoria, Australia Infosys Full timeAbout InfosysInfosys is a global leader in next-generation digital services and consulting. We enable clients in more than 56 countries to navigate their digital transformation. With over four decades of experience in managing the systems and workings of global enterprises, we expertly steer our clients through their digital journey. We do it by enabling the...
-
Site Reliability Engineer
1 week ago
Melbourne, Victoria, Australia Infosys Full timeAbout Infosys Infosys is a global leader in next-generation digital services and consulting. We enable clients in more than 56 countries to navigate their digital transformation. With over four decades of experience in managing the systems and workings of global enterprises, we expertly steer our clients through their digital journey. We do it by enabling...