Site Reliability Engineer
AI summary
Heirs Insurance Ltd is hiring a Site Reliability Engineer in Lagos, Nigeria. This full-time remote role focuses on platform reliability, monitoring, incident response, observability, automation, and chaos testing. The position reports to the SRE Lead and requires a BA/BSc/HND with at least 2 years of experience.
- Full-time remote SRE role at a growing insurance company
- Requires BA/BSc/HND and minimum 2 years experience
- Based in Lagos, Nigeria
- Focus on monitoring, incident response, and automation
- Work with NSOC, DevSecOps, and security teams
AI job guide
Use this guide to check salary signals, requirements, documents, application steps and safety before you apply.
AI salary guide
Not enough public dataNot enough public salary data is available for this exact role. Before applying, prepare to ask about gross pay, benefits, contract length, probation period, transport and any allowances.
Can you qualify for this role?
- Required2+ years of relevant experienceThe job post includes a minimum experience signal.
- RequiredEducation or certification mentioned in the postThe captured text mentions education, a diploma, certificate, or licence.
- PreferredPractical evidence in security, seguranca, site reliability engineerThe tags and summary point to skills connected with this role.
- RequiredAvailability to work in Not specifiedThe vacancy is associated with this location.
- UnclearComfort with the Full Time contract termsConfirm hours, duration, probation and benefits at the original source.
Documents to prepare
- Likely requiredUpdated CV
- Role specificCover letter or short employer message
- OptionalProfessional references
- Role specificAcademic or professional certificates
- VerifyID or passport only after verifying the employer
Application tips for this job
- Place your strongest Site Reliability Engineer evidence in the first half of your CV.
- In your cover letter or employer message, connect your experience to Heirs Insurance Ltd and the role in Not specified.
- Add concrete examples related to security, seguranca, site reliability engineer, ideally with measurable outcomes or clear responsibilities.
- Follow the instructions from MyJobMag Nigeria; avoid sending documents to unofficial contacts or copied links.
- Confirm the deadline, interview location and employer contact before sharing personal documents.
- Prepare a polite question about pay, benefits and contract terms for later interview stages.
Source and safety check
- MyJobMag Nigeria
- Original source link available
- Application method is clear
- Deadline not specified
- No major risk signal was detected in the captured text.
Never pay for interviews, shortlisting, medical checks, uniforms, or job placement. Confirm every application at the original source before sharing personal documents. Report suspicious listing.
Interview preparation
- What experience makes you a strong fit for this Site Reliability Engineer role in security, seguranca?
- How have you handled responsibilities similar to those in this job post?
- Are you available to work in Not specified under the listed contract or schedule?
- Prepare examples with clear responsibilities, tools used and measurable outcomes.
- Review the source and research Heirs Insurance Ltd before the interview.
Ask what the first priorities will be in the role and how success will be measured.
Similar jobs to consider
Use AI to apply better
After confirming the original source, use Career Assistant to check role fit, tailor your CV and prepare a cover letter or employer message.
Original source description
Heirs Insurance is a general insurance company challenging traditional insurance by providing simple and accessible protection for vehicles, homes, business and more. Site Reliability Engineer Job Type Full Time , Remote Qualification BA/BSc/HND Experience: 2 years Location Lagos Job Field Engineering / Technical About the job The Site Reliability Engineer is the hands-on operator and automation builder who keeps the platform running. Reporting to the SRE Lead, you will be responsible for the day-to-day health of the platform --- monitoring, alerting, incident response, and the continuous automation work that makes the platform easier to operate and harder to break. You will work closely with the DevSecOps engineers, cloud engineers, and security team to ensure that reliability is built into every layer of the platform, not bolted on afterwards. This is a demanding role on a high-stakes platform --- and a rare opportunity to do SRE work that genuinely matters at scale. What You'll Do NSOC Operations & Monitoring --- Work as part of the 24/7 Network & Security Operations Centre (NSOC) --- actively monitoring the platform's health across all services, cloud environments, and network layers. Triage incoming alerts, distinguish signal from noise, escalate genuine incidents promptly, and maintain the situational awareness that keeps the team ahead of problems before they become outages. Incident Response & Triage --- Serve as a first responder for platform incidents --- detecting, triaging, and containing issues in real time. Execute runbooks under pressure, coordinate with engineering teams during active incidents, and own your part of the resolution. After every significant incident, contribute to the blameless post-mortem process --- documenting what happened, what worked, and what needs to change. Observability Implementation & Maintenance --- Implement, maintain, and continuously improve the platform's observability stack --- including centralised logging, metrics collection (Prometheus, Datadog, or equivalent), distributed tracing (Jaeger, OpenTelemetry, or equivalent), and alerting rules. Instrument new services from day one, tune alerting thresholds to reduce noise without missing real issues, and build dashboards that give the team the operational visibility they need. Automation & Toil Reduction --- Identify repetitive operational tasks and build the automation that eliminates them. Write scripts, tools, and workflows in Bash, Python, or equivalent to reduce toil --- replacing manual processes with reliable, repeatable automation. Toil that exists today should not exist next quarter. This is an ongoing discipline, not a project with an end date. Reliability & Chaos Testing --- Support the SRE Lead in running the platform's reliability engineering programme --- executing load tests, stress tests, and chaos engineering experiments to validate the platform's resilience under real conditions. Document findings, flag vulnerabilities, and work with engineering teams to address weaknesses before they become production incidents. Runbook Development & Maintenance --- Write, test, and maintain operational runbooks for all known failure modes and recovery procedures. A runbook is only valuable if it works --- regularly test each runbook against real or simulated conditions, update it when the platform changes, and make sure every procedure is accurate and actionable. Runbooks are not documentation --- they are operational tools that must be ready to use at 2am. SLO Monitoring & Reporting --- Monitor SLO attainment and error budget consumption across all platform services on an ongoing basis. Proactively flag services approaching error budget limits to the SRE Lead and relevant engineering teams. Track and report DORA metrics --- deployment frequency, lead time, change failure rate, and MTTR --- as part of the team's ongoing performance visibility. Platform Health & Capacity Monitoring --- Continuously monitor resource utilisation, service performance, and infrastructure health across both cloud environments (AWS and Azure). Identify trends that suggest emerging capacity constraints or degrading performance, and surface these findings to the SRE Lead and Platform Manager before they become incidents. CI/CD Pipeline Health --- Monitor the health and performance of the platform's CI/CD pipelines in collaboration with the DevSecOps team. Identify pipeline failures, flaky tests, or degraded build performance and escalate promptly. Reliable pipelines are part of platform reliability --- treat them accordingly. Documentation & Knowledge Sharing --- Maintain accurate, up-to-date operational documentation --- procedures, architecture notes, incident timelines, and lessons learned. Share knowledge actively with your SRE colleagues and across the broader engineering team. A well-documented platform is a more reliable platform. What We're Looking For Must Have 2+ years of SRE, DevOps, or platform operations experience: in a production environment --- with hands-on responsibility for monitoring, alerting, and incident response. Hands-on experience: with observability tooling --- logging (ELK, Loki, Datadog, or equivalent), metrics (Prometheus, Datadog, or equivalent), and alerting. You have built dashboards and tuned alert rules, not just viewed them. Experience: responding to production incidents --- triaging, containing, and resolving issues under pressure. You understand incident severity classification and escalation paths. Scripting skills in Bash and Python (or equivalent) --- for writing automation, operational tooling, and toil-reduction scripts. Solid understanding of cloud infrastructure on AWS and/or Azure --- compute, managed services, networking, and the platform-level health indicators you should be monitoring. Genuine care about reliability --- you understand why SLOs and error budgets matter, and you approach operational work with the discipline that high-availability financial services demand. Nice to Have Experience: with distributed tracing --- Jaeger, OpenTelemetry, AWS X-Ray, or equivalent --- and using trace data to diagnose performance and reliability issues. Exposure to chaos engineering tools --- Chaos Monkey, Gremlin, or equivalent --- or experience: running load and stress tests in production or pre-production environments. Familiarity with Kubernetes operational patterns --- pod health, cluster observability, resource constraints, and Kubernetes-native alerting and scaling behaviours. Experience: in a regulated environment --- fintech, banking, or similar --- with an understanding of the uptime, audit, and compliance expectations financial services carry. AWS and/or Microsoft Azure certification --- SysOps, Cloud Practitioner, or equivalent entry-to-mid-level cloud certifications. Familiarity with DORA metrics and what they tell you about engineering team health and deployment risk. Check how your CV aligns with this job Method of Application Interested and qualified? Go to Heirs Insurance Ltd on heirs-technologies.breezy.hr to apply Build your CV for free. Download in different templates.