Sign up to access all features of our service
  • Job search
  • Favorites
  • Create a CV
    New
  • Subscriptions

Site Reliability / Production Engineer

Full-time

Storyteller

  Up to EUR 20,000  per year, on a full-time, contractor contract 
 Fully remote working from anywhere in Egypt!   
Shared out-of-hours UK coverage, including active evening shifts and weekday overnight pager duty 
✨ Exciting high growth product, relied on by leading global brands, particularly within sports  
 Working with the latest hardware, AI tools, and product workflows. 


We are looking for hands-on production engineers who can take ownership when live systems need attention: establish the customer impact, investigate the evidence, take safe action and keep the response moving. 

You will use AI throughout the work, but not as a substitute for judgement. You will be expected to supervise its output, understand the risk of any action and validate that the real customer outcome has recovered. 


ABOUT US 

Storyteller is a high-growth B2B SaaS platform that lets companies integrate Stories into their own apps and websites. Popularised by Instagram and Snapchat, Stories help our clients increase engagement, retention and revenue. 

Our platform includes SDKs for Web, iOS and Android, alongside publishing tools, analytics and advertising support—giving enterprises a complete Stories solution in days. We work with globally recognised sports and media brands, and your work will be used live by millions of people. 

Our production environment spans Storyteller, Storypilot and the services that support our customers’ live workflows. Reliability is therefore about more than infrastructure: we need to understand when customers are affected, respond quickly, coordinate the right people and improve our systems after every material incident. 

About the Role 

We are hiring two Site Reliability / Production Engineers. You will be the first technical response for live incidents during your coverage window.  

Working closely with Support, you will assess customer impact, investigate the system, take proportionate action and bring in product developers only when their specific knowledge or judgement is genuinely needed. 

This is not a passive escalation role. You will own the technical response, solve what you reasonably can yourself and make escalations specific and useful. Depending on the incident, you may restart or scale services, roll back deployments, change configuration, repair data, deploy a bounded fix or make a small code change. 

When incidents are quiet, you will improve the reliability system: reduce alert noise, strengthen customer-outcome monitoring, improve diagnostics, create runbooks and AI Skills, automate repeated work and make our products easier to operate. 


Working Pattern 

This role provides out-of-hours production coverage, so the schedule is a core part of the position rather than occasional overtime. 

  • The two hires will share an agreed rota that ensures one engineer is actively working from 17:00-01:00 UK time, seven days a week.  
  • Neither person will work seven days a week; the active shifts will be divided between both hires, with appropriate rest days.  
  • The two engineers will also share weekday pager coverage from 01:00-06:00 UK time. The detailed allocation of active and pager shifts will be explained during the hiring process.  
  • Weekend daytime coverage is provided separately and is not an additional expectation for these roles.  
  • When there are no live incidents, the active shift will be used for reliability-improvement work.  
  • Meetings and collaboration with management and product teams will be arranged within the agreed working pattern.  

The detailed rota, rest arrangements, leave cover, compensation and on-call terms will be confirmed clearly during the hiring process. Please consider the UK-time evening and overnight requirements carefully before applying. 


RESPONSIBILITIES 

Respond to live incidents 

  • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. 
  • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. 
  • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. 
  • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. 
  • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. 
  • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. 
  • Join customer conversations occasionally when direct technical involvement is genuinely useful. 

Coordinate the right response 

  • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. 
  • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. 
  • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. 
  • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners. 

Improve the reliability system 

  • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes. 
  • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements. 
  • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve. 
  • Create safe, supervised automation for common operational actions. 
  • Work with product teams to close observability, rollback, runbook and supportability gaps. 
  • Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner. 
  • Make reliability and on-call performance easier for the company to understand and improve over time. 


QUALIFICATIONS 

What we're looking for 

  • Agency and ownership - You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed. 
  • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded. 
  • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan. 
  • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions. 
  • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals. 
  • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work. 
  • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand. 
  • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it. 

Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response. It is not an automatic requirement: we will also consider candidates who demonstrate exceptional ownership, judgement, technical aptitude, learning velocity and performance in the practical assessment. 

You do not need experience with every technology in our stack, a previous SRE job title, people-management experience or the ability to recall every command without AI assistance. The ability to learn an unfamiliar environment, act safely and validate your work matters more than matching a long technology checklist. 


Nice to have 

  • Cloud platforms such as Azure or Cloudflare. 
  • Distributed application and API diagnostics. 
  • Databases, queues and background-processing systems. 
  • Observability, alerting and incident-management platforms. 
  • Infrastructure, deployment and release automation. 
  • Application development and safe production debugging. 
  • AI coding agents and workflow automation. 


RECRUITMENT PROCESS

We keep the process straightforward, practical and respectful of your time. 

1. Hiring Manager Conversation (20-30 mins) 

A short call to get to know you, talk through the working pattern and answer your questions. 

2. Paid Take-home Task (~60-90 mins) 

A small, bounded production-incident exercise using evidence such as a Support report, alerts, logs, metrics, deployment history, code and an imperfect runbook. We compensate you for completing it regardless of the outcome. 

You are encouraged to use AI. We are interested in how you establish impact, investigate and revise hypotheses, choose a safe response, validate the outcome and communicate the incident - not in your ability to reproduce commands or syntax from memory. 

3. Task Review and CTO Interview (60-75 mins) 

We will review your submission together, explore the decisions and trade-offs you made, and discuss how you supervised AI-generated analysis or changes. You will also meet Dave, our CTO, and talk about production judgement, escalation, validation, communication and how you improve the system after an incident. 

And that's it.

----------------------

Privacy Notice
We process your personal data for recruitment purposes in line with UK data protection law. AI tools may assist in reviewing applications, but decisions are made by our team. We retain data only as necessary for recruitment and compliance. You can request access or deletion of your data at any time by emailing View email address on storyteller.applytojob.com.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Site Reliability / Production Engineer in Egypt vacancy
  •  ...brands in US TV and entertainment ✨ A product design role with AI agents, working prototypes...  ...Work with Account Management, Product, Engineering and Solutions to choose pragmatic...  ...strong remote collaboration habits and reliable internet access. A portfolio is essential... 

    Storm Ideas

    Egypt
    15 days ago
  •  ...Mills  Production & Planning Director is Highly Needed to Join Eshratex' Cotton Spinning Company Alexandria -Egypt , Expertise in Cotton...  ...:   Education: Bachelor’s degree in Textile Engineering, Industrial Engineering, Age : 35-45 years old preferably... 

    Qureos Inc

    Egypt
    7 days ago
  •  ...seeking a highly motivated and skilled Software Testing Engineer to join our dynamic team. As a Software Testing Engineer...  ...you will play a crucial role in ensuring the quality and reliability of our software products. You will be responsible for designing, developing, and... 

    Corporatica

    Egypt
    more than 2 months ago
  •  ...We are seeking a technically skilled and customer-focused Sales Engineer with direct experience implementing Exposure Management and...  ...primary technical resource during the sales process, including product demonstrations, proof-of-concept (PoC) deployments, and technical... 

    spiderSilk

    Egypt
    11 days ago
  • 25000 - 40000 EGP per month

     ...motivated and detail-oriented Full-Stack QA Engineer (Manual & Automation) to join our...  ...automated strategies to deliver robust and reliable solutions to our clients. Salary: EGP...  ...* Collaborate closely with developers, product managers, and other stakeholders to understand... 

    Qureos Inc

    Egypt
    7 days ago
  •  ...anywhere in Egypt!    ✨ Exciting high growth product, relied on by leading global brands,...  ...Role  This is an AI-native quality engineering role. You will use AI as a core part of...  ...and are easy to diagnose;  Tests run reliably in the workflows the team actually uses;... 

    Storyteller

    Egypt
    a month ago
  •  ...ABOUT THIS ROLE We're looking for full-stack software engineers who build production applications - APIs, frontends, data pipelines, and more...  ..., and data Identify opportunities for performance, reliability, and cost improvements in client AWS environments Advise... 

    Protagona

    Egypt
    a month ago
  •  ...ABOUT THE ROLE As a Senior QA Engineer at Protagona, you will make an impact by ensuring the quality, reliability, and performance of the solutions we build for our...  ...requirements through automation, release, and production monitoring - while keeping the end user's experience... 

    Protagona

    Egypt
    a month ago
  •  ...and all the latest software, AI tools and productivity tools you will need  Who We Are At...  ...dedicated AI Operations team keeping that stack reliable and well-run. You will be stepping into...  ...AI tools” — a role where the marketing engine itself is built from AI agents, reusable... 

    Storm Ideas

    Egypt
    17 days ago
  •  ...rare opportunity to build the conversion engine of an AI-native agency  High-end, fast...  ...all the latest software, AI tools and productivity tools you will need  Who We Are At Storm...  ...AI Operations team keeping that stack reliable and well-run. You will be stepping into... 

    Storm Ideas

    Egypt
    17 days ago
  •  ...Executive on our KSA Exclusive Supply & Fulfillment team, you own the engine that makes “a life without chores” real: making sure every...  ...Weeks 9–12: lead a small efficiency improvement, and become a reliable go-to for customer escalations in your area. What you must... 
    Remote job

    Justlife

    Egypt
    1 day ago
  •  ...Entertainment   High-end, fast computer   All the latest software and productivity tools you’ll need About the Role  We are hiring a Client...  ..., contacts, and handover information current so work is reliable and repeatable What We Are Looking For Experience... 

    Storm Ideas

    Egypt
    17 days ago
  • Reliability Engineer   Sukari is located approximately 750km from Cairo and 25km from Marsa Alam...  ...poured in June 2009 and commercial production began in April 2010, making Sukari the...  ...cost of ownership.   Location: On-site in Marsa Alam Contract Type: Permanent... 

    Globe 24-7

    Marsa Alam
    14 days ago
  •  ...will involve office-based work alongside attendance at client sites, vessels and offshore locations. The position can also suit...  ...Supporting marine assurance and MWS activities. Working with Engineering and Dynamic Positioning Managers on multi-disciplinary projects... 

    Consortio Recruitment Group

    Egypt
    9 days ago
  •  ...remote working from anywhere in Egypt!  ✨ Exciting high growth product, relied on by leading global brands, particularly within...  ...AI Operations Manager to help drive the operational quality, reliability, and effectiveness of Storyteller’s AI tooling.  This role sits... 

    Storyteller

    Egypt
    17 days ago
  •  ...about how the work gets made. You direct agents that operate real production tools (design files through connectors like the Figma MCP;...  ...context and options  Keep files, naming, versions, and handovers reliable, so work can be found, reused, and picked up by people and... 

    Storm Ideas

    Egypt
    17 days ago
  • The Production Engineer is responsible for assisting the PSD team to deliver safe, efficient, and reliable product and service delivery for our customers. The Production Engineer develops and delivers optimized technical job designs for the area of responsibility. This... 

    Schlumberger

    Cairo
    15 days ago
  •  ...having earned its reputation as a reliable and respected international...  ...than a half century in the Engineering, Procurement and Construction...  ...highly motivated Mechanical Site Engineer to join our top-...  ...machinery and equipment. ~ Monitor production line performance and... 

    Archirodon Group N.V

    Suez
    9 days ago
  •  ...High-end, fast computer   All the latest software and productivity tools you’ll need About the Role  You'll work directly with...  ...work. We are more interested in your operational judgement, reliability, follow-through, and ability to use modern tools well than in... 

    Storm Ideas

    Egypt
    11 days ago
  •  ...scale, secure, and maintain production services. Write or improve...  ...call rotation, support software engineers, and debug difficult...  ...performance, scalability, and reliability of Lightspark infrastructure....  ...of experience in Production, Site Reliability, or DevOps Engineering... 

    Lightspark

    Remote
    13 days ago
  •  ...and smarter mobility worldwide, connecting cities as we reduce carbon and replace cars. Could you be the full-time Reliability Growth engineer  in Cairo we’re looking for? Your future role Take on a new challenge and apply your system-level configuration expertise... 

    Alstom

    Cairo
    23 days ago
  •  .../ Mill Manager) # Operational & Strategic Leadership Production Targets: Oversee the entire process from Raw Cotton intake to...  ...Energy Oversight (This is where an Electrical Maintenance Engineering background is a "Superpower") Energy Optimization... 

    Qureos Inc

    Egypt
    7 days ago
  •  ...Responsibilities Work across four to five products with different technology stacks and improve shared platform capabilities....  ...product decisions, and technical risks independently. Raise engineering quality through meaningful tests, code reviews, and accurate living... 

    0g Labs

    Remote
    15 days ago
  •  ...high-quality pharmaceuticals. With a long-standing heritage rooted in pharmacies, we are perceived as a reliable and trustworthy partner since 1895. With our products we help people protect and regain a dignified and able life. With our proven Generics, we ensure that... 
    Egypt
    a month ago
  •  ...high-growth businesses spanning a creative agency and a SaaS product trusted by leading global brands, particularly within sports and...  ....  You build your own tools. Airtable, HRIS admin, prompt engineering, no-code automation (Zapier/Make/Claude skills or similar). You... 

    Storyteller

    Egypt
    17 days ago
  •  ...implementations and monitoring of all mechanical related works at site as per approved drawings and methods and safety rules....  ...Good communication skills. Bachelor's Degree in Mechanical Engineering Problem solving. Creativity Job Category Job Requirements... 
    Cairo
    3 days ago
  •  ...Job Description New Job Opportunity in Construction Ekas Contracting is looking to expand its site team: Job Title: Site Engineer (Civil / Architecture) ⏳ Experience Required- 1 3 Years Key Responsibilities: • Supervise daily site operations and ensure... 
    Alexandria
    1 day ago
  •  ...enhance the safety standards, increase the number of passenger trains, and boost freight capacity. Today, we are recruiting for a  Site Engineer to join our team in Egypt. This is an excellent career opportunity for professionals looking to join a truly international team... 
    Cairo
    9 hours ago
  • 2000 - 4000 USD per month

     ...Profit & Loss (P&L) statements . Oversee bank reconciliations. Monitor the company's cash position. Maintain accurate and reliable financial information for CFO review. Team Leadership & Client Support Supervise Accounts Receivable and Accounts Payable... 

    Entrepreneur Cooperative

    Egypt
    1 day ago
  •  ...and stakeholders to meet agreed timelines. Contribute across product development, including specification, implementation, code...  ...eliminate ambiguity, and drive projects forward. Participate in engineering recruiting through interviewing, problem specification, and... 

    Opto Investments

    Remote
    17 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability / Production Engineer. Be the first to apply!