Writer

Site reliability engineer (UK)

Posted 6 Days Ago

Be an Early Applicant

In-Office or Remote

2 Locations

Senior level

In-Office or Remote

2 Locations

Senior level

As a Site Reliability Engineer, you'll ensure platform availability and reliability, automate infrastructure management, design scalable solutions, and lead incident response efforts.

The summary above was generated by AI

🚀 About WRITER

WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI.

Founded in 2020 with office hubs in San Francisco, New York City, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI.

📐 About the role

At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As a site reliability engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.

This is a hybrid position, based out of our New York City or London hubs. You'll report to our director of engineering.

🦸🏻‍♀️ What you'll do

Automate operational tasks and infrastructure management by developing robust tools and platforms using Python, Go, or similar languages, significantly reducing manual toil across our production environment
Design and implement scalable, fault-tolerant infrastructure solutions on public cloud providers (AWS, GCP, Azure) to support WRITER's rapidly expanding, high-traffic AI platform
Own the reliability, performance, and efficiency of WRITER’s core services, defining and upholding stringent Service Level Objectives (SLOs) and Error Budgets
Own the observability stack for monitoring, logging, and alerting systems to ensure rapid detection of issues across our complex distributed systems
Lead incident response, post-mortems, and root cause analyses, applying learnings to proactively prevent future outages and build a more resilient system architecture
Collaborate closely with product and engineering teams, providing expert guidance on system design for reliability, performance, and scalability from conception through launch

⭐️ What you need

A solid 7+ years of experience in site reliability engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems
Deep expertise with cloud platforms (AWS strongly preferred), containerization technologies like Docker and Kubernetes, and Infrastructure-as-Code tools such as Terraform
Strong proficiency in programming languages such as Python, Java, Go for automation and monitoring
Knowledge of monitoring and logging tools (e.g., Prometheus, Grafana, ELK Stack) to maintain system health and performance
Demonstrated ability to Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems
Excellent communication, collaboration, and problem-solving skills, with a talent for building strong relationships and Connecting with cross-functional teams
A strong sense of ownership and accountability, eager to Own mission-critical systems and drive them toward peak performance and unparalleled reliability

🍩 Benefits & perks (UK full-time employees):

Generous PTO, plus company holidays
Comprehensive medical and dental insurance
Paid parental leave for all parents (12 weeks)
Fertility and family planning support
Early-detection cancer testing through Galleri
Competitive pension scheme and company contribution
Annual work-life stipends for:
- Wellness stipend for gym, massage/chiropractor, personal training, etc.
- Learning and development stipend
Company-wide off-sites and team off-sites
Competitive compensation and company stock options

#BI-Remote

Top Skills

AWS

Azure

Docker

Elk Stack

GCP

Grafana

Java

Kubernetes

Prometheus

Python

Terraform

Similar Jobs

Zscaler

Manager, Technical Success

An Hour Ago

Remote or Hybrid

Upgate, Broadland, Norfolk, England, GBR

Senior level

Cloud • Information Technology • Security • Software • Cybersecurity

Lead a team of Technical Success Managers to enhance customer satisfaction and adoption. Engage with C-level executives and strategize account management while driving market impact through collaboration and communication.

Top Skills: Cloud Computing SolutionsNetworking TechnologiesSecurity Technologies

CrowdStrike

Corporate Growth Specialist (Remote, GBR)

3 Hours Ago

Remote or Hybrid

United Kingdom

Junior

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity

The Corporate Growth Specialist develops market routes for CrowdStrike's Falcon platform, focusing on partnership building, sales strategy, and opportunity creation.

Top Skills: Ai-Native PlatformFlex Licensing Model

CrowdStrike

Detection Engineer, Falcon Complete (Remote, GBR)

3 Hours Ago

Remote or Hybrid

United Kingdom

Senior level

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity

The role involves developing high fidelity SIEM detection rules, threat research, mentoring team members, and collaborating across teams to improve detection processes.

Top Skills: LogrhythmLogscaleQradarRegular ExpressionsSentinelSIEMSplunkSumologic

What you need to know about the Manchester Tech Scene

Home to a £5 billion digital ecosystem, including MediaCity, which consists of major players like the BBC, ITV and Ericsson, Manchester is one of the U.K.'s top digital tech hubs, at the forefront of advancements in film, television and emerging sectors like as e-sports, while also fostering a community of professionals dedicated to pushing creative and technological boundaries.