InterviewHack.ai
Start free
Jobs / Replit

Staff Site Reliability Engineer

Replit·Remote - United StatesRemotesenior$20,833–$27,083 USD/mo

In short

  • →Ingeniero SRE de nivel staff que garantiza la confiabilidad y escalabilidad de una plataforma de desarrollo en la nube con millones de usuarios.
  • →Día a día: crea sistemas de observabilidad, dirige incidentes, automatiza operaciones y optimiza Kubernetes; también enseña y guía a equipos.
  • →Lo destacado: se espera que implementes mejoras de alto impacto que transformen la resiliencia del sistema, no solo mantengas el status quo.

Required: Strong written and verbal English skills.

Apply on company site ↗Share on WhatsApp

In ~1 minute you get: who interviews you, the likely questions answered from your CV, and your CV tailored to this job. Free, no card.

🎧Land the interview? Bring the copilot. Our free extension listens to the live interview and flashes 3-4-word anchors from your resume and prep — glance, connect, talk. Get the extension →

What they ask for

  • ✓8-10 años de experiencia en SRE o roles similares (DevOps, ingeniería de infraestructura).
  • ✓Habilidades sólidas en programación con Python o Go.
  • ✓Experiencia comprobada con Kubernetes, GCP y Docker.
  • ✓Capacidad para liderar incidentes y crear sistemas de SLO/SLI.
  • ✓Experiencia en automatización, CI/CD y gestión de infraestructura como código.
  • ✓Habilidades de comunicación y mentoreo para influir en toda la organización.

Don't tick every box? That's normal — your free dossier shows your gaps and how to cover them in the interview.

PythonGoKubernetesDockerGCPTerraformPulumiCI/CDObservabilityMonitoring

Who should you write to at Replit?

Your free dossier identifies the people who'd interview you — their background, what they value, and how to reach out so you stand out before applying.

Compensation: $250K – $325K • Offers Equity. Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. ABOUT THE ROLE: Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. YOU WILL: - Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions. Create dashboards and metrics that provide real-time visibility into system health and performance, enabling proactive issue detection. - Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to monitor and report on these metrics, holding teams accountable and ensuring we maintain high reliability standards while balancing innovation speed. - Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution. Conduct thorough, blameless post-mortems and drive the implementation of preventative measures. Develop and refine runbooks and build automation to reduce Mean Time To Recovery (MTTR). - Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. - Optimize Performance on Kubernetes: Collaborate with core infrastructure and product teams to performance-tune and optimize our large-scale cloud deployments, with a deep focus on Kubernetes, Docker, and GCP. Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. - Debug and Harden Distributed Systems: Dive deep into debugging extremely difficult technical problems across the stack. Use your findings to design and implement long-term fixes that make our systems and products more robust, operable, and easier to diagnose. - Provide Staff-Level Guidance: Review feature and system designs from across the company, acting as a key owner for the reliability, scalability, security, and operational integrity of those designs. - Educate and Mentor: Educate, mentor, and hold accountable the broader engineering team to improve the reliability of our systems, making reliability a core value of the Replit engineering culture. - Build and Integrate: Write high-quality, well-tested code in Python or Go to meet the needs of your customers, whether it's building new internal tools or integrating with third-party vendors. REQUIRED SKILLS AND EXPERIENCE: - 8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering). - Strong programming skills in lang

More jobs like this

Remote DevOps Engineer jobs

Looking for something similar?

Leave your email and we'll alert you when matching jobs appear.

Don't apply unprepared

We research who's interviewing you, tailor your CV and rehearse you live — first one free.

InterviewHack.ai

Prepare for the exact interview: who's interviewing you, a tailored CV, and a real coach.

Product

JobsFree ATS checkerInterview-English checkSalary checkLATAM salary reportFree coursesBlogTailored CVSpoken practiceIt's free

Remote jobs

ReactPythonFull-StackLATAMArgentinaMexicoSee all →

Prepare

Spoken practiceFrontendBackendAI EngineerBy companySell with your CV

Company

For employersAboutContactPrivacyTerms

© 2026 InterviewHack.ai · Your CV is yours. Never used to train anything. · A product of IA-PTY

Similar open roles

Senior Site Reliability Engineer

Replit · Remote - US

→

Commercial Counsel

Replit · Foster City, CA

→

Staff Software Engineer, Identity & Authorization

Replit · Foster City, CA

→

Software Engineer - New Grad (2027)

Replit · Foster City, CA

→