Customer Reliability Engineer (CRE)

July 30, 2026

Job Description

We are seeking a Customer Reliability Engineer (CRE) to ensure the reliability, availability, and operational excellence of the organization’s Network Detection and Response (NDR) platform. The role combines Site Reliability Engineering (SRE), DevOps, cloud infrastructure, automation, and customer support to manage production environments, resolve critical incidents, automate operational tasks, and improve system reliability. The ideal candidate should have strong Linux, AWS, networking, monitoring, and programming skills, with a passion for reducing operational toil through automation.


Roles & Responsibilities

Customer Support & Operations

  • Support customer deployments, upgrades, and production environments.
  • Troubleshoot customer issues and resolve critical incidents.
  • Participate in on-call rotations and handle production escalations.
  • Provide technical support for on-premises and air-gapped customer deployments.
  • Communicate effectively with customers during incident resolution.

Site Reliability Engineering (SRE)

  • Maintain highly available and reliable production systems.
  • Perform root cause analysis (RCA) for production incidents.
  • Define and monitor Service Level Objectives (SLOs).
  • Improve system stability and operational efficiency.
  • Reduce operational toil through automation.

Automation & DevOps

  • Design and develop automation tools to eliminate manual operational tasks.
  • Build and maintain Infrastructure as Code (IaC) using Terraform.
  • Develop CI/CD pipelines for infrastructure and software deployments.
  • Automate deployments, upgrades, monitoring, and maintenance activities.
  • Collaborate with engineering teams to implement long-term automation solutions.

Cloud Infrastructure

  • Deploy and manage cloud infrastructure on AWS.
  • Manage AWS services such as EC2, VPC, IAM, and S3.
  • Maintain secure, scalable, and resilient cloud environments.
  • Optimize cloud resources for performance and reliability.

Monitoring & Observability

  • Build and maintain monitoring and observability platforms.
  • Implement metrics, logs, traces, and telemetry pipelines.
  • Monitor system health and application performance.
  • Analyze production issues using monitoring and debugging tools.
  • Improve observability to enhance incident response.

Software Development

  • Develop automation scripts and operational tools.
  • Write production-quality code using Python or Go.
  • Create Shell (Bash) scripts for system administration tasks.
  • Debug and maintain complex distributed systems.
  • Collaborate with product and platform engineering teams.

Networking

  • Troubleshoot networking issues related to production environments.
  • Support DNS, TCP/IP, routing, and network connectivity.
  • Use networking tools for debugging and performance analysis.

Required Skills

Site Reliability Engineering

  • Site Reliability Engineering (SRE)
  • DevOps
  • Incident Management
  • Root Cause Analysis (RCA)
  • High Availability
  • Reliability Engineering
  • Operational Excellence
  • Service Level Objectives (SLOs)

Cloud & Infrastructure

  • Amazon Web Services (AWS)
  • EC2
  • VPC
  • IAM
  • S3
  • Infrastructure as Code (Terraform)
  • CI/CD Pipelines
  • Cloud Infrastructure

Programming & Automation

  • Python
  • Go
  • Bash
  • Shell Scripting
  • Automation Development
  • Infrastructure Automation
  • Tool Development

Linux & System Administration

  • Linux Administration
  • Linux Command Line
  • System Monitoring
  • Performance Tuning
  • System Debugging

Networking

  • TCP/IP
  • DNS
  • Routing
  • Network Troubleshooting
  • Networking Fundamentals

Monitoring & Observability

  • Monitoring
  • Logging
  • Metrics
  • Tracing
  • Telemetry
  • Observability
  • Performance Monitoring

Debugging

  • Troubleshooting
  • Production Support
  • Log Analysis
  • Network Debugging
  • System Diagnostics

Development Tools

  • Terraform
  • CI/CD
  • Git
  • Command Line Utilities

Additional Technologies

  • Scala
  • C
  • C++
  • Rust
  • Haskell
  • PureScript

Soft Skills

  • Problem Solving
  • Analytical Thinking
  • Customer Communication
  • Incident Leadership
  • Collaboration
  • Critical Thinking
  • Decision Making
  • Time Management
  • Adaptability
  • Continuous Learning
  • Ownership & Accountability

Keywords

Customer Reliability Engineer, CRE, Site Reliability Engineering, SRE, DevOps, AWS, EC2, VPC, IAM, S3, Terraform, Infrastructure as Code, CI/CD, Linux, Bash, Shell Scripting, Python, Go, Automation, Monitoring, Observability, Logging, Metrics, Tracing, Telemetry, Incident Management, Root Cause Analysis, Production Support, DNS, TCP/IP, Routing, Network Troubleshooting, Cloud Infrastructure, Reliability Engineering, High Availability, Performance Monitoring, C++, Scala, Rust, Haskell.