Job Description
We are seeking a Customer Reliability Engineer (CRE) to ensure the reliability, availability, and operational excellence of the organization’s Network Detection and Response (NDR) platform. The role combines Site Reliability Engineering (SRE), DevOps, cloud infrastructure, automation, and customer support to manage production environments, resolve critical incidents, automate operational tasks, and improve system reliability. The ideal candidate should have strong Linux, AWS, networking, monitoring, and programming skills, with a passion for reducing operational toil through automation.
Roles & Responsibilities
Customer Support & Operations
- Support customer deployments, upgrades, and production environments.
- Troubleshoot customer issues and resolve critical incidents.
- Participate in on-call rotations and handle production escalations.
- Provide technical support for on-premises and air-gapped customer deployments.
- Communicate effectively with customers during incident resolution.
Site Reliability Engineering (SRE)
- Maintain highly available and reliable production systems.
- Perform root cause analysis (RCA) for production incidents.
- Define and monitor Service Level Objectives (SLOs).
- Improve system stability and operational efficiency.
- Reduce operational toil through automation.
Automation & DevOps
- Design and develop automation tools to eliminate manual operational tasks.
- Build and maintain Infrastructure as Code (IaC) using Terraform.
- Develop CI/CD pipelines for infrastructure and software deployments.
- Automate deployments, upgrades, monitoring, and maintenance activities.
- Collaborate with engineering teams to implement long-term automation solutions.
Cloud Infrastructure
- Deploy and manage cloud infrastructure on AWS.
- Manage AWS services such as EC2, VPC, IAM, and S3.
- Maintain secure, scalable, and resilient cloud environments.
- Optimize cloud resources for performance and reliability.
Monitoring & Observability
- Build and maintain monitoring and observability platforms.
- Implement metrics, logs, traces, and telemetry pipelines.
- Monitor system health and application performance.
- Analyze production issues using monitoring and debugging tools.
- Improve observability to enhance incident response.
Software Development
- Develop automation scripts and operational tools.
- Write production-quality code using Python or Go.
- Create Shell (Bash) scripts for system administration tasks.
- Debug and maintain complex distributed systems.
- Collaborate with product and platform engineering teams.
Networking
- Troubleshoot networking issues related to production environments.
- Support DNS, TCP/IP, routing, and network connectivity.
- Use networking tools for debugging and performance analysis.
Required Skills
Site Reliability Engineering
- Site Reliability Engineering (SRE)
- DevOps
- Incident Management
- Root Cause Analysis (RCA)
- High Availability
- Reliability Engineering
- Operational Excellence
- Service Level Objectives (SLOs)
Cloud & Infrastructure
- Amazon Web Services (AWS)
- EC2
- VPC
- IAM
- S3
- Infrastructure as Code (Terraform)
- CI/CD Pipelines
- Cloud Infrastructure
Programming & Automation
- Python
- Go
- Bash
- Shell Scripting
- Automation Development
- Infrastructure Automation
- Tool Development
Linux & System Administration
- Linux Administration
- Linux Command Line
- System Monitoring
- Performance Tuning
- System Debugging
Networking
- TCP/IP
- DNS
- Routing
- Network Troubleshooting
- Networking Fundamentals
Monitoring & Observability
- Monitoring
- Logging
- Metrics
- Tracing
- Telemetry
- Observability
- Performance Monitoring
Debugging
- Troubleshooting
- Production Support
- Log Analysis
- Network Debugging
- System Diagnostics
Development Tools
- Terraform
- CI/CD
- Git
- Command Line Utilities
Additional Technologies
- Scala
- C
- C++
- Rust
- Haskell
- PureScript
Soft Skills
- Problem Solving
- Analytical Thinking
- Customer Communication
- Incident Leadership
- Collaboration
- Critical Thinking
- Decision Making
- Time Management
- Adaptability
- Continuous Learning
- Ownership & Accountability
Keywords
Customer Reliability Engineer, CRE, Site Reliability Engineering, SRE, DevOps, AWS, EC2, VPC, IAM, S3, Terraform, Infrastructure as Code, CI/CD, Linux, Bash, Shell Scripting, Python, Go, Automation, Monitoring, Observability, Logging, Metrics, Tracing, Telemetry, Incident Management, Root Cause Analysis, Production Support, DNS, TCP/IP, Routing, Network Troubleshooting, Cloud Infrastructure, Reliability Engineering, High Availability, Performance Monitoring, C++, Scala, Rust, Haskell.