Site Reliability Engineer
Apply now »Date: Jul 21, 2026
Location: Guadalajara, JAL, MX
Company: NTT DATA Services
NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com
Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change based on client requirements. For employees near an NTT DATA office or client site, in-office attendance may be required for meetings or events, depending on business needs. At NTT DATA, we are committed to staying flexible and meeting the evolving needs of both our clients and employees. NTT DATA recruiters will never ask for payment or banking information and will only use @nttdata.com and @talent.nttdataservices.com email addresses. If you are requested to provide payment or disclose banking information, please submit a contact us.
Role Summary
We are looking for a strong, hands-on Senior Site Reliability Engineer to support and improve cloud operations for microservice-based platforms. This role requires a senior engineer who can independently manage production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments.
The ideal candidate should be technically strong, proactive, comfortable in production support, and able to reduce operational toil through automation while improving service availability, performance, scalability, and resilience.
Key Responsibilities
- Own and improve reliability of cloud-based services and supporting infrastructure.
- Participate in on-call rotation and support production systems outside normal business hours.
- Lead incident response activities including triage, escalation, mitigation, and service restoration.
- Drive blameless postmortems and ensure corrective actions are tracked to closure.
- Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis.
- Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
- Support and improve cloud/container platforms across AWS and Azure.
- Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems.
- Build automation to reduce manual effort and improve operational efficiency.
- Configure and improve monitoring, alerting, logging, diagnostics, and observability.
- Analyze performance and capacity trends to identify bottlenecks and improve scalability.
- Troubleshoot complex infrastructure, networking, application runtime, and cloud platform issues.
- Support disaster recovery planning, validation, and recovery readiness.
- Create and maintain operational runbooks, support procedures, and engineering documentation.
- Coach and guide other engineers on SRE best practices, reliability, automation, and operational excellence.
Mandatory Skills
- 6–8+ years of experience as an SRE, DevOps Engineer, Infrastructure Engineer, Cloud Engineer, or Platform Engineer.
- Strong hands-on experience with AWS and Azure cloud platforms.
- Strong experience with Terraform for Infrastructure as Code.
- Experience with Atlantis, ArgoCD, or similar infrastructure/deployment automation tools.
- Strong hands-on experience with Docker and Kubernetes.
- Experience designing, maintaining, and troubleshooting complex CI/CD pipelines.
- Strong production support experience including incident management, RCA, postmortems, and runbook creation.
- Strong observability experience: monitoring, alerting, logging, diagnostics, and performance analysis.
- Good understanding of cloud networking, security, access controls, and InfoSec practices.
- Experience with version control, branching, merging, pull requests, and conflict resolution.
- Understanding of cloud cost optimization and resource utilization.
- Strong communication skills and ability to work with DevOps, Engineering, Product, and Delivery teams.
Good-to-Have Skills
- Experience with microservice-based platforms.
- Experience with tools such as Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics, or similar.
- Scripting/programming experience using Python, Bash, Go, or Java.
- Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering.
- Experience with disaster recovery testing and production readiness reviews.
- Experience supporting customer-facing, high-availability platforms.
- Prior experience mentoring junior engineers or leading technical troubleshooting.
Job Segment:
Developer, Java, Consulting, Technology