Apply Now
Location: Dallas, TX / Scottsdale, AZ, Arizona (AZ), Texas (TX)
Contract Type: C2C,W2
Posted: 6 days ago
Closed Date: 07/31/2026
Skills: GKE and Rancher RKE2
Visa Type: Any Visa

Job Description -

AI-Enabled Platform/SRE Engineer - HYBRID

CVS

Location: Dallas, TX / Scottsdale, AZ (Hybrid)

Length of contract : 1 year (and will be renewed after that)

 

**ROPES ASSESSMENT IS REQUIRED**

**INTERVIEW WILL BE ONSITE IN DALLAS, TX or SCOTTSDALE, AZ**

 

About the Role

 

We are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale.

 

Key Responsibilities

·  Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.

·  Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.

·  Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.

·  Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.

·  Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.

·  Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.

·  Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.

Core Technical Skills

· Site Reliability Engineering (SRE) – Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.

· Kubernetes Platform Engineering – 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.

· Cloud & Infrastructure Automation – Strong experience in GCP, Terraform, Helm, GitHub, CI/CD, and production-grade automation.

· Software Development – 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).

· Observability & Monitoring – Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.

· API & Microservices Engineering – Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.

· AI-Driven Operations (AIOps) – Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.