Apply Now
Location: Austin / Charlotte
Contract Type: C2C
Posted: 6 hours ago
Closed Date: 07/31/2026
Skills: AI LLM’s – GPT - Open AI, Claude, Gemini
Visa Type: Any Visa

Job Title: SRE AI Engineer

 

Location: Austin / Charlotte

 

Job Description:

 

We are currently seeking a highly skilled SRE hands-on AI Engineer with solid experience in AI Observability and instrumentation approaches for AI systems, AI Agents development to perform detection, diagnosis and autonomous self-healing operates on AI Control Plane.

Observability data collection and automation to help lead transformational initiatives within IT operations, encompassing development as well. As a crucial figure in this role, you will participate/help with various technology domain groups and cross functional teams on unified observability gap analysis and solutioning (automation and manual fixes)

 

Responsibilities:

 

  • Incorporate GenAI tooling and agentic capabilities to strengthen reliability outcomes across monitoring/alerting, rapid incident response, change management/testing, and DevOps/deployment processes.
  • Experience building agentic workflows using LLMs, tool-calling, function-calling, multi-agent orchestration, and event-driven automation.
  • Experience with Agent-to-Agent communication, AI agent federation, and enterprise AI control-plane concepts.
  • Experience implementing AI control-plane governance, including policy-based execution, approval workflows, audit trails, guardrails, and risk-based remediation controls.
  • Expertise in Observability as a service, Dashboard as a services, monitoring as a services and alert as a service in all technology domains (application, infrastructure, database, security, middleware, network etc.,) Telemetry data collection using Dynatrace APM, SolarWinds, CISCO Switches, F5, Databases, Open-Source tools (Prometheus and Grafana), Log Aggregations (Kibana or Splunk) and AIOPS Tools.
  • Practical experience implementing Golden Signals (latency, traffic, errors, saturation) using related telemetry sources.
  • Configure application performance monitoring (APM), infrastructure monitoring, synthetic monitoring, RUM, and log monitoring.
  • Integrate Dynatrace with CI/CD pipelines, alerting tools, ITSM systems, and incident automation frameworks.
  • Tune alert thresholds, baselines, and AI-driven anomaly detection to reduce noise and improve actionable insights.
  • Deeper understanding of Login authentication mechanisms using Ping, ForgeRock and SiteMinder technologies (session management and cookie management)
  • Define best practices and principles for SRE, including monitoring, alerting, and automation.
  • Collaborate with development teams on resiliency to ensure that services and applications are designed with operational reliability in mind.
  • Implement monitoring systems to assess the performance of applications and infrastructure and proactively identifying areas for optimization.
  • Ability to develop close relationship with other operational teams to integrate SRE practices and drive overall operational improvements across enterprise.
  • Stay up to date on industry trends, new technologies, and best practices in SRE and applying relevant advancements to the organization.
  • Ability to build strong working relationships across different levels, client focus mindset.

 

 

 

Qualifications:

  • Around 7-10 years of SRE hands on experience with AI OPS, cloud technologies, development, SRE toolsets and automation
  • Hands-on experience implementing Retrieval-Augmented Generation using approved enterprise knowledge sources such as runbooks, SOPs, RCA documents, incident history, architecture documents, and knowledge articles.
  • Expertise SPEC driven and Prompting using Ai IDE tools – Cursor, Kiro and Good Antigravity
  • Hands-on experience AI LLM’s – GPT - Open AI, Claude, Gemini etc.,
  • Experience with LangChain, LangGraph, Bedrock Agents, Azure AI Foundry, or Vertex AI Agent Builder
  • Experience with vector databases such as OpenSearch, Pinecone, Chroma, Redis Vector, or pgvector.
  • Experience performing Observability current-state assessments, gap analysis and solutioning (automation and manual fixes) in all technology domains (application, infrastructure, database, security, middleware, network etc.,),
  • Strong hands-on automation experience in Observability as a code, dashboard as a code, monitoring as a code, alert as a code (Instrumentation, templates, automatic deployment, visualization and alerting)
  • Strong hands-on experience with any Cloud Technology (AWS): Control Tower, Project Setup, Creating Accounts, RDS, SSO
  • Monitoring & alerting setup experience with Splunk, Prometheus, Grafana, Kibana, ELK, with pref. for APM (Dynatrace).
  • Strong skills in APM, distributed tracing, synthetic & real user monitoring, log monitoring, and Davis AI configuration
  • Own the design, configuration, CICD deployment, and optimization for enterprise-wide observability tools.
  • Experience integrating, automation, and cloud platforms (AWS, Azure, GCP).
  • Extended experience instrumenting OTEL Framework.
  • Hands on experience with Dynatrace Plug-and-play observability modules (OKit) development for Observability Developers Java and .Net applications.
  • Define monitoring standards, best practices, and governance to ensure consistency and scalability.
  • Experience to deploy and tune OneAgent, build end-to-end PurePath tracing, and leverage Smartscape topology for proactive performance monitoring and root-cause analysis.
  • Collaborate with application and infrastructure teams to troubleshoot performance issues and implement permanent fixes.

Good to have:

  • Any of the relevant professional certifications – AIOPS related certifications, Certified Site Reliability Engineer (CSRE), Certified Kubernetes Administrator (CKA), AWS Certified DevOps Engineer Professional, , Google Cloud Professional; DevOps Engineer