Agentic AI Automation for IT Infrastructure:…
When deterministic playbooks meet probabilistic reasoning
Jatin Gupta Eighteen years building and running the systems enterprises depend on to know when something is broken — monitoring, event correlation, incident automation — across telecom, BFSI, retail, and manufacturing. The last three of those years have gone into a specific question: which parts of that work should an LLM do, and which parts should stay deterministic. Below is the path that got me here.
Experience
My Professional Resume
I design agentic AI systems for IT infrastructure automation. In practice: agent workflows that ingest signals from observability and ITSM platforms, reason over them with an LLM, and hand execution to deterministic automation — BigFix, Ansible, Terraform, Jenkins, scripted runbooks — rather than generating commands at inference time. Human approval gates sit in front of anything with production blast radius. The design decision I defend most often is that split. Using the model for triage, summarization, and routing while keeping remediation deterministic trades away flexibility and buys a bounded failure mode: a bad classification costs a misroute, a bad generated command costs an outage. Retrieval over runbooks and historical incidents grounds the reasoning layer, and cases without supporting evidence abstain and escalate rather than guess. Prompts are versioned and regression-tested against a held-out incident set before anything ships. Alongside that, I architect BigFix-based solutions across the HCL software portfolio — endpoint lifecycle and patch management, Runbook AI automation, cloud lifecycle management, ITSM integration, job scheduling, and reporting and business observability. I position the suite in customer engagements and then extend it, building agentic automation on top of the BigFix APIs so the agent layer has a real execution surface rather than a demo one. The foundation under all of it is a consolidated signal plane: Splunk, Dynatrace, Datadog, New Relic, OpsRamp, Grafana, and ServiceNow, stitched together with custom event-correlation pipelines in place of fragmented legacy tooling. Alert noise reaching on-call dropped by [40%], and recurring manual toil. I also lead solution architecture and pre-sales for these engagements — discovery, technical proposal, demo, delivery handoff — run release and upgrade programs for live customer deployments, and review designs before they reach the customer.
The work I'd point to first: I architected and built a greenfield monitoring platform on Splunk Enterprise for Accenture's internal managed-services practice. Custom app and dashboard development, index and data-onboarding design, the reference architecture, multi-tenant onboarding, integration into the shared NOC. Then I ran the enhancement roadmap that turned it from one team's platform into a standard managed-service offering used across accounts. Building it multi-tenant from the start was the decision that made the difference. It's the reason the platform could be handed to new accounts without a rebuild each time. Beyond that, I architected observability and AIOps platforms for Fortune 500 BFSI and telecom clients across infrastructure, database, cloud, and application layers on Splunk, Dynatrace, and AppDynamics. I owned the proposal-to-delivery loop end to end — wrote the technical proposal, defended it in the bid, then led the team that built it, which is the only reliable way I know to keep what sales promises and what engineering ships in the same room. I also replaced siloed monitoring tools with correlated event pipelines so on-call attention went to real incidents rather than duplicates, and built an internal documentation and runbook system adopted across delivery teams. That corpus later became retrieval grounding for the agent workflows I build now, which was not the plan at the time.
Owned monitoring platform delivery on AppDynamics, CA Nimsoft, CA Spectrum, Cisco Prime, and Netmapper for enterprise clients, and served as primary technical contact through implementation and support. Ran multi-disciplinary teams through scope, design, and rollout — project planning, risk review, customer approvals. SLA closure times improved by roughly 20%. I also wrote the delivery playbooks that became the team standard for how monitoring engagements get scoped and shipped.
Ran the monitoring stack for Vodafone India's IT estate: HP OVO, HP NNM, HP OVPI, BAC/BSM. Standing member of the Change Advisory Board, reviewing every change that touched production monitoring before it shipped, and an RCA member for alarm optimization work. The one I still tell people about: I diagnosed and fixed root-cause defects in HP OVO that vendor support couldn't reproduce, then designed and rolled out internal enhancements on top of the platform. I also defined the technical requirements and access and security models for monitoring tooling across the org.
- Provided support for clients' infra monitoring platform across multiple regions utilizing HP OVO suite. - Efficiently managed and maintained the existing setup, ensuring minimal downtime for clients. - Conducted thorough evaluations of root causes and implemented fixes using vendor-provided patches. - Collaborated with team members to optimize system performance and streamline processes. Skills: HP Openview
- Created detailed business requirement documents to guide technical specifications development, ensuring accuracy and consistency throughout the process. - Fostered strong client relationships by serving as the primary point of contact for product knowledge and implementation efforts, ultimately driving customer satisfaction and retention. - Provided technical expertise to solve complex software engineering problems, utilizing creativity and ingenuity to develop innovative solutions. - Developed technical Pre-Sales solutions and collaborated closely with the Sales team to effectively communicate proposal knowledge and win new business opportunities.
- Managed a team of 5-10 members while effectively collaborating with application domain experts within established process framework to develop fundamental web development skills. - Demonstrated proficiency in various systems including CA Spectrum 8.0, HP OpenView NNM 7, CA eHealth 6.0, HP OVPI 7, CA NSM 3.1, CA ServiceDesk r11, MS SQL, and Veritas Netbackup. - Gathered user needs and worked with technical resources to develop, test, and document software while providing technical assistance to other developers. - Coordinated with functional users and IT staff to find solutions to problems identified during testing, resolved issues during systems upgrades, and ensured easy translation of requirements documentation into UAT. Skills: HP Openview · NNMi
- Conducted thorough research to identify opportunities for improving efficiency of IT applications and techniques in the industry - Successfully managed the network for 30-40 branches and maintained LAN for 150-200 systems - Collaborated with clients to gather requirements, conduct system analysis, and finalize technical specifications and high level design documents for projects - Demonstrated expertise in network printers, active directory, SQL server installation and maintenance, and Windows NT and 2003 server - Managed mail servers, backup servers, and program packages on database server with Veritas backup and Ultrim drive Autoloader.
We share our news and blog
When deterministic playbooks meet probabilistic reasoning
Monitoring vs. Observability: Unveiling the Power of Modern System Insights
In this blog, we will discuss what a AppDynamics Observability is, its benefits, and how it works.
Need Some Help?