AI Platform Reliability Engineer (PRE) /Architect
HCLTech
Job Description
Location - Noida, Bengaluru, Chennai, Hyderabad and Pune.
Role
An AI PRE Engineer (L4) acts as a Platform Reliability Architect and Production Readiness Leader responsible for defining, governing, and scaling enterprise-grade AI/ML platforms. This role ensures AI systems are production-ready, resilient, observable, secure, compliant, and cost-efficient at scale, while driving standardization, automation, and platform-wide governance across AI engineering, SRE, DevOps, and MLOps disciplines. The L4 PRE acts as a strategic bridge between platform engineering, AI teams, and business leadership, ensuring reliability and scalability of AI solutions.
Responsibilities
Architecture & Production Readiness Governance
- Define and govern enterprise-wide production readiness frameworks across: Platform, data, model, application, and security layers
- Establish organization-wide PRE standards, policies, and maturity models
- Define and drive release certification governance (PRR frameworks) across all AI workloads
- Own production readiness lifecycle from design → deployment → scale
SLO/SLI Strategy & Reliability Engineering
- Define and govern: SLO/SLI/SLA frameworks for latency, availability, quality, safety, drift
- Error budget strategy at platform and business levels
- Establish reliability engineering models for AI systems
- Align SLOs with: Customer experience, business KPIs, and cost targets
Reference Architecture & Platform Design
- Define and publish: Enterprise AI reference architectures for: LLM applications, RAG pipelines, vector stores, agent frameworks
- Batch, real-time, and streaming inference systems
- Standardize: Platform architecture patterns across multi-cloud and hybrid environments
- Drive platform abstraction and reusable AI services
- Design and architect the NVIDIA Enterprise AI Software stack and deployment of GPU enabled Kubernetes clusters.
Deployment & Release Engineering Governance
- Define enterprise standards for: Safe deployment strategies (canary, shadow, blue-green, A/B testing)
- Prompt/model versioning, rollback, and release controls
- Govern: Release pipelines, CI/CD, GitOps, and policy-as-code frameworks
- Establish release risk assessment frameworks for AI workloads
Observability & AI Ops Strategy
- Define enterprise observability strategy covering: Metrics, logs, traces, and AI-specific telemetry: Token usage, latency, model drift, GPU utilization, quality metrics
- Standardize: Observability patterns for AI systems across platform
- Drive adoption of: AIOps and predictive reliability models
- Establish: SLO burn alerts, cost anomaly detection frameworks
Capacity Engineering & FinOps for AI
- Own: Capacity planning and forecasting models for: Token usage
- GPU/CPU scaling
- Vector databases and caching layers
- Define: Cost optimization strategies (FinOps for AI platforms)
- Govern: Resource allocation, autoscaling policies, and utilization efficiency
Resilience Engineering & System Design
- Define: Enterprise resilience patterns, including: Circuit breakers, fallbacks, retries, timeouts
- Semantic caching and prompt optimization
- Architect: High-availability AI systems with: Multi-region / multi-zone failover
- Establish: DR/BCP strategies with defined RTO/RPO targets
AI Security, Risk & Compliance Governance
- Define and govern: AI security frameworks: Secrets management
- Private networking
- Egress controls
- Model and prompt security
- Establish: Red-team validation and AI safety testing standards
- Own: Compliance frameworks mapping (ISO 27001, SOC 2, GDPR/DPDP, HIPAA)
- Ensure: Data privacy, auditability, and governance across AI platforms
Platform Engineering & Developer Enablement
- Define: Enterprise delivery frameworks: CI/CD pipelines
- SDKs, APIs, Helm charts, Terraform templates
- Drive: Platform standardization and self-service enablement
- Build: Reusable automation frameworks and PRE accelerators
- Lead: Enablement programs (trainings, brown-bags, playbooks)
Cross-Functional Leadership & Stakeholder Engagement
- Partner with: AI/ML teams, platform teams, DevOps, and business stakeholders
- Act as: Technical advisor to leadership and customers
- Validate: Architecture, NFRs (scale, latency, cost, safety) across programs
- Drive: Enterprise-wide adoption of PRE practices
Incident Management & Reliability Operations
- Own: Enterprise incident response frameworks for AI systems
- Lead: Critical incident command and escalations
- Govern: RCA, postmortems, and systemic reliability improvements
- Establish: Continuous improvement and
Qualifications & Experience
- Bachelor’s degree in Computer Science / Engineering (mandatory)
- Master’s preferred (Distributed Systems / Cloud / AI Systems)
- 12–18 years overall IT / platform engineering experience
- 7–10 years in AI / platform / cloud engineering at scale
- 5+ years in architecture / PRE / SRE strategy roles
Core Experience:
- Enterprise-scale:
- Production readiness frameworks, PRR governance
- SLO/SLI/SLA design and reliability engineering
- Deep expertise in:
- Kubernetes platforms (EKS/GKE/AKS) at scale
- GPU-aware AI infrastructure
- Multi-cluster, multi-tenant environments
- Strong experience in:
- MLOps / LLMOps / AI production systems
- Deployment strategies and release governance
Technology Experience:
- Cloud: AWS / Azure / GCP (architect-level understanding)
- Infra: Kubernetes, Docker, Terraform, GitOps
- Observability: Prometheus, Grafana, OpenTelemetry
- AI stack:
- LLM deployment, RAG pipelines, vector DBs
- Frameworks like LangChain, OpenAI
- NVIDIA AI Enterprise (NVAIE)
- Programming:
- Python / Go / Java (platform engineering focus)
Certifications Required
- NVIDIA Certified Professional – AI Infrastructure & Operations
- NVIDIA DLI – Advanced AI Infrastructure / GPU platforms
- Kubernetes Certifications (CKA / CKS – mandatory)
- Cloud Architect Certifications (AWS/Azure/GCP – Professional level preferred)
- DevOps / CI-CD certifications (preferred)
- Linux Certifications (RHCE / LFCS – advanced preferred)
Required Skills
Similar Jobs
View all →
DevOps Engineer
SecLogic.ai
DevOps Engineer
Raft
Ripik AI - Engineering Manager - Backend Architecture
Ripik.AI
DevOps Engineer
BitQcode Capital
Senior Devops Engineer Lead
Amura Health
Share
Quick Apply
Upload your resume to apply for this position