Bestkaam Logo
Back to Jobs
HCLTech

AI Platform Reliability Engineer (PRE) /Architect

Actively Reviewing

HCLTech

Noida Full-Time 4–8 yrs exp Posted 7 hours ago  · Apply by Sep 22, 2026

Location - Noida, Bengaluru, Chennai, Hyderabad and Pune.


Role


An AI PRE Engineer (L4) acts as a Platform Reliability Architect and Production Readiness Leader responsible for defining, governing, and scaling enterprise-grade AI/ML platforms. This role ensures AI systems are production-ready, resilient, observable, secure, compliant, and cost-efficient at scale, while driving standardization, automation, and platform-wide governance across AI engineering, SRE, DevOps, and MLOps disciplines. The L4 PRE acts as a strategic bridge between platform engineering, AI teams, and business leadership, ensuring reliability and scalability of AI solutions.



Responsibilities


Architecture & Production Readiness Governance


  • Define and govern enterprise-wide production readiness frameworks across: Platform, data, model, application, and security layers
  • Establish organization-wide PRE standards, policies, and maturity models
  • Define and drive release certification governance (PRR frameworks) across all AI workloads
  • Own production readiness lifecycle from design → deployment → scale


SLO/SLI Strategy & Reliability Engineering


  • Define and govern: SLO/SLI/SLA frameworks for latency, availability, quality, safety, drift
  • Error budget strategy at platform and business levels
  • Establish reliability engineering models for AI systems
  • Align SLOs with: Customer experience, business KPIs, and cost targets


Reference Architecture & Platform Design


  • Define and publish: Enterprise AI reference architectures for: LLM applications, RAG pipelines, vector stores, agent frameworks
  • Batch, real-time, and streaming inference systems
  • Standardize: Platform architecture patterns across multi-cloud and hybrid environments
  • Drive platform abstraction and reusable AI services
  • Design and architect the NVIDIA Enterprise AI Software stack and deployment of GPU enabled Kubernetes clusters.


Deployment & Release Engineering Governance


  • Define enterprise standards for: Safe deployment strategies (canary, shadow, blue-green, A/B testing)
  • Prompt/model versioning, rollback, and release controls
  • Govern: Release pipelines, CI/CD, GitOps, and policy-as-code frameworks
  • Establish release risk assessment frameworks for AI workloads


Observability & AI Ops Strategy


  • Define enterprise observability strategy covering: Metrics, logs, traces, and AI-specific telemetry: Token usage, latency, model drift, GPU utilization, quality metrics
  • Standardize: Observability patterns for AI systems across platform
  • Drive adoption of: AIOps and predictive reliability models
  • Establish: SLO burn alerts, cost anomaly detection frameworks


Capacity Engineering & FinOps for AI


  • Own: Capacity planning and forecasting models for: Token usage
  • GPU/CPU scaling
  • Vector databases and caching layers
  • Define: Cost optimization strategies (FinOps for AI platforms)
  • Govern: Resource allocation, autoscaling policies, and utilization efficiency


Resilience Engineering & System Design


  • Define: Enterprise resilience patterns, including: Circuit breakers, fallbacks, retries, timeouts
  • Semantic caching and prompt optimization
  • Architect: High-availability AI systems with: Multi-region / multi-zone failover
  • Establish: DR/BCP strategies with defined RTO/RPO targets


AI Security, Risk & Compliance Governance


  • Define and govern: AI security frameworks: Secrets management
  • Private networking
  • Egress controls
  • Model and prompt security
  • Establish: Red-team validation and AI safety testing standards
  • Own: Compliance frameworks mapping (ISO 27001, SOC 2, GDPR/DPDP, HIPAA)
  • Ensure: Data privacy, auditability, and governance across AI platforms


Platform Engineering & Developer Enablement


  • Define: Enterprise delivery frameworks: CI/CD pipelines
  • SDKs, APIs, Helm charts, Terraform templates
  • Drive: Platform standardization and self-service enablement
  • Build: Reusable automation frameworks and PRE accelerators
  • Lead: Enablement programs (trainings, brown-bags, playbooks)


Cross-Functional Leadership & Stakeholder Engagement


  • Partner with: AI/ML teams, platform teams, DevOps, and business stakeholders
  • Act as: Technical advisor to leadership and customers
  • Validate: Architecture, NFRs (scale, latency, cost, safety) across programs
  • Drive: Enterprise-wide adoption of PRE practices


Incident Management & Reliability Operations


  • Own: Enterprise incident response frameworks for AI systems
  • Lead: Critical incident command and escalations
  • Govern: RCA, postmortems, and systemic reliability improvements
  • Establish: Continuous improvement and


Qualifications & Experience


  • Bachelor’s degree in Computer Science / Engineering (mandatory)
  • Master’s preferred (Distributed Systems / Cloud / AI Systems)
  • 12–18 years overall IT / platform engineering experience
  • 7–10 years in AI / platform / cloud engineering at scale
  • 5+ years in architecture / PRE / SRE strategy roles


Core Experience:


  • Enterprise-scale:
  • Production readiness frameworks, PRR governance
  • SLO/SLI/SLA design and reliability engineering
  • Deep expertise in:
  • Kubernetes platforms (EKS/GKE/AKS) at scale
  • GPU-aware AI infrastructure
  • Multi-cluster, multi-tenant environments
  • Strong experience in:
  • MLOps / LLMOps / AI production systems
  • Deployment strategies and release governance


Technology Experience:


  • Cloud: AWS / Azure / GCP (architect-level understanding)
  • Infra: Kubernetes, Docker, Terraform, GitOps
  • Observability: Prometheus, Grafana, OpenTelemetry
  • AI stack:
  • LLM deployment, RAG pipelines, vector DBs
  • Frameworks like LangChain, OpenAI
  • NVIDIA AI Enterprise (NVAIE)
  • Programming:
  • Python / Go / Java (platform engineering focus)


Certifications Required


  • NVIDIA Certified Professional – AI Infrastructure & Operations
  • NVIDIA DLI – Advanced AI Infrastructure / GPU platforms
  • Kubernetes Certifications (CKA / CKS – mandatory)
  • Cloud Architect Certifications (AWS/Azure/GCP – Professional level preferred)
  • DevOps / CI-CD certifications (preferred)
  • Linux Certifications (RHCE / LFCS – advanced preferred)