Bestkaam Logo
Back to Jobs
Unknown

AI/ML Engineer (LLMOps & Model Deployment)

Actively Reviewing

Unknown

India Contract 4–8 yrs exp Posted 4 days ago  · Apply by Sep 22, 2026

Key Responsibilities

  • Model Deployment & API Development:Host & Scale Open-Weights Models: Deploy and maintain state-of-the-art open-weights models (e.g., Llama, Mistral, Phi) on Azure infrastructure.
  • API Engineering: Design, build, and secure high-performance, low-latency REST/gRPC APIs (using frameworks like FastAPI) to serve models to downstream cloud applications.
  • Inference Optimization: Implement advanced optimization techniques (e.g., vLLM, TensorRT-LLM, DeepSpeed) to maximize throughput and minimize time-to-first-token (TTFT).
  • LLMOps & Infrastructure: Pipeline Automation: Build and manage end-to-end LLMOps pipelines for continuous integration, deployment, and monitoring of models.
  • Azure Cloud Architecture: Leverage Azure AI Studio, Azure Machine Learning (Azure ML), Azure Kubernetes Service (AKS), and managed compute (GPUs like A100/H100) efficiently.
  • Monitoring & Observability: Implement comprehensive logging, tracing, and evaluations for LLM outputs (tracking drift, latency, costs, and hallucination rates).
  • Model Efficiency & Distillation:Knowledge Distillation: Train smaller, task-specific student models from larger, high-performing teacher models to reduce operational costs and latency without sacrificing accuracy.
  • Quantization & Fine-Tuning: Apply quantization techniques (AWQ, GPTQ, GGUF) and parameter-efficient fine-tuning (PEFT/LoRA) to adapt models to specific business domains.

Required Skills & Qualifications

  • Experience: 4+ years of professional experience as an ML Engineer, Data Scientist, or Backend Engineer with a heavy focus on AI deployment.
  • Cloud Proficiency: Strong hands-on experience with Microsoft Azure (Azure ML, AKS, Azure Container Apps, Key Vault).
  • AI/ML Frameworks: Deep proficiency with PyTorch, Hugging Face ecosystem (Transformers, Accelerate), and LangChain or LlamaIndex.
  • API & Backend: Strong Python programming skills and experience with containerization (Docker, Kubernetes) and API development (FastAPI).
  • Model Optimization: Demonstrated experience with model distillation, pruning, quantization, and utilizing inference engines like vLLM or TGI.
  • Soft Skills & Culture Fit
  • Problem Solver: Ability to triage infrastructure bottlenecks, memory constraints (OOM errors), and CUDA-related issues independently.
  • Collaborator: Comfortable working cross-functionally with backend engineers, product managers, and security teams.
  • Cost-Conscious Mindset: A sharp focus on balancing model accuracy with cloud spend and compute efficiency.

Preferred Qualifications

  • Certifications such as Azure AI Engineer Associate or Azure Solutions Architect.
  • Experience implementing secure guardrails (e.g., NeMo Guardrails, Azure AI Content Safety). Contributions to open-source ML/LLMOps projects.