Back to Jobs
AI/ML Engineer (LLMOps & Model Deployment)
Actively Reviewing
Unknown
Job Description
Key Responsibilities
- Model Deployment & API Development:Host & Scale Open-Weights Models: Deploy and maintain state-of-the-art open-weights models (e.g., Llama, Mistral, Phi) on Azure infrastructure.
- API Engineering: Design, build, and secure high-performance, low-latency REST/gRPC APIs (using frameworks like FastAPI) to serve models to downstream cloud applications.
- Inference Optimization: Implement advanced optimization techniques (e.g., vLLM, TensorRT-LLM, DeepSpeed) to maximize throughput and minimize time-to-first-token (TTFT).
- LLMOps & Infrastructure: Pipeline Automation: Build and manage end-to-end LLMOps pipelines for continuous integration, deployment, and monitoring of models.
- Azure Cloud Architecture: Leverage Azure AI Studio, Azure Machine Learning (Azure ML), Azure Kubernetes Service (AKS), and managed compute (GPUs like A100/H100) efficiently.
- Monitoring & Observability: Implement comprehensive logging, tracing, and evaluations for LLM outputs (tracking drift, latency, costs, and hallucination rates).
- Model Efficiency & Distillation:Knowledge Distillation: Train smaller, task-specific student models from larger, high-performing teacher models to reduce operational costs and latency without sacrificing accuracy.
- Quantization & Fine-Tuning: Apply quantization techniques (AWQ, GPTQ, GGUF) and parameter-efficient fine-tuning (PEFT/LoRA) to adapt models to specific business domains.
Required Skills & Qualifications
- Experience: 4+ years of professional experience as an ML Engineer, Data Scientist, or Backend Engineer with a heavy focus on AI deployment.
- Cloud Proficiency: Strong hands-on experience with Microsoft Azure (Azure ML, AKS, Azure Container Apps, Key Vault).
- AI/ML Frameworks: Deep proficiency with PyTorch, Hugging Face ecosystem (Transformers, Accelerate), and LangChain or LlamaIndex.
- API & Backend: Strong Python programming skills and experience with containerization (Docker, Kubernetes) and API development (FastAPI).
- Model Optimization: Demonstrated experience with model distillation, pruning, quantization, and utilizing inference engines like vLLM or TGI.
- Soft Skills & Culture Fit
- Problem Solver: Ability to triage infrastructure bottlenecks, memory constraints (OOM errors), and CUDA-related issues independently.
- Collaborator: Comfortable working cross-functionally with backend engineers, product managers, and security teams.
- Cost-Conscious Mindset: A sharp focus on balancing model accuracy with cloud spend and compute efficiency.
Preferred Qualifications
- Certifications such as Azure AI Engineer Associate or Azure Solutions Architect.
- Experience implementing secure guardrails (e.g., NeMo Guardrails, Azure AI Content Safety). Contributions to open-source ML/LLMOps projects.
Required Skills
Similar Jobs
View all →
Data Scientist
Valtech
Bengaluru
Deep Learning
Retrieval-Augmented Generation
Vertex AI
+39
Artificial Intelligence Engineer
Rosemallow Technologies Pvt Ltd
Coimbatore
Machine Learning
Retrieval-Augmented Generation
REST API
+18
Material+ - Lead AI Engineer
Srijan: Now Material
Gurugram
Machine Learning
Design patterns
Adobe Illustrator
+16
AI Engineer - Agentic
Srijan: Now Material
Gurugram
Retrieval-Augmented Generation
Machine Learning
REST API
+18
Lead AI Engineer
Srijan: Now Material
Gurugram
Machine Learning
Design patterns
Adobe Illustrator
+15
Share
Quick Apply
Upload your resume to apply for this position
–