Find The RightJob.
OVERVIEW
We are hiring an ML Systems Engineer to design and deliver cutting-edge AI solutions for enterprise
clients at the frontier of agentic AI, inference engineering, and ML systems architecture. You will go
beyond applied ML - dissecting how AI systems are built, optimized, and scaled - designing production-
grade architectures spanning retrieval systems, inference pipelines, and agentic workflows. You will
translate state-of-the-art capabilities into robust, performant solutions, operating at the intersection of ML
research awareness and engineering discipline.
KEY RESPONSIBILITIES
LLM inference pipelines, and retrieval-augmented architectures.
cache management, batching strategies, and memory optimization - alongside ML performance
engineering to profile bottlenecks, benchmark throughput/latency, and evaluate quantization
strategies (GPTQ, AWQ, GGUF).
ranking through to multi-agent orchestration, tool use, and memory architectures.
make principled architectural trade-off decisions for client contexts.
encoding best practices across engagements.
assessments - in client-facing contexts.
TECHNICAL QUALIFICATIONS
Core Requirements
on with the PyTorch ecosystem including Hugging Face Transformers, PEFT, Accelerate, and
Datasets.
hands-on with at least one inference runtime (vLLM, TGI, TensorRT-LLM, SGLang, or similar).
working with vLLM for high-performance LLM serving, optimization, and large-scale inference.
using LangGraph, LlamaIndex Workflows, AutoGen, or CrewAI - including tool use, memory, and
multi-agent coordination.
provider SDK usage across OpenAI, Anthropic, Mistral, Hugging Face, and similar.
packaging, deploying, and debugging AI systems.
and productivity enhancement.
Preferred
of when fine-tuning is the right lever vs. prompting or RAG.
knowledge of C++ or Rust.
LangSmith, Arize, W&B, Phoenix, or Prometheus/Grafana for ML observability.
WAYS TO STAND OUT FROM THE CROWD
design decisions that only emerge at runtime.
custom batching logic, or optimizing a quantization pipeline to hit real SLAs.
llama.cpp, TGI, LangChain, LlamaIndex, or similar - demonstrating work that holds up to community
scrutiny.
domain-specific evals, or human-in-the-loop feedback loops - beyond off-the-shelf metrics.
context retrieval, or inference hardware tradeoffs based on hands-on experimentation.
Pay: ₹500,000.00 - ₹1,800,000.00 per year
Benefits:
Work Location: In person
Similar jobs
Infosys
India
about 1 month ago
Citi
Chennai, India
about 1 month ago
Mindsprint
Chennai, India
about 1 month ago
NTT DATA
India
about 1 month ago
OneMagnify
Chennai, India
about 1 month ago
Programming.com
India
about 1 month ago
App Innovation Technologies
India
about 1 month ago
© 2026 Qureos. All rights reserved.