ML Engineer

Rahul Sharma

I build and fine-tune vision-language models, retrieval pipelines, and agentic systems — and ship them into production.

Rahul Sharma

Hey, I'm Rahul — an ML engineer. I work across the stack that turns models into shipped products: fine-tuning vision-language models, building retrieval and RAG pipelines, and wiring up agentic systems that hold up in production.

I'm currently an SDE-2 AI/ML Engineer at ZysecAI, focused on large-scale document understanding — OCR over huge unstructured corpora, the fine-tuning behind it, and the retrieval and agent plumbing that makes it trustworthy.

Before that I did undergraduate research on biomedical visual question answering (the basis of our IJETS paper), plus NLP work in finance and EdTech. I studied CS at VTU, and I write here to think out loud about what I build.

Selected Writing

All articles →
Aug 30, 2026
Article
Serving LLMs: what actually decides your latency
TTFT, prefill vs decode, the KV-cache bill, quantization, multi-LoRA serving — and the runtime + hardware matrix behind them.