ishan

i currently work on

My focus is primarily on efficient LLM routing policies and optimizing both the orchestration and engine for RL and agentic inference. You can read more about this here

I am also the author and maintainer of srt-slurm which provides a k8s style deployment experience on SLURM. It is used extensively for benchmarking inside and outside of NVIDIA.

previously

more of me