NVIDIA
Deep Learning Algorithms Engineering Intern, Dynamo
- Designed Dynamo's sidecar architecture across vLLM, SGLang, and TensorRT-LLM, preserving native engine entrypoints while adding out-of-process routing, lifecycle management, and distributed orchestration.
- Contributed upstream to the vLLM, SGLang, and TensorRT-LLM open-source projects while co-designing sidecar lifecycle, failure-recovery, multimodal, and KV-cache contracts with each engine team.
- Created and open-sourced OpenEngine, a vendor-neutral gRPC/Protobuf protocol defining typed APIs for inference, discovery, health, abort, drain, and KV coordination across engines and distributed frameworks.