Senior System Software Designer
AMD India Pvt Ltd · Mar 2024 – Present- Cluster Observability: Prometheus, Grafana, Slurm telemetry; auto-deploy observability runners and agents for large GPU fleets.
- Agentic Memory Network: Obsidian design for agent memory index, knowledge base, and skills auto-sync (human-readable and agent-readable).
- Telemetry Pipeline: Secure log upload and cluster validation telemetry shipping to AWS S3.
- R&D Leadership: Next-generation AMD GPU accelerators (MI2XX, MI3XX series).
- Performance Engineering: GPU benchmarking with modern C++ and Python.
- Deep Learning: PyTorch on large-scale GPU clusters — GEMM and distributed training pipelines.
- HIP/ROCm: Compute partitioning, HIP kernel profiling, RAS injection, parallel execution.