Cluster observability
Prometheus, Grafana, and Slurm telemetry across large GPU networks — with auto-deployed observability runners and agents.
Senior System Software Designer · AMD India
HPC Development Engineer with 9+ years of R&D in system software, GPU computing, and distributed systems — specializing in AI/ML workloads on accelerators and cluster observability.
Current role
Prometheus, Grafana, and Slurm telemetry across large GPU networks — with auto-deployed observability runners and agents.
Obsidian-based agent memory index, knowledge base, and skills auto-sync for humans and agents.
Secure cluster validation log shipping to AWS S3, integrated into CI for heterogeneous systems.
MI2XX / MI3XX benchmarking, PyTorch workloads, HIP/ROCm profiling, RAS injection, and SMI metrics.
Selected projects
Industry-standard monitoring stacks, agent auto-deploy, and S3 telemetry for cluster validation CI.
Memory index and knowledge base with skills sync optimized for human- and agent-readable use.
Context-aware issue routing with FAISS semantic search, GitHub/Jira connectors, and LLM APIs.
Multi-tool skills, agents, and hooks (Claude, Cursor) with security and token spend controls.
Recognition