Senior System Software Designer · AMD India

Gahan
Saraiya

HPC Development Engineer with 9+ years of R&D in system software, GPU computing, and distributed systems — specializing in AI/ML workloads on accelerators and cluster observability.

Building reliable GPU fleets at scale

Cluster observability

Prometheus, Grafana, and Slurm telemetry across large GPU networks — with auto-deployed observability runners and agents.

Agentic memory network

Obsidian-based agent memory index, knowledge base, and skills auto-sync for humans and agents.

Telemetry pipelines

Secure cluster validation log shipping to AWS S3, integrated into CI for heterogeneous systems.

GPU performance R&D

MI2XX / MI3XX benchmarking, PyTorch workloads, HIP/ROCm profiling, RAS injection, and SMI metrics.

Recent engineering

2026

Cluster Management and Observability

Industry-standard monitoring stacks, agent auto-deploy, and S3 telemetry for cluster validation CI.

2026

Obsidian Agentic Memory Network

Memory index and knowledge base with skills sync optimized for human- and agent-readable use.

2025

LangGraph Task Routing & Triage

Context-aware issue routing with FAISS semantic search, GitHub/Jira connectors, and LLM APIs.

2025

Agentic AI for Engineering Workflows

Multi-tool skills, agents, and hooks (Claude, Cursor) with security and token spend controls.

See full portfolio, talks, and certifications →

Highlights

Resumes by focus

Canonical profile is R&D. Variants tailored for ML/HPC, DevOps, and firmware roles.