gpu-telemetry. GPU Observability with workload attribution. One OTLP agent per node ties hardware metrics (NVIDIA, AMD, Intel Gaudi) to the K8s pod or Slurm job burning the GPU.

github.com/last9/gpu-telemetry

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.