Catalog
ml-ops skills
1 skills. Every one is readable here; downloads and live use need a subscription.
remote-gpu-trainer
Deploy, monitor, and debug long GPU jobs on RENTED/remote instances (AutoDL, RunPod, vast.ai, Lambda, Slurm, K8s): teardown/billing safety, spot resilience, resumable checkpointing, OOM/NaN triage.
