Idle GPU Policy and Guidance

Idle GPU Policy and Guidance

To maintain efficiency and resource availability for all users, UFIT Research Computing has a policy prohibiting the allocation of GPUs and leaving them idle (see paragraph 2 of the Scheduler/Job policy).

Automated systems are now in place to terminate jobs where GPU utilization falls below a defined threshold. Although this threshold is subject to change, it is currently set to 0% utilization for 1 hour. We recognize that adhering to this policy may require users to adjust their workflows; however, high-demand resources such as the B200 and RTX Pro 6000 GPUs are highly valuable and too scarce to remain inactive for extended periods.

FAQ

I have a workflow that needs a GPU for part of the time, but has other parts that do not use the GPU. How can I run this on HiPerGator given this policy?

We suggest breaking these workflows into separate jobs that form a chain, where the next job is submitted when a job in the chain finishes: GPU processing in one job and CPU-only processing in another. This can be automated using Slurm’s dependency flag for sbatch.

My LLM workload involves stages where inference is performed and then evaluated, during the evaluation stage, there is no GPU load. What can I do to prevent my jobs from being canceled?

For this type of workload, we suggest cancelling your LLM inference job during the evaluation stage.

Re-working your workflow to batch inference steps may be able to make this more efficient. You may consider exploring Slurm's sbatch job dependency options. This would allow, for example, submitting a GPU job to start the LLM service, and a dependant job to run inference. Since it may take some time for the GPU job to start, using -d after:job_id for the inference job would allow that to start once the service has started. Multiple inference jobs could be run, with the last one shutting down the LLM service job.

I launch an LLM inference service, user requests come in at unpredictable rates. How can I make this work?

HiPerGator should not be used as an interactive LLM inference server. Please explore NaviGator Toolkit for this use.

I work interactively. How can I prevent my job from being canceled?

We suggest that interactive users cancel jobs/sessions when they are not working and request a new job/session when they return.

How can I get help managing GPU use better?

Please open a support request, and we will gladly assist you.

Does this policy apply to the L4 GPUs?

Yes, though at this time, we are not canceling jobs on those GPUs.

How can I monitor GPU utilization?

The jobnvtop command can be used (see here).

How does this apply to multi-GPU jobs?

If any GPU allocated to a job is idle, the job will be terminated.

How about if I run something just to keep the GPU from being idle?

Nice try, but no!

HiPerGator GPUs are heavily subsidized by the University and the State of Florida. It is our responsibility to ensure the resources are used efficiently. While the policy has helped reduce queue times and improve throughput, some users have found ways to run fake workloads to intentionally violate the policy.

This is a willful act of circumventing UFIT-RC policy and tying up valuable research hardware. University and UFIT-RC acceptable use policies are in place to ensure compliance with regulations and efficient use of university resources. Users who intentionally run fake workloads or otherwise actively circumvent enforcement mechanisms will have their accounts suspended for 2 weeks and will be required to have their sponsor acknowledge the violation and take responsibility for preventing a recurrence. Any further violations will be escalated appropriately. For students, this includes referral to the SCCR (Student Conduct and Conflict Resolution). For faculty and staff, this includes referral to the appropriate Chair, Dean or HR specialist. For non-UF users, this includes a permanent ban from accessing HiPerGator and referral to the appropriate authorities.