Skip to main content
LLM
Toolkit
Search
/
Home
/
Troubleshooting
Troubleshooting
Common failures and how to fix them. Select your symptom to get targeted guidance.
WHAT WENT WRONG?
◈
CUDA Out of Memory
Training crashes with memory errors
⏱
Training Too Slow
Low throughput, GPU underutilized
▲
Loss is NaN
Loss becomes NaN and training breaks
▽
Loss Not Decreasing
Training loss stays flat or increases
▣️
GPU Not Detected
torch.cuda.is_available() returns False
◇
Dataset Error
Data loading fails or produces bad batches
▦
Checkpoint Error
Cannot save or load checkpoints
⬡
Distributed Training Failure
NCCL errors, scaling issues
◐
Poor Evaluation
Model scores lower than expected
○
Model Outputs Are Bad
Generated text is incoherent or off-target