nvidia-smi can show a healthy GPU while an Ollama request still runs partly—or entirely—on CPU. That contradiction does not have one universal cause. The loaded model's PROCESSOR field is the starting evidence: it separates full GPU placement, full CPU placement and a mixed split before you touch ...