Pattern
Memory-Capped Model Inference
memory-capped-inference-oom
Text encoder dequantization and CPU offload still fail because embedding expansion demands large peak RAM, and cgroup/systemd or other environment constraints terminate the process before swap can help.