Pattern

Memory-Capped Model Inference

memory-capped-inference-oom

Text encoder dequantization and CPU offload still fail because embedding expansion demands large peak RAM, and cgroup/systemd or other environment constraints terminate the process before swap can help.