Wednesday at 03:18 AM2 days I have an Nvidia RTX 2070 Super in my server.I also have the Dynamic S3 Sleep Plugin (this plugin is not responsible for the issue).When I put my server to sleep though the UI then wake it up, CUDA crashes within milliseconds, containers that leverage AI either stop working or start using the CPU+System ram.I have replicated this with both the open source driver and the latest driver.When I enter dmesg | grep -B20 -A20 "Xid" in the Unraid terminal, this is the output:dmesg | grep -B20 -A20 "Xid"[ 1820.752171] ata6.00: revalidation failed (errno=-5)[ 1820.752190] ata1.00: qc timeout after 5000 msecs (cmd 0xec)[ 1820.752200] ata1.00: failed to IDENTIFY (I/O error, err_mask=0x4)[ 1820.752203] ata1.00: revalidation failed (errno=-5)[ 1824.032060] ata4: SATA link up 6.0 Gbps (SStatus 133 SControl 300)[ 1824.103300] ata4.00: configured for UDMA/133[ 1824.239854] ata6: SATA link up 6.0 Gbps (SStatus 133 SControl 300)[ 1824.307231] ata6.00: configured for UDMA/133[ 1824.743864] ata1: SATA link up 6.0 Gbps (SStatus 133 SControl 300)[ 1824.810567] ata1.00: configured for UDMA/133[ 1825.616118] ata2: SATA link up 6.0 Gbps (SStatus 133 SControl 300)[ 1825.689735] ata2.00: configured for UDMA/133[ 1844.924052] docker0: port 3(veth791c2cd) entered blocking state[ 1844.924059] docker0: port 3(veth791c2cd) entered disabled state[ 1844.924067] veth791c2cd: entered allmulticast mode[ 1844.924159] veth791c2cd: entered promiscuous mode[ 1844.929874] eth0: renamed from vethf7a46f8[ 1844.930313] docker0: port 3(veth791c2cd) entered blocking state[ 1844.930319] docker0: port 3(veth791c2cd) entered forwarding state[ 1852.706530] NVRM: GPU at PCI:0000:08:00: GPU-9619125e-6ae1-e026-121a-d310f1334081[ 1852.706536] NVRM: Xid (PCI:0000:08:00): 31, pid=6298, name=modprobe, channel 0x08000005, intr 00000000. MMU Fault: ENGINE HOST9 HUBCLIENT_HOST faulted @ 0x1_21010000. Fault is of type FAULT_PDE ACCESS_TYPE_VIRT_READ[ 1852.706742] NVRM: nvGpuOpsReportFatalError: uvm encountered global fatal error 0x60, requiring os reboot to recover.[ 1852.708980] NVRM: Xid (PCI:0000:08:00): 154, GPU recovery action changed from 0x0 (None) to 0x2 (Node Reboot Required)[ 1854.008687] NVRM: Xid (PCI:0000:08:00): 31, pid=6298, name=modprobe, channel 0x09000007, intr 00000000. MMU Fault: ENGINE HOST10 HUBCLIENT_HOST faulted @ 0x1_21070000. Fault is of type FAULT_PDE ACCESS_TYPE_VIRT_READ[ 1854.141606] docker0: port 3(veth791c2cd) entered disabled state[ 1854.141693] vethf7a46f8: renamed from eth0[ 1854.158438] docker0: port 3(veth791c2cd) entered disabled state[ 1854.159249] veth791c2cd (unregistering): left allmulticast mode[ 1854.159254] veth791c2cd (unregistering): left promiscuous mode[ 1854.159260] docker0: port 3(veth791c2cd) entered disabled state[ 1857.029693] docker0: port 3(vethfd4a51e) entered blocking state[ 1857.029700] docker0: port 3(vethfd4a51e) entered disabled state[ 1857.029710] vethfd4a51e: entered allmulticast mode[ 1857.029843] vethfd4a51e: entered promiscuous mode[ 1857.035564] eth0: renamed from vethb56fa0b[ 1857.035994] docker0: port 3(vethfd4a51e) entered blocking state[ 1857.035999] docker0: port 3(vethfd4a51e) entered forwarding stateDiagnostics attached.tower-diagnostics-20260729-0455.zip Edited Wednesday at 03:20 AM2 days by Stubbs
Wednesday at 07:01 PM2 days Community Expert Honestly its probably the nvidia driver not playing nice with sleep. Not really much anyone can do about it.Sleep is also not officially supported by unraid either.
Join the conversation
You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.