September 24Sep 24 Hello, happy Unraiders!I’m trying to track down a recurring stability issue that started recently on a server that had otherwise been stable for several months.The system becomes partially or completely unresponsive after anywhere from roughly 1 to 4 days of uptime. WebGUI, SSH and normal services stop responding. On at least one occasion a VM was still alive and able to send a notification while the Unraid host itself was unreachable.Persistent syslog is enabled and archived externally. Unfortunately, there is usually nothing useful immediately before the failure — logging simply stops after ordinary Docker/service messages.When the machine is dead and I press the power button the screen shows that the shutdown procedure starts but it never finishes.Unraid 7.3.2 / kernel 6.18.38hangs every ~1–4 dayshost WebUI/SSH/services dieVM can sometimes remain alivesyslog often ends with nothing abnormalbad HDD replacedold RAM removedentirely new RAM kit still reproduces the hangboot USB corruption happened after some of the forced resetsCurrent platform:ASUS TUF GAMING B550M-PLUSRyzen 7 5700GLSI SAS2308 HBA, firmware 20.00.06.00Unraid 7.3.2 / kernel 6.18.38-Unraid2* 16 GBFirst fail on 28th August killed the boot drive (not bootable and not visible in BIOS) so I fixed it - didn't stick so I changed my boot drive to a new SanDisk.I had third RAM stick there, but it's out as a suspect - didn't help.I swapped old 2*16 GB kit for a new one - didn't help.One drive was turning bad so I swapped it - didn't help. Some other drives have some historical SMART warnings.In the log you can see that today I started monitoring health of the docker system (unrelated project). I am just parsing "docker info" in n8n and reading the warnings section. You can see that this happened:Normally: SSH connects, authenticates, runs the health check, exits within seconds.Around 11:54, one session suddenly takes tens of seconds.Around 12:06, multiple SSH connections begin piling up before authentication/session handling completes.Around 12:07, the system partially recovers — exactly the hiccup you noticed.Subsequent checks work again briefly.At 12:09:52–53, the final health-check connection gets as far as authentication/session start.Then the host essentially stops making progress until you reset it.The diagnostic was captured during startup so the load might seem high but it usually hovers much lower. I might try to rollback to 7.3.1. I dont recall when I upgraded but I usually dont do first day updates so it might also line up with update from 7.3.1 to 7.3.2.Do you have any suggestions?Viktor clearsky-syslog-10.10.1.202-20260924-1050.zip clearsky-diagnostics-20260924-1228.zip
September 24Sep 24 Community Expert The syslog you are using isn't actually external. The server sends syslog to its own IP, and the local syslog server writes to a user share. If the problem is a user-share or storage stall, which is what this looks like, logging stops at the same moment. That's why the log ends on normal messages.Some things point to a stall rather than a hard crash: the SSH checks got slower over a few minutes before the hang, a VM kept running, and the power button starts a shutdown that can't finish. A hard crash wouldn't react to the power button.Please do this:1. In Settings > Syslog Server, send the log to a different device (another PC, NAS or Raspberry Pi running a syslog server). If that's not possible, enable "Mirror syslog to flash" until we catch the next hang. Don't use the server's own IP as the remote target.2. Next time it hangs, don't reset it first. Log in at the console with your NanoKVM and run:ps -eo pid,stat,wchan:32,args | awk '$2 ~ /D/'Take a screenshot of the output. Then type diagnostics, and if it finishes, post the zip (it's saved to the logs folder on the flash drive).3. If you're done testing the webgui-pr-2708 plugin, remove it so there's one less thing to rule out.Rolling back to 7.3.1 is also a good test, especially if the problem started around the update.
September 24Sep 24 Author Thank you for your reply!I can point to my pfsense router. That has that capability if I recall correctly.I could never login in that console when the server is stuck. If I recall I could enter new lines but login didn't work. SSH also doesn't work when it happens.I can certainly remove that PR but then my logs would be flooded with what the PR solves and wouldn't be much use. The log was full just after few hours. Should I remove it?I haven't rolled back yet. Still waiting for another stuck and some possible feedback form the forum that I could apply. Edited September 24Sep 24 by vitis forgot to say thanks
September 24Sep 24 Community Expert 44 minutes ago, vitis said:I can certainly remove that PR but then my logs would be flooded with what the PR solves and wouldn't be much use. The log was full just after few hours. Should I remove it?Best to leave it then, and it's almost certainly unrelated.
Join the conversation
You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.