Skip to content
View in the app

A better way to browse. Learn more.

Unraid

A full-screen app on your home screen with push notifications, badges and more.

To install this app on iOS and iPadOS
  1. Tap the Share icon in Safari
  2. Scroll the menu and tap Add to Home Screen.
  3. Tap Add in the top-right corner.
To install this app on Android
  1. Tap the 3-dot menu (⋮) in the top-right corner of the browser.
  2. Tap Add to Home screen or Install app.
  3. Confirm by tapping Install.

Unraid 7.3.2 host-wide stall during parity check after 2 months uptime; hard reset required

Featured Replies

Hi,

My server had been stable for approximately two months before this incident. On 28 September 2026 it became almost completely unresponsive and eventually required a hard reset.

Symptoms

- The WebGUI and all hosted applications stopped responding.

- The server still answered ICMP ping consistently.

- TCP connections to HTTP and SSH could be established, but HTTP returned no usable data and SSH never presented a usable session/banner.

- A local SSH/Telnet attempt from another device on the LAN also failed.

- Remote access through Tailscale failed in the same way.

- The BMC/IPMI interface remained reachable and its sensor overview looked normal.

- A graceful ACPI shutdown requested through the BMC did not complete. The BMC reported that the host OS was taking unusually long to shut down.

- A hard reset was ultimately required.

- Ping returned after approximately 4 minutes 42 seconds, followed by normal SSH/Telnet and WebGUI access.

This did not look like a simple WAN, NIC, Tailscale or WebGUI problem: the network stack still handled ping and TCP handshakes, while multiple unrelated userspace services and the ACPI shutdown path stopped making progress.


System context

- Unraid 7.3.2, upgraded from 7.3.1

- Linux 6.18.38-Unraid

- Intel Core i9-14900K, 128 GiB RAM

- ASUS Pro WS W680-ACE IPMI, BIOS 3101 (2023)

- Broadcom/LSI SAS3224 SAS-3 HBA with expanders

- 28 data drives, dual parity and four NVMe pool devices

- Dual Intel I226-LM 2.5 GbE in an active/backup bond

- A scheduled non-correcting parity check was active before the incident


Evidence after reboot

- Unraid recorded an unclean shutdown/system crash.

- The pre-reset parity check was reported as cancelled with zero parity errors recorded. It had approximately 11 hours of actual runtime.

- Unfortunately, /var/log was RAM-backed and neither remote syslog nor flash mirroring was enabled. The diagnostics therefore start with the new boot and contain no pre-crash kernel messages, blocked-task output or call trace.

- The first diagnostics were collected shortly after reboot while the encrypted array was still stopped, so Docker, VM and filesystem runtime information in that package is incomplete.

- After the array was unlocked, the array and Docker services started again. The automatic parity check caused by the unclean shutdown was visible but paused at 0.0%.

The preserved post-boot diagnostics look healthy:

- load average 2.02 / 1.53 / 0.77 on 32 logical CPUs;

- approximately 98.6% CPU idle, 0.0% I/O wait and no D-state processes;

- ample free memory and no OOM evidence;

- no kernel panic, call trace, hung task, lockup, MCE/EDAC, filesystem-corruption, block-I/O or NIC-watchdog errors;

- SAS HBA/expanders and both 2.5 GbE links initialized normally;

- all HDD and NVMe SMART health assessments passed;

- no current ATA pending, offline-uncorrectable or CRC errors;

- all NVMe devices reported zero critical warnings, zero media/data-integrity errors and zero error-log entries.

There are a few small historical SMART counters on older HDDs, but no current error or timing correlation that identifies a failed drive. The boot USB also initialized normally.

One separate maintenance warning is present: NUT reports that the UPS battery should be replaced. I saw no evidence of a power loss during this incident, because the BMC and host stayed online, but I will service the battery and export the BMC System Event Log.

The motherboard BIOS is old for a 14th-generation Intel CPU. Linux loaded microcode revision 0x133 during boot and there are no MCE or thermal errors, but I will also verify the appropriate ASUS BIOS/Intel ME update path after the storage investigation.

Current assessment

My leading hypothesis is a host-wide kernel/storage-I/O stall, possibly involving the md/parity path, mpt3sas, an expander, a disk path or a mounted filesystem. This is based on the symptom pattern and the active parity workload, not on a preserved pre-crash trace.

A kernel/driver/plugin deadlock or a slow resource leak is also possible. I am not claiming that the parity check, HBA, a disk or Unraid 7.3.2 is proven to be the cause.

This appears similar to the recently reported case here:

Questions

1. Does this symptom pattern suggest an md/storage-I/O or mpt3sas stall, given that ICMP and TCP handshakes still worked while userspace services stopped responding?

2. Are there known Unraid 7.3.2 / Linux 6.18.38 issues involving parity checks, mpt3sas, expanders or host-wide userspace stalls?

3. Would you recommend testing Safe Mode, temporarily rolling back to 7.3.1, or checking/updating HBA firmware first if it recurs?

4. Is the older motherboard BIOS/ME a meaningful suspect despite the newer Linux-loaded microcode and absence of MCE/EDAC errors?

5. Which Magic SysRq traces or console commands would be most useful during the next hang?

6. Besides remote syslog to a different physical machine and the BMC SEL, what additional persistent diagnostics should I enable?

I plan to enable remote syslog to another device. If the problem recurs and the console responds, I plan to capture D-state tasks with:

ps -eo pid,stat,wchan:32,args | awk '$2 ~ /D/'

unraid-diagnostics-sanitized-20260928-1638.zip

Edited by casperse

  • JorgeB changed the title to Unraid 7.3.2 host-wide stall during parity check after 2 months uptime; hard reset required
  • Community Expert

Thanks for the detailed report. Unfortunately, as you noted, there's nothing from before the reset. The diagnostics only show a clean boot, so they can't show what caused it. I agree it looks similar to https://forums.unraid.net/topic/200676-unraid-732-intermittent-host-hangs-every-14-days/. Both have ping still working, all services stopping, a shutdown that starts but never finishes, and no crash signature. That topic has no cause yet either, and it's on a completely different platform (AMD, SAS2308).

1. The pattern tells us the kernel was still partly alive. Ping replies and TCP handshakes are handled inside the kernel, but the SSH banner, the HTTP reply and the ACPI shutdown all need userspace processes to make progress. So something blocked userspace system-wide. A storage or filesystem stall, with tasks stuck in D state, is one common cause. A kernel lock problem or an unstable platform can look exactly the same. Without a trace from the hang, we can't tell which it was.

2. There's no confirmed 7.3.2 bug that matches this. The topic above is the closest report. Your check averaged about 500 MB/s, so there's nothing unusual there.

3/4. I'd start with the BIOS, before Safe Mode, a rollback or HBA firmware. BIOS 3101 is from December 2023 and has microcode 0x11f. That's older than Intel's microcode fixes for 13th/14th-gen instability (Vmin shift) and older than the BIOS releases that added Intel's default power profile. Linux loads 0x133 early at boot, so the runtime microcode is current. But the board's power limits still come from the old BIOS, and if the CPU ran on older microcode for a long time, some damage may already be done. This kind of instability often causes random hangs with no MCE logged, so I'd treat it as a real suspect. Update to the latest stable BIOS for the board, load defaults, and choose the Intel default power profile (no MCE, undervolt or unlimited power limits). After that, check the HBA firmware version. Your sanitizer replaced it in the syslog (FWVersion(IP_REDACTED)). Make sure the 9305-24i is on the latest P16 firmware.

With one event in about two months, a Safe Mode test or a rollback to 7.3.1 would need months to prove anything. I'd only try those if it happens again after the BIOS update.

5. In the other topic, console login didn't work during the hang: new lines echoed, but login never finished. So don't count on getting a shell. SysRq keys are handled by the kernel and don't need a login:

- Enable SysRq now, because you may not be able to when it hangs: echo 1 > /proc/sys/kernel/sysrq

- During a hang, from the BMC KVM console, press Alt+SysRq+w (blocked tasks) and then Alt+SysRq+l (backtraces from all CPUs).

- If you do get a shell, change your ps command to use comm instead of args. Reading the command line of a process stuck in D state can block too, and then your shell hangs:

ps -eo pid,stat,wchan:32,comm | awk '$2 ~ /D/'

Also run cat /proc/pressure/io /proc/pressure/memory and dmesg | tail -100.

6. Remote syslog to another machine is good, but rsyslog is a userspace process and may stop during the same hang. Netconsole is better for this. It sends kernel messages, including hung-task warnings and the SysRq output above, straight to another machine over UDP without going through userspace. It's available in 7.3.3-rc.2 and 7.4.0-beta.3 (not 7.3.2). If you're willing to run the RC, I'd do the BIOS update first. Otherwise, a missing hang afterwards won't tell us which change helped. Also save the BMC SEL, and tell us whether the parity check was running or paused by Parity Check Tuning when it hung.

  • Author

Thanks @JorgeB
I tried something like this on my Unraid server some years ago on some old version and they disappeared in future releases!
I havent installed anything new in a very long time things just worked! and dockers get auto updated,
So maybe I am back to scheduling a monthly reboot, better than a crash.

Join the conversation

You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.

Guest
Reply to this topic...

Account

Navigation

Search

Search

Configure browser push notifications

Chrome (Android)
  1. Tap the lock icon next to the address bar.
  2. Tap Permissions → Notifications.
  3. Adjust your preference.
Chrome (Desktop)
  1. Click the padlock icon in the address bar.
  2. Select Site settings.
  3. Find Notifications and adjust your preference.