5 hours ago5 hr Unraid [7.3.2], kernel 6.18.38-Unraid. i7-14700K, W680, 128GB ECC. Diagnostics attached.Summary: A userspace process (Plex Media Scanner) got permanently stuck in the kernel inside munmap(), waiting on a per-VMA lock that nothing holds. Every ps/pgrep that touches that process then blocks forever, php-fpm workers pile up behind them, and the WebGUI dies. SSH and the other containers stay up. Only a reboot fixes it. It's happened a few times, at most every couple of weeks; this is the first time I've had the time to catch and troubleshoot it properly.I reboot with "powerdown -r" via SSH. I need to send this twice as the first one never finishes shutting down. On restart, a parity check needs to run. Plex doesn't seem to be the bug, it's the trigger. Userspace shouldn't be able to wedge the kernel like this.Plex container details: plexinc/pms-dockerTimeline:- Oct 2 04:20:17: Plex launches credits detection (Plex Media Scanner -C) on one episode. It never finishes. Plex logs it giving up at 04:50.- From then on: ps and pgrep hang. The first hung pgrep's PID is ~2,400 after the scanner's, so it started within minutes.- Oct 3 04:05: AppData Backup stops the Plex container. The scanner gets the signal and every thread enters do_exit, but none can finish. The container stays "Up" with its init gone, and the healthcheck fails with "OCI runtime exec failed: setns process: exit status 1".- WebGUI: login spins forever. nginx logs "upstream timed out" on php-fpm.sock for /Dashboard, and /Main hangs too.- Load average ~28 with ~2 runnable, i.e. ~26 tasks in D state.The stuck thread (the only scanner thread not in do_exit):TID 1629308 Plex Media Scan state=D wchan=__vma_enter_locked __vma_enter_locked+0x85/0xb0 __vma_start_write+0x13/0x40 __split_vma+0x132/0x220 vms_gather_munmap_vmas+0x81/0x1e0 do_vmi_align_munmap+0x11f/0x1a0 do_vmi_munmap+0x109/0x130 __vm_munmap+0xd0/0x110 __x64_sys_munmap+0x17/0x20All 17 other scanner threads are stuck at do_exit+0x2b9. Every ps/pgrep is stuck here: mmap_read_lock_killable __access_remote_vm environ_read vfs_readEvery other thread in the process is in do_exit, so nothing alive holds that VMA lock. It leaked.Same signature as these reports on other 6.18.x kernels: https://github.com/microsoft/WSL/issues/41708https://github.com/microsoft/WSL/issues/41113One of the WSL reports shows the same pattern on 6.18.40, so I wouldn't assume a point release fixes it. I haven't found an upstream fix.Questions:1. Is this known, and is there a kernel with a fix on the roadmap?2. Can CONFIG_DETECT_HUNG_TASK be enabled? /proc/sys/kernel/hung_task_timeout_secs doesn't exist, so the kernel logged nothing despite threads being stuck for 24+ hours. I had to dig everything out by hand.Workaround in the meantime: turning off Plex credit-marker generation, since that's the workload that tripped it. The bug will presumably find another victim eventually.Attached: diagnostics (taken while wedged), the hang captures, the per-thread stacks, and the Plex logs. 00-commands-used.txt 01-hang-capture.txt 02-dstate-scan.txt 03-scanner-threads.txt define-diagnostics-20261003-1141.zip Plex Media Scanner Credits logs.txt Plex Media Server.log
1 hour ago1 hr Community Expert Thanks for the detailed captures. They make this much easier to look at.I agree with your reading. The scanner thread is stuck inside the kernel's memory-management code during munmap(), and the rest follows from that. Anything that reads that process's /proc/<pid>/environ or cmdline (ps, pgrep, and the WebGUI's own process checks) blocks behind it, so the PHP workers pile up and the WebGUI stops responding. The process can't exit, so the array can't unmount cleanly. That is why the first powerdown -r doesn't finish and a parity check runs after the reboot. Plex only triggers it. Userspace should not be able to leave a task in that state. From the diagnostics, I don't see a hardware angle. Microcode is 0x133 (current for Raptor Lake), and the syslog has no MCE/EDAC or other kernel warnings. Other users also report the same stack on WSL 6.18.x kernels on different hardware.1. I don't have a confirmed upstream fix for this specific signature yet. 7.3.3-rc.2 uses Linux 6.18.52. If you're OK with running a release candidate, it would be a useful comparison, since it includes many mm fixes after 6.18.38. Please keep in mind that it may not fix this. If you test it, it would be best to turn credit detection back on at some point, because otherwise we can't tell whether the kernel or the workaround made the difference.2. I'll pass on the hung-task detector request. It would only make the kernel log the stall automatically. It would not prevent it.For now, keeping credit-marker generation off is a reasonable workaround. If it happens again, before rebooting please run echo w > /proc/sysrq-trigger, which writes the stacks of all blocked tasks to the syslog even without the hung-task detector, and then grab diagnostics. Also note the kernel version and whether credits were enabled.
Join the conversation
You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.