2 hours ago2 hr Hello everyone,I'm looking for some guidance from our community. I have been running the server fine for about 2 years now without issue. Recently over the last few months I've noticed the server hangs. All dockers pause, GUI freezes, cant log into the server and majority or the time requires a hard reset.Its hard to pull diagnostics when everything locks up.Every few weeks it locks up. I am looking for some help to troubleshoot this.I have attached 2 diagnostics files, first one is from today a little after the restart of it hanging, and the second from a month ago.I also push the syslog to a synology box and noticed after 713AM the syslog stops till I restart the server. I believe this is when the server hung.The syslog is filled with avahi-daemon warning.10/10/202611:43:31InfoWanderTowersyslogrsyslogdrsyslogd[internal_messages]: 391 messages lost due to rate-limiting (500 allowed within 5 seconds)10/10/202611:43:31InfoWanderTowersyslogrsyslogdaction 'action-2-builtin:omfwd' resumed (module 'builtin:omfwd') [v8.2102.0 try https://www.rsyslog.com/e/2359 ]10/10/202611:43:31NoticeWanderToweruserflash_backupchecking for changes every 1800 seconds10/10/202607:13:07WarningWanderTowerdaemonavahi-daemonRecord [WanderTower._device-info._tcp.local#011IN#011SRV 0 0 0 WanderTower.local ; ttl=120] not fitting in legacy unicast packet, dropping.10/10/202607:08:07WarningWanderTowerdaemonavahi-daemonRecord [WanderTower._device-info._tcp.local#011IN#011SRV 0 0 0 WanderTower.local ; ttl=120] not fitting in legacy unicast packet, dropping.10/10/202606:59:07WarningWanderTowerdaemonavahi-daemonRecord [WanderTower._device-info._tcp.local#011IN#011SRV 0 0 0 WanderTower.local ; ttl=120] not fitting in legacy unicast packet, dropping.10/10/202606:50:07WarningWanderTowerdaemonavahi-daemonRecord [WanderTower._device-info._tcp.local#011IN#011SRV 0 0 0 WanderTower.local ; ttl=120] not fitting in legacy unicast packet, dropping.Thank you wandertower-diagnostics-20260913-0156.zip wandertower-diagnostics-20261010-1502.zip
1 hour ago1 hr Had a look at both diagnostics. The useful one is the 10/10 set, taken 3h20 after reboot with everything running. The September set was taken while the array was already stopping (VM force-killed and containers stopped about a minute before), so it isn't a baseline for normal operation.Facts from 10/10:- 32 GB RAM, no swap. 29 GB used, 2.1 GB available.- Windows10 VM: 12 GB configured, 12.3 GB resident, started within seconds of libvirt coming up.- rootfs (lives in RAM, 16 GB max) is at 11 GB; Shmem 11.7 GB. In September it was 2.2 GB, but see the caveat above.- A /tmp/Transcode directory exists (also present in September).- php-fpm children killed with SIGKILL: at least 13 between 14:44 and 14:56 (the log collapses repeats), with nginx 502 errors on the GUI at 14:45. The September log has 21 of the same kills on Sep 12, 21:18-21:27.- No kernel "Out of memory" line in either local syslog, so I can't prove it's the OOM killer. It does fit a box with 2 GB free and no swap.- Both boots start with "unclean shutdown detected". SMART passes on all array and pool drives (one external USB drive reports no SMART data). No MCE lines.What I'd check:1. du -xsh /tmp/* /var/lib/* 2>/dev/null | sort -h | tail If /tmp/Transcode is your Plex transcode dir, it competes with the VM and containers for the same RAM and can grow to 16 GB. Move it to the plex pool/cache or mount a dedicated tmpfs with size=4G.2. Synology syslog for Oct 10, 07:00-07:13: look for "Out of memory", "oom-kill", "Killed process" or any kernel trace. That decides it.3. Per the September libvirt log the VM is started every Saturday 09:00 UTC and stopped Monday 09:00 UTC. Oct 10 was a Saturday, so if that schedule is still active the VM was running at 07:13 local.4. Dynamix File Integrity (bunker) verified all 7 disks, overlapping, from 07:00 on Sep 2. Same time window as your hang. Could be coincidence, but easy to stagger or pause for a test.5. Until the memory picture is clear: VM to 8 GB, and --memory limits on the big containers (Nextcloud AIO, immich, Postgres).6. docker.img shows 861 btrfs write errors ("BTRFS info (device loop2): errs: wr 861"). The counter is unchanged since August, so historical, but worth recreating the image at some point.7. For a RAM test, boot memtest86+ rather than using the in-OS Live Memory Tester plugin - with 2 GB free a userspace test only adds to the pressure.Minor: r8169/RTL8125 dropped link twice on Aug 27 and once on Aug 28. Edited 1 hour ago1 hr by UnraidSecretaryOffice
Join the conversation
You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.