Skip to content
View in the app

A better way to browse. Learn more.

Unraid

A full-screen app on your home screen with push notifications, badges and more.

To install this app on iOS and iPadOS
  1. Tap the Share icon in Safari
  2. Scroll the menu and tap Add to Home Screen.
  3. Tap Add in the top-right corner.
To install this app on Android
  1. Tap the 3-dot menu (⋮) in the top-right corner of the browser.
  2. Tap Add to Home screen or Install app.
  3. Confirm by tapping Install.

JorgeB

Moderators
  • Joined

  • Last visited

Everything posted by JorgeB

  1. Syslog is after a reboot, so we cannot see the failure, but SMART looks fine, so most likely a power/connection issue. Recommend checking/replacing cables for that disk and then rebuilding on top. After that upgarde Unraid (note that youds wolnd need to conver the the filesystem of all array disks)
  2. This boot looks much healthier. I do not see any of the kernel oopses, process faults, or decompression errors found in the previous log. The BCM57810S is also absent. The Broadcom messages that remain concern the motherboard’s Wi-Fi and Bluetooth devices, not the 10GbE adapter. Removing the suspect RAM is therefore a strong lead, but it does not yet prove that those modules caused the problem. You also reset the BIOS, recreated the Unraid files, and removed the NIC. Any of those changes can affect the result. Also confirm the number of completed Memtest86+ passes and the total test duration. A clean test of the remaining RAM shows that this configuration passed. It does not prove that the removed modules are faulty unless they failed separately. This log covers only about 87 seconds. One earlier NIC-free fault occurred after approximately 14 hours, so this is not yet a sufficient stability test. Recommend keeping the suspect RAM and 10GbE card removed. Then: 1. Disable array auto-start. 2. Boot Unraid OS Safe Mode (no plugins). 3. Leave the array, Docker, and VMs stopped. 4. Keep remote syslog enabled. 5. Leave this configuration running for at least 24 hours. 6. Then post full diagnostics and the remote syslog. The restored network configuration still references absent eth2 and eth3 interfaces. We should also clean that configuration before testing the 10GbE card. The BIOS reset disabled hardware virtualization. Leave it disabled during this control test. If you use VMs, you can enable it again after the system proves stable. The log also shows an unclean shutdown. A correcting parity check started and was then cancelled. Keep the array stopped during the hardware test. After the hardware is stable, run a complete non-correcting parity check. Do not reinstall the 10GbE card yet. If the 24-hour Safe Mode control remains clean, we can reintroduce it in stages: card installed without the cable, link connected without traffic, network-only traffic, and storage activity last.
  3. btrfs is detecting data corruption, start by runing memtest.
  4. This was not a system-wide OOM. It was a Docker container related OOM because that container hit its configured 512M RAM limit.
  5. Iperf results confirm it's a network problem. This is typically the NIC (or its driver), cable, switch , or client PC
  6. Are you booting from a flash drive or internal boot pool?
  7. SimonF from LT already created a bug report about this, it will need to be fixed upstream and then make its way to current kernels: GitLabWindows 11 VM fails to resume from sleep if guest agent i...Host environment Operating system: Unraid OS/kernel version: Linux 6.12.30-Unraid x86_64 Architecture: x86...
  8. If you cannot type anything in the console, you will need to reboot, by force if needed. Data should be OK, make sure to enable the Sylog server after in case it happens again.
  9. This scrub confirms a real data-integrity problem on disk2, but it still does not explain the whole-server resets. The increase from corrupt 89 to corrupt 90 means the scrub detected this checksum mismatch. It does not tell us when the data became corrupt. First, protect the data. If the named file is available from a known-good backup, replace it. If it is expendable, it can be deleted and recreated. If it is important and has no backup, do not delete it before deciding whether recovery is needed. After replacing or deleting the file, run another scrub on disk2 and confirm that it reports no uncorrectable errors. Replacing or rebuilding the disk should not be relied upon to repair that file. Parity does not know which version satisfies the Btrfs checksum. The attached SMART report does not presently condemn the disk. It passes, contains no logged ATA errors, and reports zero reallocated, pending, offline-uncorrectable, reported-uncorrectable errors. You can also run an extended SMART test. The WD60EZAZ is an SMR model. SMR can cause poor sustained-write and rebuild performance, but this does not show that SMR caused the server resets. Smartmontools also prints a warning about reported problems with some older WD SMR drives. Good backups and a planned replacement with a suitable CMR disk would therefore be reasonable precautions, but this is risk management rather than a confirmed diagnosis. I would not move it to another bay yet. That will not repair the affected file and changes another storage-path variable. If the extended test fails, another scrub error appears, or the counter rises again after the file is corrected, replace the disk and then investigate the bay, cable, expander connection, and HBA one change at a time. Treat the HBA update to 16.00.12.00 as the beginning of a new observation period, continue the remaining scrubs, complete a full offline memory test, and collect the IPMI SEL/reset-cause information after another reset. If the resets continue, the onboard-NIC-only test with the Mellanox card physically removed remains the next controlled comparison.
  10. Thanks, this log changes the diagnosis. I would not reintroduce 10GbE yet. The server recorded a kernel page fault before the replacement-card work and while bnx2x was not loaded. The log later contains four more unrelated kernel oopses, repeated decompression failures, and many losetup processes faulting at address zero. This points to broader system corruption rather than a confirmed NIC or NFS problem. The latest boot in the attached log also still detects the Broadcom card and faults again. It does not show a clean final boot with the card removed. Recommend to physically remove the 10GbE card and boot a freshly recreated release in Safe Mode. Keep the array, Docker, and VMs stopped. Leave remote syslog active and post fresh diagnostics from that boot. Also run multiple complete Memtest86+ passes at BIOS defaults. If practical, test one matched DIMM set. Do not test the NIC again until the NIC-free baseline has no oopses, process faults, or decompression errors.
  11. You should post in the existing support thread: https://forums.unraid.net/topic/98978-plugin-nvidia-driver/
  12. Enable the syslog server and post that after an array tsrat attempt to see if there's something there.
  13. Those ICRC abort errors are usually a bad SAT cable; replace just that cable one more time to make sure that's not the problem.
  14. See if you can get the syslog at least: cp /var/log/sylog /boot/syslog.txt
  15. Run a single thread iperf test in on direction and then reversed (-R) and post the results.
  16. Type tail -n 250 /proc/spl/kstat/zfs/dbgmsg > /boot/zlog.txt Then attach that file here
  17. Please post the complete diagnostics.
  18. Stop array, unassign parity2, start array, stop array, assign it as parity1, start array to begin parity sync, array will be unprotected until the sync is done.
  19. The cache device is nvme1n1, and as far as I can see, it's working normally; it looks like an FCP false alarm.
  20. Thanks, that clarification is important. It confirms that 14:33:47 was the failure itself and not a shutdown that you requested. This was not a kernel crash or complete host lockup. Instead, many core services received SIGTERM together. This stopped networking, user shares, DNS, UPS monitoring, ACPI handling, and Samba. The physical server and local console remained active. This resembles an incomplete system teardown. However, the log does not identify the process that initiated it. The UPS faults still need attention, but the event order does not show that NUT initiated this teardown. There is no nearby on-battery, forced-shutdown, or automatic power-fail message. NUT receives its termination signal after the broader teardown starts. Connect the UPS data cable directly to another motherboard USB port. Then monitor whether the battery and USB communication warnings return. Run the Safe Mode test for at least the usual one-week recurrence period. Safe Mode removes third-party plugins, including NUT, from the active path. It also removes automatic UPS shutdown protection, so do not leave the server unattended if the utility power is unreliable. If the failure returns and the local console still works, do not reboot immediately. Run the commands from my previous reply and add: runlevel ps -p 1 -o pid,ppid,stat,etime,comm,args last -x | head -n 20 mount Then run diagnostics and attach the resulting file. These results can show whether the host entered a shutdown runlevel or whether another process sent broad termination signals while the normal runlevel remained active. Also confirm exactly how you restarted the server at approximately 17:57. Did you type reboot, press the reset button, or cycle the power? If you used reboot, did it complete normally? Also describe what the console showed before the restart. One nginx reload started two seconds before the service termination. Reloads were occurring about every five minutes. This timing is relevant, but it does not yet prove that nginx or Unraid Connect caused the teardown.
  21. The best bet is to look for the support link for that container
  22. The previous syslog does not show an Unraid crash. It ends after Mover completes normally at 02:10, and the next boot at 07:23 is marked unclean. There is no panic, call trace, machine check, HBA timeout, disk-I/O error, or thermal event before it. This looks more like a platform-level reset that occurred before Linux could save the cause, although the responsible component is not yet known. The most actionable finding is the storage path. The SAS9300-4i is running firmware 15.00.02.00, and four separate Btrfs array filesystems retain corruption counters of 4, 89, 255, and 5. These are historical counters and do not prove the HBA caused the reboot, but all affected disks share the HBA and SAS expander. After confirming that important data is backed up, recommend to verify and update the HBA using the correct supported IT firmware for that exact card. Then run Btrfs scrubs on the four affected disks, save the results, record the existing counters, and monitor whether they increase. Do not clear the counters until their current values and scrub results have been saved. Also export the complete IPMI System Event Log and any reset-cause information after the next occurrence, before clearing anything, and include the installed BMC firmware version. Separating the two mains supplies helps exclude an external outage, but it does not rule out a PSU, power-distribution board, motherboard, or BMC-triggered reset. Can you clarify which memory test you ran? The diagnostics show that the live dwmemtester plugin is installed, but neither log contains a test start or result. If this was boot-time Memtest running all tests for three days, that is already a useful result and makes a straightforward repeatable DIMM fault less likely. If it was the live plugin, disable it and use offline boot-time Memtest instead. If the reset returns after checking the HBA and filesystems, use an onboard Intel i350 interface and physically remove the Mellanox card for one observation interval. After that, test Safe Mode with Docker and VM Manager disabled. Change only one item per interval because a stable week is not conclusive when the normal interval can be a month. Keep the independent remote syslog receiver enabled and add external ping or heartbeat monitoring. Netconsole may capture something that normal userspace syslog misses. I would not add more speculative PCIe or power-management boot options, the current C-state restriction is applied and did not resolve the problem. The diagnostics only identify 7.3.1 as the previous Unraid version. Older diagnostics or flash-backup metadata would be needed to reconstruct the version that was running when the resets began.
  23. There something creating /mnt/user/system before the array is started that is not supported. Start by booting in safe mode to rule out a plugin.
  24. Thanks, the log does not show a kernel crash. At 14:33:47, several core services receive termination signals together. The network interface is removed, user shares stop, and the UPS and DNS services exit. Can you confirm whether you pressed the power button or requested a shutdown or reboot at that time? Also, how did you restart the server at approximately 17:57? Please describe what happened during the intervening period. If 14:33:47 was your recovery attempt, this log does not contain the original failure. If you did not initiate it, an unidentified process started the system teardown. The UPS also needs attention. It repeatedly reports: - Low battery - Battery replacement required - No battery installed - USB communication errors Recommend to check or replace the UPS battery. Also connect its USB cable directly to another motherboard port. The log separately shows an nginx reload approximately every five minutes. Some reloads fail because the generated configuration is invalid. This can affect the WebGUI and Unraid Connect, but it does not explain the loss of SSH or the coordinated service termination. A Safe Mode test is reasonable. Run it for at least the normal recurrence period. Safe Mode disables plugins, including UPS monitoring, so do not leave the server unattended if the utility power is unreliable. If the problem occurs again and the local console still responds, run the following commands before restarting anything: date uptime ps -eo pid,ppid,stat,wchan:32,etime,comm,args ip -br addr ip route ss -lntp dmesg -T | tail -n 200 diagnostics Please attach the generated diagnostics and describe which local commands still worked.

Account

Navigation

Search

Search

Configure browser push notifications

Chrome (Android)
  1. Tap the lock icon next to the address bar.
  2. Tap Permissions → Notifications.
  3. Adjust your preference.
Chrome (Desktop)
  1. Click the padlock icon in the address bar.
  2. Select Site settings.
  3. Find Notifications and adjust your preference.