Skip to content
View in the app

A better way to browse. Learn more.

Unraid

A full-screen app on your home screen with push notifications, badges and more.

To install this app on iOS and iPadOS
  1. Tap the Share icon in Safari
  2. Scroll the menu and tap Add to Home Screen.
  3. Tap Add in the top-right corner.
To install this app on Android
  1. Tap the 3-dot menu (⋮) in the top-right corner of the browser.
  2. Tap Add to Home screen or Install app.
  3. Confirm by tapping Install.

Frequent Server Freeze

Featured Replies

I have become quite frustrated as the server no longer runs more than 2 days without a hard lockup. I have run memtest over night with no errors. I have followed the guidance for Ryzen power states and RAM. The server had previously been running quite stably for a few years with the current hardware. When I checked the video output at the most recent crash, the display is attached. I can not pinpoint any specific activity prior to the crash and the system log has no errors leading up to the crash.

Any help would be greatly appreciated!

Screenshot 2026-08-28 231351.png

tower-diagnostics-20260828-2319.zip syslog-192.168.50.6 (4).log

  • Community Expert

I dont see anything obvious logged other than your jetkvm device seems to have the usb connection reset. Do you recall messing with your JetKVM device? Enabling/disabling options that would essentially cause the USB connection to reset?

I suggest running unraid in safe boot mode and keeping containers/VMs disabled for a few days to ensure it's not being caused by a rogue plugin or container/VM.

If it can stay running as a basic NAS start enabling services 1 at a time, containers as well.

  • Community Expert

There is something significant in the diagnostics. On the boot following one of the freezes, the kernel reports:

x86/amd: Previous system reset reason [0x08000800]:an uncorrected error caused a data fabric sync flood event

mce: [Hardware Error]: CPU 6: Machine Check: 0 Bank 5:bea0000000000108

This means that at least one of the lockups ended with a CPU-reported uncorrected hardware/platform error. It does not identify the exact defective component, but it points toward the CPU or memory controller, RAM, motherboard/BIOS, or power delivery rather than the JetKVM.

The JetKVM USB connection is resetting, but the device successfully reconnects each time, there is no indication that the xHCI host controller itself died, and those events do not occur near either retained freeze boundary. I would treat that as a separate issue for now.

Recommend loading BIOS defaults and disable A-XMP/XMP, PBO, Curve Optimizer, any overclocking, undervolting, or manually adjusted CPU/RAM voltages. You can then apply only the required boot/storage settings and the recommended Ryzen idle-power setting.

Was the current BIOS installed shortly before these freezes began? If so, that is another important variable.

Safe Mode with Docker and VM Manager disabled is still a useful test, but I would perform it with those stock BIOS settings and let it run longer than the normal two-day failure window. If it freezes again or another machine check appears, test with one matched pair of DIMMs at JEDEC speed, followed by a known-good PSU/power path. A passed memtest does not rule out CPU, memory-controller, motherboard, firmware, or load-dependent power problems.

Keep remote syslog enabled and, after another freeze, collect diagnostics immediately after reboot, the previous-reset reason and machine-check record may only appear during that following boot.

  • Author

Thank you both for your replies. I was not messing with JetKVM leading up to the crashes. They seem to be very random, timing wise.

Interesting catch from the diagnostics, JorgeB. I hadn't noticed that one. I installed the current BIOS trying to deal with the crashing that has started. I don't believe any other BIOS settings were messed with, other than the guidance from the FAQ, and the server had been running comfortably for some time prior to this issue coming up. Is there a way I can more specifically identify which of those components could be the failure point?

I currently have the server running, have deleted most plugins, deleted the docker image, and only have rebuilt docker with plex installed and running. I'll leave it for a while and see if it stays running for a few days, otherwise it'll be a full safe mode for a few days.

Edited by flymcb

  • Community Expert

CPU or RAM would be my main suspects for that type of error.

  • Author

After roughly 6 days of stability, I had another crash today. I have removed most plugins and docker containers and it had been running quite well, leading me to think I was in the clear ☹️

I've attached the diagnostics from immediately after the crash that occurred around 13:41. The following error occurred in sys-log on the reboot:


Sep 4 13:44:02 Tower kernel: x86/amd: Previous system reset reason [0x08000800]: an uncorrected error caused a data fabric sync flood event

Sep 4 13:44:02 Tower kernel: microcode: Current revision: 0x08701035

Sep 4 13:44:02 Tower kernel: mce: [Hardware Error]: CPU 8: Machine Check: 0 Bank 5: bea0000000000108

Sep 4 13:44:02 Tower kernel: mce: [Hardware Error]: TSC 0 ADDR 1ffff81ecc156 MISC d01a000000000000 SYND 4d000000 IPID 500b000000000

Sep 4 13:44:02 Tower kernel: mce: [Hardware Error]: PROCESSOR 2:870f10 TIME 1788543759 SOCKET 0 APIC 5 microcode 8701035

Is the next step running in safe mode only? Do you have any further recommendations? Does the error above confirm CPU failure?

As always, thanks for all the help!

tower-diagnostics-20260904-1351.zip

Edited by flymcb

  • Community Expert

This recurrence is significant because the system again reports:

an uncorrected error caused a data fabric sync flood event

It also repeats the same Bank 5 status and other machine-check fields as the earlier crash. This time it was reported by CPU 8/APIC 5 instead of CPU 6/APIC 1.

This strongly supports a hardware/platform stability problem and makes a plugin or container the less likely root cause. It does not conclusively prove that the CPU itself is defective, since RAM/IMC instability, motherboard or VRM issues, BIOS settings, and PSU/power instability can all result in a CPU-reported machine check.

Safe Mode is still a useful control, but I would not limit the next testing to Safe Mode alone. Recommend first confirming that the BIOS is at defaults, with A-XMP/XMP, PBO, Curve Optimizer, overclocking, undervolting, and manual CPU/RAM voltages disabled, and that the RAM is running at JEDEC speed.

Since all four DIMM slots are populated, the most useful hardware test would be to run with one matched pair only in the motherboard-recommended A2/B2 slots. Because the latest failure took approximately six days, the test should run for longer than that before being considered stable. An overnight Memtest pass does not rule out runtime memory-controller or platform instability.

If it freezes again with two DIMMs at stock settings, the next steps would be to test with the other two, or use a known-good PSU/power path and then, if possible, testing with a compatible known-good CPU. The motherboard would remain another possibility.

You can also test Safe Mode with Docker and VM Manager disabled, and a test on 7.3.2 would remove the beta release as another variable. Ideally, change one variable at a time. Also you shoudl keep the remote syslog enabled and post fresh diagnostics immediately after any further reboot.

  • Author

Thanks as always for the incredibly helpful post, JorgeB. I will try all that and report back.

  • Author

Hi again. I have a quick follow up question/thought.

A while back I was getting crashes when mover was running. I changed from turbo to normal write mode and the crashes stopped, but I thought my SATA cards may have been failing or getting overloaded.

Would any of the errors I've been experiencing recently be caused by failure of either of these two cards?

https://a.co/d/0ghUwSWU

I imagine there would be other error messages associated with that, but am having difficulty remembering what happened with my previous issue. Thanks again

  • Community Expert

It seems very unlikely to me that the server crashes are related to SATA controllers; if there were issues with those, I would expect to see them logged on the syslog, in the form of ATA errors, timeouts, or similar, but if you have a way to retest without them, it may still be worth doing, but would not be my first concern.

  • 2 weeks later...
  • Author

Hi Again,

So I completely reset BIOS settings and set them per the recommendations for Ryzen CPUs.

Still crashed and put RAM pair in A2/B2. Crashed again and tried the other pair of RAM in A2/B2 with another crash.

Plex put out an update with the following information:

Plex Media Server 1.43.4.10903 is now available to Plex Pass users in the Beta update channel.

NEW

  • (GPU) Added support for HEVC transcode on devices that only support 8 bit HEVC encoding (PM-5595)

  • (Library) Add library filter for HDR10+ media (PM-5617)

  • (Metadata) Add HDR10+ metadata detection (PM-5355)

  • (Transcoding) Added support for HW Tonemapping on MacOS (PM-4883)

  • (Transcoding) Added support for HW burn in of image based subtitle formats (such as PGS) on MacOS (PM-4882)

  • (Transcoding) Added support for HW decoding AV1 videos on MacOS (M3 and newer) (PM-4884)

FIXES

  • (Crash) Certain malformed webp images can crash the server (PM-5565)

  • (Crash) Handling some POST requests could crash the server (PM-5592)

  • (Crash) In rare instances, a timeline request can crash the server (PM-5590)

  • (Database) Fixed a corrupt search index that made the database fail its integrity check and blocked backups (PM-5868)

  • (HTTP) The server could crash after making certain HTTP requests (PM-5964)

  • (Metadata) In rare instances, triggering an item rematch could crash the server (PM-5591)

  • (Metadata) Reject invalid custom provider prefixes before they get added (PM-5275)

  • (Metadata) Year value stored in TYER id3 tags would not be read (PM-5532)

  • (Search) Fixed albums, tracks and episodes whose parent item was missing being left out of search results (PM-5868)

  • (Shield) Tonemapping was being applied incorrectly (PM-5748)

  • (Tuner) The Plex Tuner Service can crash during startup and shutdown (PM-5627)

After completing the update, the server was running for 9 days and I thought I may be in the clear, but it eventually locked up again.

I've now created a script that is logging the sensor data to monitor voltages on the PSU. Would this exercise be a waste of time? I am trying to avoid spending money on a new PSU or replacing the CPU+MB unless they are truly the culprit. I've also started a post on the Plex support forum with my logs to help diagnose the potential problem there.

The server crashed at 09:25 on Sep 18. Nothing is jumping out at me, but do you see anything in the diagnostics that provide any insight?

Thanks as always!

tower-diagnostics-20260918-0954.zip syslog-192.168.50.6 (7).log

  • Community Expert

The new logs do not identify the cause of this freeze. Both the previous syslog and remote log stop at 09:22:19 with routine disk spin-down messages, shortly before your reported 09:25 lockup.

Unlike the earlier crashes, the following boot does not contain a recovered machine check or data-fabric reset reason. That does not rule out hardware instability, but we cannot describe this as another confirmed occurrence of the same hardware error.

One detail needs clarification: both boots in this diagnostic report all four RAM slots populated. Did you reinstall all four modules after testing each pair separately? Please confirm the memory speed and how long each pair ran before crashing.

This failed session also had numerous containers running, including Plex, with third-party plugins and the NVIDIA driver loaded. Before buying replacement parts, I would complete the Safe Mode test with Docker and VM Manager disabled. Keep the hardware settings unchanged during that test and allow longer than the latest nine-day interval before judging the result.

Voltage logging can provide clues, but normal readings cannot clear the PSU: motherboard sensors require correct scaling, and periodic readings can miss brief drops. There is no voltage log in these files to assess.

The Plex release notes describe application crashes; they do not establish a cause for the whole-host freezes or earlier machine checks. Plex release notes

You are now running 7.4.0-beta.2. A separate test on a supported stable release would also be useful. If the machine still freezes in a minimal configuration, a borrowed known-good PSU would be a useful next test before purchasing hardware. Use the replacement PSU’s own cables.

  • Author

Yes, I tried each pair separately before reinstalling. The first pair last 4 days before crashing. The second pair was around 6 hours. The memory speed is 2667

I've tried to keep notes, so I think I have the right diagnostics uploaded. The two attached now are right after the crashes on each pair.

I had previously disabled most plugins and all docker containers except Plex. I didn't want to lose too much uptime on Plex. I'm currently spinning up an old laptop to migrate plex to, so I can do a proper Safe Mode test.

I was also experiencing crashing on 7.3.2 - would an even earlier version be more stable for any reason?

Thanks as always!

tower-diagnostics-20260909-1658.zip tower-diagnostics-20260909-1006.zip

Edited by flymcb

  • Community Expert

Thanks for clarifying. The new diagnostics confirm two successive runs with two DIMMs installed, followed by reinstallation of all four modules. The log timelines fit the roughly four-day and six-hour intervals you reported.

Neither post-crash boot contains the earlier machine-check or data-fabric reset error. The logs end with routine disk activity, without identifying the cause.

Both pairs failing separately makes a problem confined to one pair, or to using four DIMMs, less likely. It does not rule out the CPU/memory controller, motherboard, firmware, or power supply.

Moving Plex to the laptop will allow the most useful remaining software test. Boot Unraid in Safe Mode, disable Docker, and keep VM Manager disabled. The supplied captures still show normal mode with Docker and the NVIDIA driver loaded.

Keep the BIOS, RAM, and Unraid version unchanged during this test. Allow longer than the previous nine-day interval before drawing conclusions. If it remains stable, restore services in stages.

Since you also experienced crashes on 7.3.2, there is no specific evidence here that an even older release would help. I would complete the Safe Mode test before changing versions again.

If it still freezes in that configuration, a borrowed known-good PSU would be a useful next test, followed by a compatible known-good CPU. Use the replacement PSU’s own cables.

Keep remote syslog enabled and collect diagnostics immediately after the next reboot if another freeze occurs.

Join the conversation

You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.

Guest
Reply to this topic...

Account

Navigation

Search

Search

Configure browser push notifications

Chrome (Android)
  1. Tap the lock icon next to the address bar.
  2. Tap Permissions → Notifications.
  3. Adjust your preference.
Chrome (Desktop)
  1. Click the padlock icon in the address bar.
  2. Select Site settings.
  3. Find Notifications and adjust your preference.