November 11, 20241 yr Hello, I'm having strange problems with my platform. Namely, when copying small files, say up to 15GB, or larger files (movies) up to 30GB. I have no problem and everything works great. But when I start backing up my drives, which have a lot of small files (250GB+), sometimes server simply freezes and the only thing I can do is hard reset it. As an additional fact, I will add that this is not a rule because there are days when the entire copying process goes without any problems. At the beginning, when looking through the system log, it seemed as if it was some problem related to "TPM". So I turned it off in the BIOS, and it certainly improved stability, but the problems still occur (it used to be impossible to make a backup, now it only freezes occasionally) Now when I copy data in the syslog, other errors appear, e.g.: CPU: 3 PID: 14513 Comm: notify Tainted: P B O 6.1.106-Unraid #1 Hardware name: Default string Default string/Default string, BIOS 5.27 09/11/2024 I did a memtest and it passed without any problems. I checked the BIOS settings and everything seems to be fine I have 2.5 Gb network cards and switches and "mtu" is set to 9000 everywhere I have a brand new 1TB NVME as a cash drive I will be very grateful for your ideas. Because I regret to admit that I have no idea how to properly configure network devices or UNRAID settings. (but it's interesting science ; ) unraid-syslog-previous-20241110-2231.zip
November 11, 20241 yr Community Expert I doubt that an MTU of 9000 will improve anything. it is a myth. It just creates trouble now an then. You should reset it to the default (1500 or 1504 depending on the device) everywhere and see if your problem is gone. There are known problems with large MTUs in UNRAID. Edited November 11, 20241 yr by MAM59
November 12, 20241 yr Author Hi, Thanks for the advice! Today I changed all mtu settings to default, but unfortunately it did not solve the problem :(. And today, while making a backup copy, the server froze again and I had to hard reset it. I am adding fresh "syslog" unraid-syslog-previous-20241112-1557 (1).zip
November 12, 20241 yr Community Expert There are constant call traces logged, looks more like a hardware issue.
November 12, 20241 yr Author My gut tells me it's probably some processor settings in the BIOS, but I have no idea what... because stability improved by 50% if only I turned off TPM. So that's part of the puzzle Other ideas assume it might be something to do with the switch (I know it shouldn't matter) but apart from the server hanging, the switch hangs too. The only question is which device is hanging because of which device. I will change the switch to another one and then "bombard" it with a few hundred gigs of small stuff. As for the hardware problem, I changed the motherboard and RAM
November 12, 20241 yr Community Expert if it hangs again next time, just pull out the lan cable, wait 5 seconds and plug it in again. Transfers will abort, but machine/switch may recover. If this is the case, your problem is "not working flow control". Check settings in the computer and the switch. It may happen that someone says "STOP", the other does not honor the request and finally the first one locks up the port to prevent being overrun. Check this out
November 12, 20241 yr Author Unfortunately, it's not a network problem. I've changed the switches and network cards in both my computer and the server. "The little things are still killing it." As JorgeB wrote, it's probably a hardware problem (or its configuration in the BIOS), but I can't quite figure out what device the problem is. Looking at the logs, you can conclude that it's the processor. But I can't change it because it's a soldered N100. I add syslog after changing cards and switch. unraid-syslog-previous-20241112-1750.zip
November 19, 20241 yr Author Solution I think I found the culprit. It wasn't the processor but the nvme disk responsible for the cache. Generally, I replaced the disk in the meantime, but with a different capacity and not from a different manufacturer. Namely, I had an ADATA LEGEND 800 disk installed, the first one was 1TB and the second 2TB. It turns out that the controller on my board doesn't really like the controller in the disks... For now, I changed the cache disk to a regular Samsung SATA 2TB(I didn't have any free nvme. I just have to buy one for Black Friday). The only strange thing is that the instability was also when reading and transferring small files from the platter disks in my NAS. But I assume that the transfer from them was also done with the help of cache, or the mere presence of these disks in the system caused instability in the operation of the whole thing. When I install a disk from another manufacturer, I'll let you know! (damn, I hope it's not a problem with the pcie lines to the processor :D)
December 4, 20241 yr Author Hello, Sorry I haven't replied for so long, but I've been waiting for parts from China ;). In general, everything has been working stably since I removed the ADATA NVMe drive, so I assume it's simply not compatible with my board or unrai. In place of the drive, I've mounted an adapter to the PCI-E 4x port (my board is one of those cheap Chinese N100s and doesn't have a PCI-E port on the board). I've now connected an LSI 2308 controller to this port, and 16TB drives to it, and I'm happy to report that everything is working very stably and uses really little power (below 40W). Thank you all very much for your tips! See you next time!
Join the conversation
You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.