Monday at 12:58 PM2 days Hi there,I thought I'd be done with this, but hey. Another crash, after a long history of crashes.hard data:today, 13:54h, I got an automatic notification from Tautulli that Plex is offline. This means, the Plex container stopped before Tautulli (if Tautulli stopped at all)from the log of my Tasmota plug (attached below; time is -2h for whatever reason) I can see that power-consumption rose around 13:50h, rather quickly to ~130W, staying around that valuean externally recorded syslog (also attached below) stops at 13:56h, reading SMART datathe syslog is also spammed with a reoccurring message regarding "docker0", but did this since early in the morning; however, those stopped around 13:49h; coincidence?WebUI completely unresponsive, SSH as well. I had the same form of crashes before (server starting to run wild with power consumption, being unresponsive); I am not at home, and won't be for hours, so I can't tell if pushing the power button would have shut it down safelyI hard-cut its power and switched the Tasmota back on, server boots normally (unclean shutdown lead to Parity check of course)that docker0 still spams the log after rebootI attached logs and diagnositics, but something tells me it's as obscure as it was in the past :(Things I tried since all those earlier crashes this year:switched RAM (passed memtest 5 times)switched CPU (to rule out 13th gen issues)recently, I switched the whole mainboard, newest BIOS installed (to get more SATA ports/throw out extenders)threw out all SATA extenders, all disks now attach natively to the board (as I thought it could've been the ASM cards)generally removed disks, and replaced SATA SSD cache with NVMeThe mainboard was switched some weeks ago, and on July 4th the Server went back to its designated place and ran there well till now. There were no crashes in this config before.Before the mainboard switch, I deactivated all ASPM/C-State stuff inside the BIOS; it crashed less, but still did.I also tweaked C-States/ASPM with this new Mainboard, to be able to save some power again (got great help in the german part of the forum)Messages in syslog nearly always related to reading SMART data. I am not sure how accurate all timestamps are, the SMART-readings came a bit after it "went wild" on the power consumption (I never know if it's the disks or the CPU; during the Parity check now it uses 95-99W so I kinda assume both, to a degree). The disk that gets SMART-read last is never the same as the last crash.After I changed and switched so much, I really don't know what else to try :( lunas-diagnostics-20260720-1432.zip 260720_tasmotaHistory.csv 260720_lunas_syslog.log
Monday at 02:04 PM2 days Community Expert There's nothing really relevent logged in the logs, however the syslog file you posted shows a container boot-looping which may potentially be the source of this issue.On the docker page you can see the docker uptimes, find out which one is 0 seconds or relatively short uptime and stop it.Ill be honest you might even want to consider running unraid for a while with all dockers stopped and see if the behavior goes away.
Monday at 02:06 PM2 days Community Expert Another thing is you are using plugins and 3rd party packages which alter the system in ways which might be unintended such as powertop and autotweak.These could cause system instability.
Monday at 05:40 PM2 days Author 3 hours ago, MowMdown said:On the docker page you can see the docker uptimes, find out which one is 0 seconds or relatively short uptime and stop it.I found the rogue Docker, it's Homebox. The Log crashes when I try to open it. I closed it down, looked at the Logs, some kernel panic because of the API Key or so. Will investigate this later, would suck if it's broken, that's my whole DIY inventory :( anyway, this is at least settled!3 hours ago, MowMdown said:Ill be honest you might even want to consider running unraid for a while with all dockers stopped and see if the behavior goes away.Ah well, this time it took roughly 3 Weeks till it crashed, that's not really an option I am afraid :( 3 hours ago, MowMdown said:Another thing is you are using plugins and 3rd party packages which alter the system in ways which might be unintended such as powertop and autotweak.These could cause system instability.I am not sure why I installed Autotweak, I guess to play with the power profile; I removed it now.Powertop I use for the mentioned power saving, or at least the attempt to do so. It's recommended a lot in this forum, didn't know it could lead to instabilities :( Best,Rick
Monday at 05:44 PM2 days Community Expert 1 minute ago, CameraRick said:I found the rogue Docker, it's Homebox. The Log crashes when I try to open it. I closed it down, looked at the Logs, some kernel panic because of the API Key or so. Will investigate this later, would suck if it's broken, that's my whole DIY inventory :( anyway, this is at least settled!Based on what I saw, I would have guessed that one in particular was being problematic. I just wasn't 100% sure.1 minute ago, CameraRick said:Powertop I use for the mentioned power saving, or at least the attempt to do so. It's recommended a lot in this forum, didn't know it could lead to instabilities :( This is more of a general statement, enabling settings that is outside of stock unraid carries the potential to cause instability. This applies to anything the BIOS manufacturer didnt ship stock too. Going back to the defaults is just a step of troubleshooting to make sure the changes you made didnt affect it in some way/shape/form.
Monday at 09:15 PM2 days Author 3 hours ago, MowMdown said:Based on what I saw, I would have guessed that one in particular was being problematic. I just wasn't 100% sure.Forgot to mention: I found it because of your suggestion to see which Docker was started seconds ago.I had a look at the syslog now, the spam stopped 19:33h. At around 19:39h, I can see the messages for "unapply autotweak settings" when I uninstalled it, where all CPU settings went back to "performance", which huts a little :) not sure what I did in the minutes between those events, but my last post was from 19:40h, so it checks out3 hours ago, MowMdown said:This is more of a general statement, enabling settings that is outside of stock unraid carries the potential to cause instability. This applies to anything the BIOS manufacturer didnt ship stock too. Going back to the defaults is just a step of troubleshooting to make sure the changes you made didnt affect it in some way/shape/form.I understand. I don't have the powertop --auto-tune as part of my go-file, so it's currently not active. But I rarely had crashes during high system usage, mostly in idle. I also have ASPM helper installed, for the NVMe and NIC if I recall correctly. Is that also something I shouldn't be doing? I will wait for the parity check to be over (will take ~1½ days) till I can see how this affects the power consumption.I also had a look at my syslinux config:kernel /bzimage acpi_enforce_resources=laxappend initrd=/bzroot usb-storage.quirks=059f:105e:u The usb-storage entry is for two Backup-HDD-cases, which have a hard time mounting under UAS; this forces them into the legacy USB mode. The acpi-enforce is for reading out systemps (but I don't recall when I put that in, could be there for some time and may not be even needed anymore, not sure). Could those be issues?Also, I am still on unRAID 7.2.6, I am not that fast with updating. Maybe I should.Best,Rick
Join the conversation
You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.