Everything posted by Michael_P
-
OOM Errors started about 5 days ago
Try leaving it off for a bit and see if it stops going OOM, I've seen these processes in a couple OOM reports in the last week and both times it was nginx proxy manager (its node process in particular) Oct 9 08:07:39 unServerD kernel: [ 22401] 0 22401 320314 8283 7968 315 0 1306624 0 0 node Oct 9 08:07:39 unServerD kernel: [ 22447] 0 22447 13121 1127 961 166 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22448] 0 22448 13121 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22449] 0 22449 13188 1484 961 523 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22450] 0 22450 13153 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22451] 0 22451 13155 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22452] 0 22452 13112 1141 1041 100 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22453] 0 22453 13112 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22454] 0 22454 13078 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22455] 0 22455 13078 1015 961 54 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22456] 0 22456 13011 1463 961 502 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22457] 0 22457 12978 1321 961 360 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22458] 0 22458 13011 1123 961 162 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22459] 0 22459 13011 961 961 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22460] 0 22460 13011 965 961 4 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22461] 0 22461 13011 1067 961 106 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22462] 0 22462 13012 961 961 0 0 81920 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22463] 0 22463 13012 1126 961 165 0 81920 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22464] 0 22464 13011 961 961 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22465] 0 22465 13064 961 961 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22466] 0 22466 13054 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22467] 0 22467 13012 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22468] 0 22468 13011 1146 961 185 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22469] 0 22469 13011 961 961 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22470] 0 22470 13011 1314 961 353 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22471] 0 22471 13011 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22472] 0 22472 13011 961 961 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22473] 0 22473 13011 1234 1041 193 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22474] 0 22474 13088 1049 961 88 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22475] 0 22475 13088 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22476] 0 22476 13012 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22477] 0 22477 13046 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22478] 0 22478 13120 1582 961 621 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22479] 0 22479 13087 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22480] 0 22480 13046 961 961 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22481] 0 22481 13046 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22482] 0 22482 13087 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22483] 0 22483 13046 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22484] 0 22484 13089 1041 1041 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22485] 0 22485 13045 1237 1041 196 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22486] 0 22486 13121 961 961 0 0 86016 0 0 nginx Oct 9 08:07:39 unServerD kernel: [ 22487] 0 22487 12720 1555 835 720 0 81920 0 0 nginx
-
OOM Errors started about 5 days ago
Just for shits and giggles, put a memory limit on your nginx proxy manager and see if that gets killed before the whole host goes oom - start with 1GB
-
OOM Errors started about 5 days ago
Seeing as how the kernel was reaping different processes, that suggests it wasn't something running away and instead was just the sever being over-utilized - instead of trying to swap it out, just configure either your tdarr or plex to transcode to disk since that's where it will end up anyway with the swap file.
-
OOM Errors started about 5 days ago
Looks like you have tdarr and plex transcodes going at the same time, if they're both set to transcode to RAM try running it without tdarr for a bit and see if it stops running oom
-
/var/log is getting full
Do you have something bound to port 80?
-
Out of Memory Error From Fix Common Problems
Probably not related Go to the docker tab and toggle the advanced view and see which container has this ID (or close to it) 7d3934da0c4f - it may be nginx proxy manager as another user recently had the same node process running way going OOM and being killed. I can't offer any advice as to why it's doing that, I have been running it for years and have never seen it do that, best I can advise is to re-install it. You will need to restart the sever to clear the log You should also try to figure out what's going on with these, it's flooding your log Oct 3 14:37:35 Tower dhcpcd[1976]: br0: fe80::be24:11ff:fe4c:4bc3 is reachable again Oct 3 14:47:46 Tower dhcpcd[1976]: br0: fe80::be24:11ff:fe4c:4bc3 is unreachable Oct 3 14:47:46 Tower dhcpcd[1976]: br0: fe80::be24:11ff:fe4c:4bc3 is reachable again Oct 3 14:47:55 Tower dhcpcd[1976]: br0: fe80::be24:11ff:fe4c:4bc3 is unreachable Oct 3 14:48:40 Tower dhcpcd[1976]: br0: fe80::be24:11ff:fe4c:4bc3 is reachable again Oct 3 14:48:48 Tower dhcpcd[1976]: br0: fe80::be24:11ff:fe4c:4bc3 is unreachable Oct 3 14:49:22 Tower dhcpcd[1976]: br0: fe80::be24:11ff:fe4c:4bc3 is reachable again Oct 3 14:49:30 Tower dhcpcd[1976]: br0: fe80::be24:11ff:fe4c:4bc3 is unreachable Oct 3 14:49:30 Tower dhcpcd[1976]: br0: fe80::be24:11ff:fe4c:4bc3 is reachable agai
-
Parity errors but which files
The parity errors don't tell you if a drive is bad or not, either parity or array, the errors are the number of differences from what the parity drive has calculated for the value of a particular sector. The errors can be from bad RAM, cables, random corruption or other hardware issues, or usually power disruption while writing. You'll need to interrogate your disks to make sure they're OK hardware wise, then re-sync parity - hoping whatever was in those 87 sectors was just out of sync or not important if the data was actually corrupted. This is a good reminder that parity is for redundancy, not data protection. It has no idea what's on your disks, and can only rebuild a disk based on whatever sector info it's got stored.
-
Parity errors but which files
No, parity doesn't contain any file information, it only knows it was expecting a 1 and got a 0 from a particular sector on the disk, not what part of which file occupied that sector (it won't know which is the correct bit, the 1 on parity or the 0 from the array disk). If you've had an unclean shutdown, you can just run a correcting parity check to sync parity again, chance for corruption is lower. But, if it just suddenly started showing errors, you need to figure out why (bad disk, bad cable, bad RAM) before you continue.
-
Parity errors but which files
Unless you have checksums, you have no way of knowing which files are 'broken'. If you're showing errors, it's because the zeros and ones on the disk don't match what the parity calculation is returning.
-
MCE errors
Looks like maybe the RAM going bad, run memtest to see if it finds any errors
-
MCE errors
nothing in /var/log/mcelog ?
-
MCE errors
What's the MCE log say?
-
Out of memory error detected
Whatever container is running node - you can run ps -auxf to figure out which one is kicking it off. It'll be one you have memory limited, since the host itself isn't running OOM, the container is hitting whatever memory limit you have set for it
-
Out of memory error detected
Looks like back on the 22nd node ran over its memory limit and was killed, twice. You can allocate more memory to it, or change whatever you were doing to use less RAM. Reboot to clear the log so FCP stops warning you Sep 22 21:03:13 Mercury kernel: Memory cgroup out of memory: Killed process 24538 (node) total-vm:1292120kB, anon-rss:48364kB, file-rss:996kB, shmem-rss:0kB, UID:0 pgtables:1536kB oom_score_adj:0 Sep 22 21:54:07 Mercury kernel: Memory cgroup out of memory: Killed process 724928 (node) total-vm:1280704kB, anon-rss:37004kB, file-rss:652kB, shmem-rss:0kB, UID:0 pgtables:1276kB oom_score_adj:0
-
Seg faults on server
Check to make sure you didn't bend any pins on the the LGA
-
Seg faults on server
Try turning off your deluge container, too - it looks like that's what's running the python process
-
Seg faults on server
CPU could be faulty, have you removed it lately?
-
Seg faults on server
Figure out which container is running a python3 instance. running ps -auxf can help you narrow it down
-
Seg faults on server
If it's only segfaulting with the python script you're running, start there
-
Out Of Memory errors detected
Looks like your LLM container hit its limit back on the 26th and was killed. Reboot to clear the log so FCP stops warning you
-
High CPU usage in DOCKER
ps -auxf can give you a more granular report
-
High CPU usage in DOCKER
toggle the advanced view in the docker tab, that will show you how much CPU each container is using
-
Why are default permissions 777 and 666?
Samba shares probably
-
Older Server - Remote Access to Content
Depends, remote access for media or just for general file access?
-
Out Of Memory errors detected on your server
You'll need to reboot to clear the log or FCP will continue to warn you