Skip to content
View in the app

A better way to browse. Learn more.

Unraid

A full-screen app on your home screen with push notifications, badges and more.

To install this app on iOS and iPadOS
  1. Tap the Share icon in Safari
  2. Scroll the menu and tap Add to Home Screen.
  3. Tap Add in the top-right corner.
To install this app on Android
  1. Tap the 3-dot menu (⋮) in the top-right corner of the browser.
  2. Tap Add to Home screen or Install app.
  3. Confirm by tapping Install.

rutherford

Members
  • Joined

  • Last visited

  1. The facial recognition wrapped up over night. Immich version 3.2.2, unraid 7.3.2. from the immichdb queries: face_count is 273k, person_count is 10k We are stretching a bit past the limits of my skill here! I've attached three files that I hope answer your request there Jorge "timestamped server logs and memory readings from your current test." I also got a related, and possibly helpful, response over at Immich Discord: oom-memory-windows.tsv first-oom-immich-server.log oom-events.txt
  2. The below is AI slop, but I'm working outside of what I'm capable of here. I hope some of this helps someone. It's not an answer, it's where I'm at right now. chatgpt AI status update Update / more troubleshooting: I've brought Immich back online and have been reproducing the facial-recognition rebuild with considerably more instrumentation. I also put hard Docker memory limits on the Immich containers so that, hopefully, I can observe the failure without allowing one container to consume the entire host. So far this is pointing much more strongly toward memory growth in immich_server during the facial-recognition rebuild. immich_machine_learning does not appear to be running away. Its cgroup memory peaked around 2.6 GB and has since fallen toward ~1 GB while facial recognition continues processing. There have been no ML cgroup OOM events. immich_server is a very different story. I initially limited it to 6 GB. It reached that ceiling and the cgroup recorded 1,106 memory.max events, although there were no OOM/OOM-kill events. I raised the limit to 10 GB and subsequently 12 GB so I could continue observing it without immediately killing the workload. Once given additional headroom, immich_server continued growing steadily rather than immediately leveling off. Over one measured ~4.5-minute interval its cgroup memory increased from roughly 6.3 GB to 8.0 GB. A few minutes later it was around 9.6 GB. Importantly, most of the increase is anonymous/application memory rather than just reclaimable file cache — at ~9.6 GB cgroup usage, approximately 7.1 GB was anonymous memory. At the same time, the host itself remains responsive. It still had roughly 14 GB MemAvailable, no swap, CPU was mostly idle, and I/O wait was around 6–8%. The NVMe cache devices have not shown anything resembling 100% utilization or huge latency, there are no persistent D-state processes in the spot checks, and I've seen no OOM, hung-task, BTRFS error, NVMe reset/timeout, or I/O-error messages in the live kernel log. Facial Detection and Facial Recognition job concurrency are both set to 1. This makes my current working hypothesis that the original "whole server grinds to a halt" behavior is caused by immich_server continually increasing its memory footprint during a large facial-recognition rebuild. Without a Docker memory limit, it may eventually consume enough of this 32 GB machine to put the host into severe reclaim/thrashing. That would also explain why I wasn't necessarily seeing a clean OOM kill before the machine became effectively unusable. I don't want to label this a confirmed Immich memory leak yet because I'm still letting the rebuild run to see whether the memory eventually plateaus. For now I've left immich_server capped at 12 GB and immich_machine_learning at 10 GB, and the Unraid server is still snappy while the rebuild continues. I'm continuously logging host memory/reclaim stats, vmstat, iostat, Docker stats, D-state/top processes, kernel messages, ML/server cgroup memory (including anon/file breakdown and OOM/max events), and the Immich server/ML logs. So if it deteriorates again I should have much better evidence than I did during the previous crashes.
  3. I have been tooting through some troubleshooting with chatgpt to get some initial troubleshooting done. I did a full reset of facial recognition with immich, and I think that's what initially led to instability and crashing. My system as 32GB DDR5 RAM, 2GB reserved for a Homeassistant VM. That's off. I stopped all the immich dockers (four of them: server, machine_learning, redis, postgres) and I'll leave those off for now. I did turn on mirror logs to flash drive, looks like log_previous is in cluded in this attached diagnostics zip. Crafty 4 (minecraft server) is also a memory hog, as well as the seafile-elasticsearch docker. Those are all stopped now as well. If this is immich slowly gobbling up all the memory and crashing the server, well that shouldn't be happening. I can start to chase that down over at the immich forums, or here. Thanks in advance! mayorgoodway-diagnostics-20260921-1046.zip
  4. Welcome Jake! I think we'll get on famously.
  5. I love our forums, and our community. It's what steered me to unraid on day one from a couple other options at the time. AND: no shade to our newly formed paid experts! I've used them once before and it was 100% worth it. You know when you're using chatgpt or claude, you can tell it to "remember this conversation"? What if we had our own claude/chatgpt that remembered everything from all the users talking with it. It would know all the official docs, and put much more weight on the official docs, less weight on the wiki, and less weight again on other forum entries. it can point directly to other users or posts that have had this problem too maintain it's own documentation for those who'd like the written word vs the interactive type suggest edits to the wiki and official documentation (a wish for all chatbots) get to a point where it would say "we're diving too deep into configuration modification. I suggest you make a post on the forums and wait for a human to get back to you. Here's the meat and potatoes of what we've been working on [code data]." Something like unraid.claude.ai that users can share in the cost; or maintain that free query/day thing. Or better yet, robochat.unraid.com for a self-hosted thing!
  6. I'd put all that hardware into pcpartpicker and see if it flags any compatibility stuff. Also do a search on pcpartpicker for anyone who's built a computer into that Jonsbo chassis. They'll frequently write up issues they had with the build, things to look out for. I'd get a better power supply. Something more name brand. https://www.newegg.com/p/pl?d=500+watt+psu Seasonic, Thermaltake, and Corsair all jump out at me from this US based place. RAM pricing is stupid right now. How about building for DDR4? Three windows 11 VMs, wow! I've never put my system through that kind of ringer. This will be your primary machine I take it? I couldn't say how the CPU will hold up. When it comes to power requirements, I think pcpartpicker has that built in. But it wouldn't hurt to check it. Make sure your added up wattage comes to something like 80% of what that PSU can handle. There is probably a more smart percentage to work with, you'd have to google that. Post some pics when it's done!
  7. @bmartino1 I appreciate your efforts in this community. I'm positive your contributions have helped a lot of folks.
  8. I don't think this works for the machine learning OCR right? That has different steps over here: https://docs.immich.app/features/ml-hardware-acceleration/
  9. Run to failure? I do have one extra SATA wire coming off my HBA I could swap. SMART Attributes Data Structure revision number: 10 Vendor Specific SMART Attributes with Thresholds: ID# ATTRIBUTE_NAME FLAGS VALUE WORST THRESH FAIL RAW_VALUE 1 Raw_Read_Error_Rate POSR-- 078 064 044 - 63435056 3 Spin_Up_Time PO---- 091 089 000 - 0 4 Start_Stop_Count -O--CK 099 099 020 - 2029 5 Reallocated_Sector_Ct PO--CK 100 100 010 - 0 7 Seek_Error_Rate POSR-- 086 060 045 - 386327305 9 Power_On_Hours -O--CK 066 066 000 - 30394 10 Spin_Retry_Count PO--C- 100 100 097 - 0 12 Power_Cycle_Count -O--CK 100 100 020 - 70 18 Head_Health PO-R-- 100 100 050 - 0 187 Reported_Uncorrect -O--CK 100 100 000 - 0 188 Command_Timeout -O--CK 100 100 000 - 0 190 Airflow_Temperature_Cel -O---K 053 044 040 - 47 (Min/Max 40/48) 192 Power-Off_Retract_Count -O--CK 100 100 000 - 92 193 Load_Cycle_Count -O--CK 054 054 000 - 93516 194 Temperature_Celsius -O---K 047 056 000 - 47 (0 15 0 0 0) 195 Hardware_ECC_Recovered -O-RC- 009 001 000 - 63435056 197 Current_Pending_Sector -O--C- 100 100 000 - 0 198 Offline_Uncorrectable ----C- 100 100 000 - 0 199 UDMA_CRC_Error_Count -OSRCK 200 200 000 - 142 200 Pressure_Limit PO---K 100 100 001 - 0 240 Head_Flying_Hours ------ 100 253 000 - 7667h+33m+43.187s 241 Total_LBAs_Written ------ 100 253 000 - 18003672784 242 Total_LBAs_Read ------ 100 253 000 - 1269847927595 ||||||_ K auto-keep |||||__ C event count ||||___ R error rate |||____ S speed/performance ||_____ O updated online |______ P prefailure warning mayorgoodway-diagnostics-20251126-1052.zip
  10. I've got a pikvm hooked up via HDMI to an Sparkle Intel Arc A310 ECO. The BIOS post works, I changed a few settings in the BIOS set to make this happen: Primary video: IGFX Video Integrated Graphics: Force UMA Frame Buffer: 1G When it was posting I could see BOTH the pikvm and H5Viewer. As soon as unRaid started to boot and go through it's flipping text stuff, H5Viewer went black, but the pikvm kept showing stuff. I made it to the command line, but I can only see the prompt on the pikvm, nothing on the IPMI H5Viewer. Oddly, it will take my keystrokes, I can see evidence of that on the pikvm going to next line with empty input for login: The CPU is AMD Ryzen 7 7700 8-Core. I wanted to check here with humans before going down the AI rabbit hole with editing odd files forcing this thing to display on the right internal GPU. root@mayorgoodway:~# lspci -nnk | grep -i vga -A3 07:00.0 VGA compatible controller [0300]: Intel Corporation DG2 [Arc A310] [8086:56a6] (rev 05) Subsystem: Device [172f:4019] Kernel driver in use: i915 Kernel modules: i915, xe -- 0a:00.0 VGA compatible controller [0300]: ASPEED Technology, Inc. ASPEED Graphics Family [1a03:2000] (rev 52) Subsystem: ASPEED Technology, Inc. ASPEED Graphics Family [1a03:2000] Kernel driver in use: ast Kernel modules: ast -- 11:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Raphael [1002:164e] (rev c5) Subsystem: ASUSTeK Computer Inc. Device [1043:8877] Kernel driver in use: amdgpu Kernel modules: amdgpu root@mayorgoodway:~# ls /dev/dri by-path/ card0 card1 card2 renderD128 renderD129 root@mayorgoodway:~# mayorgoodway-diagnostics-20251123-1444.zip
  11. I'm also looking into this. How about offloading that to an Intel ARC A310... looks like it should work! https://www.reddit.com/r/unRAID/comments/1n2m7h7/comment/nbc5phh/
  12. @Cuissedemouche figure out what the name of the owner of that process is and do a chown on the folder. sudo chown -R postgres_user:postgres_user $PGDATA/globalIt's odd that it didn't initialize correctly - was it running fine for a while, but then started throwing this error? Or has it been this error since the beginning? Might be worth backing up the database, deleting the /mnt/user/appdata/immich folder, then reinstalling the docker and postgres database. Are you running two dockers: Immich and the PostgresDB docker? Shouldn't the PostgresDB docker have it's appdata in a different folder from the immich appdata folder?
  13. AI pulled me through this one. For those of you who hate it (I almost do hate it!) Here's what I did to make it start working. I was coming from an nvidia gpu. When I pulled the card, Plex docker (lscr.io/linuxserver/plex) stopped working. I had to remove all the custom entries I'd put in years ago for that nvidia card - including a --runtime=nvidia thing that was behind the Advanced toggle. AI wanted me to change my Network from Host to Bridge. That broke port forwarding and I couldn't connect to my Plex server, so I changed that back to Host. I had to add two Devices to the docker template. Add another Path, Port, Variable, Label or Device Select Device Name: /dev/dri/card0 Value: /dev/dri/card0 If you want to make sure these things are present, like I did, you can run: root@mayorgoodway:~# ls -la /dev/dri total 0 drwxrwxrwx 3 root root 100 Oct 1 03:42 ./ drwxr-xr-x 17 root root 4080 Oct 1 10:46 ../ drwxrwxrwx 2 root root 80 Oct 1 03:42 by-path/ crw-rw---- 1 root video 226, 0 Oct 1 10:44 card0 crwxrwxrwx 1 root video 226, 128 Oct 1 03:42 renderD128 root@mayorgoodway:~#(formatting blah) (formatting blah) (formatting blah) I did add this into Extra-parameters "--group-add 44". This is what AI said about it: C. Add the video Group Access Docker containers often need to be part of the correct host group to access the GPU device nodes. Since your ls -la output shows the devices belong to the video group, you must grant the Plex container access to this group. AI asked me to double check I had some boxes ticked in the Plex settings. They were already checked because of the previous nvidia stuff I had in there. They were: Go to Settings (wrench icon) → Transcoder. Check the box for "Use hardware acceleration when available". Check the box for "Use hardware-accelerated video encoding". Click Save Changes. I put on a South Park, forced it to transcode, and badda bing: (HW) showed up and I could see the card doing it's thing. Doing it's thing. I optionally installed a Community Apps > "Intel-GPU-TOP; ich777" This added a command line interface program called intel_gpu_top. That was slick for watching the card do it's thing. it cal also give you a list of your available devices. Here's mine root@mayorgoodway:~# intel_gpu_top -L card0 Intel Dg2 (Gen12) pci:vendor=8086,device=56A6,card=0 └─renderD128 root@mayorgoodway:~# So anyhow, no really biggie. Ah another slightly unrelated thing. I have a PiKVM v3 Pre-assembled to help manage this headless server. Because my CPU has no onboard GPU, the HDMI on my motherboard is dead. When I put the new Sparkle Intel Arc A310 ECO in there, and it detected there was no monitor attached: no POST/boot. I had to swipe the monitor mouse and keyboard from my gaming rig to get the unraid server to BOOT. I ordered one of those dummy HDMI things. I also posted over at pishop.com forums about WTH. I'll probably end up buying the V4 of the PiKVM so I can still remote access this puppy.
  14. any updates here @Jltoro ??
  15. I was looking all over the place for where to post this comment and even here, doesn't seem like quite the right spot. I don't particularly want to participate in the YouTube comments, reddit is fine, these forums are better. Tiffany Jones said she wanted to hear from us if we liked this content and I wanted to say Yes! Inlike this content! I enjoy listening to you guys while I'm out for runs, then I'll swap over to some Zeppelin on the way back home. I've been with you guys for years and years. It's been refreshing to see new faces. I like and agree with the direction you're taking the company. Spending some time doing things like this, IMO, is a correct way to spend some of your time. I'm excited to continue to see the software refined, simplified and enhanced. I use it daily so it's a big deal! How about some stats on which dockers are the most popular? How about exploring some coordinated relationship with those popular, or cool, or interesting would be partners? Then again, maybe that's the beauty of docker 😉 Keep em coming, One of your fans, drew

Account

Navigation

Search

Search

Configure browser push notifications

Chrome (Android)
  1. Tap the lock icon next to the address bar.
  2. Tap Permissions → Notifications.
  3. Adjust your preference.
Chrome (Desktop)
  1. Click the padlock icon in the address bar.
  2. Select Site settings.
  3. Find Notifications and adjust your preference.