May 1, 20206 yr I double checked in the config field, no there are no space, neither at the start nor at the end
May 1, 20206 yr if you look at the end of the docker error line device error: unknown device id: GPU-de8ab77e-8fff-db12-f93d-ebe991944a85\\n\ \\n\ is what is causing your issue
May 1, 20206 yr 1 hour ago, scottc said: if you look at the end of the docker error line device error: unknown device id: GPU-de8ab77e-8fff-db12-f93d-ebe991944a85\\n\ \\n\ is what is causing your issue I using the edit to put the data in .... and when putting the other UUID from the 1050, I get no errors ? how to remove this \\n\ ?
May 1, 20206 yr On 5/1/2020 at 9:26 AM, saarg said: Can you post your diagnostics and I'll see if I find anything there. Yes of of course please see the attached Edited May 7, 20206 yr by Solverz
May 1, 20206 yr 6 hours ago, chocorem said: I double checked in the config field, no there are no space, neither at the start nor at the end I'd you try to delete it and toe it in manually? You might not see that there is a space. There is a bug that makes the same issue when copy pasting from the forum and typing it manually works.
May 2, 20206 yr 15 hours ago, Solverz said: darth-nas-diagnostics-20200501-1952.zip 66.59 kB · 0 downloads Yes of of course please see the attached 1. Your bios is very old. You have 1.8.2 (From 2011) and the newest is 1.13.0 (2018). 2. You have this error in the syslog about the P400. pci 0000:04:00.0: can't claim BAR 3 [mem 0xbe000000-0xbfffffff 64bit pref]: no compatible bridge window kernel: pci 0000:04:00.0: BAR 3: no space for [mem size 0x02000000 64bit pref] kernel: pci 0000:04:00.0: BAR 3: trying firmware assignment [mem 0xbe000000-0xbfffffff 64bit pref] kernel: pci 0000:04:00.0: BAR 3: [mem 0xbe000000-0xbfffffff 64bit pref] conflicts with System RAM [mem 0x00100000-0xbf698fff] kernel: pci 0000:04:00.0: BAR 3: failed to assign [mem size 0x02000000 64bit pref] 3. You got the below lines in your syslog, which I don't have. Might be different revisions of the card or that your's a Dell card? kernel: nvidia: loading out-of-tree module taints kernel. kernel: nvidia: module license 'NVIDIA' taints kernel. 4. It seems it's a different GPU ID also. Yours is [drm] [nvidia-drm] [GPU ID 0x00000400] Loading driver And mine is [drm] [nvidia-drm] [GPU ID 0x00008200] Loading driver 5. You have the below when the plugin is installed, so there is something going on. Hardware/bios issue is my guess. kernel: resource sanity check: requesting [mem 0xdd700000-0xde6fffff], which spans more than PCI Bus 0000:04 [mem 0xdc000000-0xddffffff] kernel: caller _nv030928rm+0x5d/0xd0 [nvidia] mapping multiple BARs kernel: NVRM: GPU 0000:04:00.0: RmInitAdapter failed! (0x26:0xffff:1227) kernel: NVRM: GPU 0000:04:00.0: rm_init_adapter failed, device minor number 0 So try to update your bios and see if that works.
May 2, 20206 yr 24 minutes ago, saarg said: 1. Your bios is very old. You have 1.8.2 (From 2011) and the newest is 1.13.0 (2018). 2. You have this error in the syslog about the P400. pci 0000:04:00.0: can't claim BAR 3 [mem 0xbe000000-0xbfffffff 64bit pref]: no compatible bridge window kernel: pci 0000:04:00.0: BAR 3: no space for [mem size 0x02000000 64bit pref] kernel: pci 0000:04:00.0: BAR 3: trying firmware assignment [mem 0xbe000000-0xbfffffff 64bit pref] kernel: pci 0000:04:00.0: BAR 3: [mem 0xbe000000-0xbfffffff 64bit pref] conflicts with System RAM [mem 0x00100000-0xbf698fff] kernel: pci 0000:04:00.0: BAR 3: failed to assign [mem size 0x02000000 64bit pref] 3. You got the below lines in your syslog, which I don't have. Might be different revisions of the card or that your's a Dell card? kernel: nvidia: loading out-of-tree module taints kernel. kernel: nvidia: module license 'NVIDIA' taints kernel. 4. It seems it's a different GPU ID also. Yours is [drm] [nvidia-drm] [GPU ID 0x00000400] Loading driver And mine is [drm] [nvidia-drm] [GPU ID 0x00008200] Loading driver 5. You have the below when the plugin is installed, so there is something going on. Hardware/bios issue is my guess. kernel: resource sanity check: requesting [mem 0xdd700000-0xde6fffff], which spans more than PCI Bus 0000:04 [mem 0xdc000000-0xddffffff] kernel: caller _nv030928rm+0x5d/0xd0 [nvidia] mapping multiple BARs kernel: NVRM: GPU 0000:04:00.0: RmInitAdapter failed! (0x26:0xffff:1227) kernel: NVRM: GPU 0000:04:00.0: rm_init_adapter failed, device minor number 0 So try to update your bios and see if that works. Thank you so much for finding these issues. I am unsure if it is a Dell P400 or not as I just got it off of ebay. Would this cause any issues that you know of if it is a Dell p400? I also just found out that the Pci-e x16 slot on the motherboard is PCIE Gen2 and performs in x8 mode not x16.... Would this relate to any of the above issues from what you understand with your experience? I will also get the BIOs updated and report back Thanks again!!!
May 2, 20206 yr 57 minutes ago, Solverz said: Thank you so much for finding these issues. I am unsure if it is a Dell P400 or not as I just got it off of ebay. Would this cause any issues that you know of if it is a Dell p400? I also just found out that the Pci-e x16 slot on the motherboard is PCIE Gen2 and performs in x8 mode not x16.... Would this relate to any of the above issues from what you understand with your experience? I will also get the BIOs updated and report back Thanks again!!! I don't think there should be any issues if it's a dell p400. Running the card in 8x gen2 shoudln't be an issue as far as I know.
May 2, 20206 yr Just checking after reading the last 10 pages to verify. Is a dummy plug required for an Nvidia card being used exclusively for transcoding. Right now I have it set to work on the Plex docker from ls.io and I may also look at passing it over to a handbrake docker for retranscoding some media that quality isn't a huge concern (cartoons, etc.). Right now I don't have it installed, but I've only just switched over to the nvidia transcode yesterday. Don't want to hit any snares as we go and I can just order the dummy plug ASAP if I should use it. Thanks! Love the work on the nvidia support LS.io, greatly appreciated.
May 2, 20206 yr I don't think the card will initialize if it doesn't have some type of output detected.
May 2, 20206 yr Just now, david279 said: I don't think the card will initialize if it doesn't have some type of output detected. I have it installed and transcoding without any output currently. I just wasn't sure if I'd run into some sort of low power state issues or anything. I never even removed the plastic covers that come installed in the HDMI/Displayport connectors on the GPU when I installed it into my server.
May 2, 20206 yr 5 minutes ago, DaClownie said: I have it installed and transcoding without any output currently. I just wasn't sure if I'd run into some sort of low power state issues or anything. I never even removed the plastic covers that come installed in the HDMI/Displayport connectors on the GPU when I installed it into my server. Ok that's good to know. I know VMs with a passed through GPU that don't have anything connected will freak out sometimes.
May 5, 20206 yr Just a quick update on my 1660 super getting 'lost'.. I've now ruled out the HW (I think) - Booting directly to a windows 10 native install (I just removed the HBA controller and unRAID USB Stick and used a spare 120GB SSD with a fresh WIN 10 install) I have run GPU tools for 3 days solid with no issues seen - I've then booted to unRAID (non-nvidia) and passed it through as the primary GPU to a Win10 VM and that's run diagnostics for 2.5 days with no issues) All I can think is the 440.59 linux drivers don't sit nicely with my Asus 1660 Super OC Phoenix GPU / Ryzen 3600 / Asus B450F Motherboard. I guess in the spirit of this plugin all I can do now is not use the GPU until the next unraid build is released and hopefully the plugin can be updated to the latest drivers. I appreciate LinuxServer are busy so I can't expect anything more than they've stated. I have popped a +1 post in the feature request for native unRAID support for NVidia Drivers. I'll focus on fixing the one or two niggles I've not yet resolved and sit patiently with fingers crossed!
May 5, 20206 yr On 5/2/2020 at 6:00 AM, david279 said: Ok that's good to know. I know VMs with a passed through GPU that don't have anything connected will freak out sometimes. is this just the consumer grade cards or the P series as well?
May 7, 20206 yr On 5/2/2020 at 1:08 PM, saarg said: I don't think there should be any issues if it's a dell p400. Running the card in 8x gen2 shoudln't be an issue as far as I know. I put the p400 in another system and it is now working fine, so I think I am going to use this new system instead. However I noticed when rebooting the gpu UUID is not shown in the NVIDIA UNRAID settings and to get it to show I have to uninstall the NVIDIA UNRAID plug in and reinstall it. Any ideas?
May 7, 20206 yr 1 hour ago, Solverz said: I put the p400 in another system and it is now working fine, so I think I am going to use this new system instead. However I noticed when rebooting the gpu UUID is not shown in the NVIDIA UNRAID settings and to get it to show I have to uninstall the NVIDIA UNRAID plug in and reinstall it. Any ideas? Probably a race condition. You could use the command chbmb posted. I don't think the UUID will change so ou only need to get it once.
May 7, 20206 yr 2 minutes ago, saarg said: Probably a race condition. You could use the command chbmb posted. I don't think the UUID will change so ou only need to get it once. Yes that command works also thank you. Do you mind me asking what you mean by a race condition?
May 7, 20206 yr I hope you don't mind me asking... Is it possible to get the two kernel modules: 'joydev' and 'uinput' with this plugin/images? I recently released a container to play Steam games with In Home Streaming (even over the internet) inside a container but the only problem is that these two kernel modules are needed to enable 'real' controller support, now i have to map the controller buttons to keyboard or even mouse inputs and that's really frustrating for some games. EDIT: Not needed anymore, got it solved. Edited May 27, 20206 yr by ich777
May 8, 20206 yr thanks for work on this, I'm trying to get nvidia-smi to work in telegraf docker but just can't get it to run, it passes nvidia devices over as expected but just fails to run nvidia-smi altogether
May 8, 20206 yr 21 hours ago, Solverz said: Yes that command works also thank you. Do you mind me asking what you mean by a race condition? @saarg I looked into race condition and found out what it means. I also found out why the uuid was not showing. It was because jellyfin was set to autostart, when I disabled this it showed the gpu uuid again after and reboot Thanks for all your help!!!
May 8, 20206 yr 2 hours ago, Solverz said: @saarg I looked into race condition and found out what it means. I also found out why the uuid was not showing. It was because jellyfin was set to autostart, when I disabled this it showed the gpu uuid again after and reboot Thanks for all your help!!! I didn't see your post until now. It's weird jellyfin would make the UUID disappear.
May 8, 20206 yr Hi All, Just recently installed this (the 6.8.3 nvidia build) for the first time and I'm getting the "NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running. " error. I have two cards, the first is a GT 240 for console access and wouldn't expect to show up, but the second card is a 1660 GTX. I previously used the card for gpu passthrough but its used far and few so I'd rather repurpose it. I removed the stub on the card and have disabled the VM manager. After rebooting I'm still not seeing the card show up in the Nvida Build settings. Is there anything else I need to disable besides the stub and the vm manager? I believe its still in the XML for the VM however I would assume with the VM manager disabled it shouldn't be holding on to the card still no? Any help is appreciated! Thanks, ~chiefo Edited May 8, 20206 yr by chiefo
May 8, 20206 yr 32 minutes ago, chiefo said: Hi All, Just recently installed this (the 6.8.3 nvidia build) for the first time and I'm getting the "NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running. " error. I have two cards, the first is a GT 240 for console access and wouldn't expect to show up, but the second card is a 1660 GTX. I previously used the card for gpu passthrough but its used far and few so I'd rather repurpose it. I removed the stub on the card and have disabled the VM manager. After rebooting I'm still not seeing the card show up in the Nvida Build settings. Is there anything else I need to disable besides the stub and the vm manager? I believe its still in the XML for the VM however I would assume with the VM manager disabled it shouldn't be holding on to the card still no? Any help is appreciated! Thanks, ~chiefo You can check which module is loaded for the 1660 with lspci -k. There have been others with a 1660 that have had issues getting the card recognised. Might be a driver version issue.
May 8, 20206 yr 5 minutes ago, saarg said: You can check which module is loaded for the 1660 with lspci -k. There have been others with a 1660 that have had issues getting the card recognised. Might be a driver version issue. I'll check that out in a minute. Ran lspci -v and saw it was still being stubbed. Turns out while most of the components ended in 10d<letter> there was another one still stubbed that I thought was from the wireless card I had stubbed. Servers rebooting now. *EDIT* Yeah jumped the gun. Its showing up now that i removed the 4th stub item that was hanging behind Edited May 8, 20206 yr by chiefo More info
May 9, 20206 yr 18 hours ago, saarg said: I didn't see your post until now. It's weird jellyfin would make the UUID disappear. I know, jellyfin only makes it disappear when jellyfin is set to autostart and the server reboot with this setting on. The gpu is assigned to the jellyfin container so not sure if anything is going on when the server boots with jellyfin autostarting? Maybe? 😅 Edited May 9, 20206 yr by Solverz
Archived
This topic is now archived and is closed to further replies.