April 27, 20206 yr Hi All, I am new to this forum so please if you need anymore information please let me know and I will post it, as I am unsure what information will be helpful I have the unraid nividia version installed but my Quadro P400 does not show in the UNRAID NVIDIA area of settings. Yet I am able to see it in System Devices and pass it through to a VM. I have not tried passing it through to a docker as I do not know what the UUID is as it does not show up in the NVIDIA UNRAID area of setting. Please can you help with this? If you need anymore information please let me know My server is a Dell Poweredge t310 btw!
April 27, 20206 yr 43 minutes ago, Solverz said: Hi All, I am new to this forum so please if you need anymore information please let me know and I will post it, as I am unsure what information will be helpful I have the unraid nividia version installed but my Quadro P400 does not show in the UNRAID NVIDIA area of settings. Yet I am able to see it in System Devices and pass it through to a VM. I have not tried passing it through to a docker as I do not know what the UUID is as it does not show up in the NVIDIA UNRAID area of setting. Please can you help with this? If you need anymore information please let me know My server is a Dell Poweredge t310 btw! The problem is that you are passing it through your a vm. That means the nvidia driver is not loaded, and the plugin doesn't see the card.
April 27, 20206 yr 13 minutes ago, saarg said: The problem is that you are passing it through your a vm. That means the nvidia driver is not loaded, and the plugin doesn't see the card. Sorry I forgot to state that I only tested if it could be passed through to a VM after seeing that it was not visible in NVIDIA UNRAID settings. It is currently not being passed to a VM at all but is visible in System Devices, However it does not show in NVIDIA UNRAID settings still.
April 27, 20206 yr 1 hour ago, Solverz said: Sorry I forgot to state that I only tested if it could be passed through to a VM after seeing that it was not visible in NVIDIA UNRAID settings. It is currently not being passed to a VM at all but is visible in System Devices, However it does not show in NVIDIA UNRAID settings still. Is it chosen in the VM template? If so, unraid automatically binds the GPU to vfio so the Nvidia diver can't be used.
April 27, 20206 yr 1 minute ago, saarg said: Is it chosen in the VM template? If so, unraid automatically binds the GPU to vfio so the Nvidia diver can't be used. No it is not chosen in any vm templates in fact I don't have any vm templates at all. I also noticed when I remove the gpu and go into nividia unraid settings it states the driver could not be loaded. But when I have the gpu installed it just lists the nvidia driver version but no information about the gpu at all.
April 28, 20206 yr I just got notified that the plugin is not known to the community apps plugin or the fix common problems plugin, and when I search for it now, it doesn't show up in the list of plugins that can be downloaded.... When I search for "NVIDIA" in the community apps plugin, it no longer shows up... Rebooting the server, and uninstalling the GPU statistics plugin changed nothing... Did this plugin just go unsupported? or is there a glitch in the community apps plugin? if it is now unsupported, I will be very sad to see it go... 😢 Still thank you to the author no matter which is the case...
April 28, 20206 yr 2 hours ago, Warrentheo said: I just got notified that the plugin is not known to the community apps plugin or the fix common problems plugin, and when I search for it now, it doesn't show up in the list of plugins that can be downloaded.... When I search for "NVIDIA" in the community apps plugin, it no longer shows up... Rebooting the server, and uninstalling the GPU statistics plugin changed nothing... Did this plugin just go unsupported? or is there a glitch in the community apps plugin? if it is now unsupported, I will be very sad to see it go... 😢 Still thank you to the author no matter which is the case... Glitch on github
April 28, 20206 yr 23 hours ago, aptalca said: I'm kinda seeing a trend here. Most if not all of the people experiencing these issues are also using the gpu stats plugin. Did you try without it? Thanks for the help. I uninstalled the GPU Stats plugin and rebooted. The issue happened again within 10 minutes (Checking with nvidia-smi I get the same GPU Lost message) I rebooted, and it maybe lasted 4 hours or so before happening again. I've used the nvidia-bug-report.sh that is mentioned when nvidia-smi loses the GPU and also carefully checked the syslog 1. Despite my GTX1660 Super being on the 440.59 supported list (checked on nvidia.com). the nvidia-bug-report.log files states "WARNING: You do not appear to have an NVIDIA GPU supported by the 440.59 NVIDIA Linux graphics driver installed in this system".. 2. Trawling through various logs, I found the error code XID79 just before the GPU went missing on one occasion, on the Nvidia developer site, this unfortunately can be attrutable to pretty much anything, HW error, Driver Error, Temperature etc.. 3. I've been checking the temperatures / HW state of the card, after boot it's in P0 (12W out of 125W) @ 33C, it them occasionally bumps up to P0 (26W/125W)@44C, so even when plex uses the card, 44C is barely ticking over, so pretty sure it's not temperature. 4. I think (looking at logs) there could possibly be some correlation between drives spinning down and the GPU crashing (or it may well be coincidence), I would like to try bulk spinning down/up the drives to see if power spikes might be upsetting the GPU, as I know HDD's draw the most power when they are spinning up.. [edit] - I found example user scripts to spin down/up all disks and tried those several times, whilst the GPU is idle and whilst transcoding a 4k HDR, no issues found. 5. I did at some point (more of a quick trial) have some User Scripts to 'tweak the driver for obvious reasons' and also to bump the card back to it's lowest power setting.. I haven't had these enabled for some time, so I've deleted the scripts entirely , and re-installed the unRAID-NVidia 6.8.3 from the plugin just to 'clear' things out.. 6. with 100% repeatability, I can trigger the "caller _nv000908rm+0x1bf/0x1f0 [nvidia] mapping multiple BARS" and hte associated memory spanning message by just running nvidia-smi to check the GPU is still there, I did this every 5-10 minutes over lunch and everytime I get an associated message in syslog. So nothing conclusive yet, some observations, some clutching at straws, but I sense maybe some experimentation and discussion might prompt something of note.. One test I 'may' do is to go back to the normal unRAID build, and pass the GPU through to my windows 10 VM (it's only spun up once in a blue moon) and run something GPU intensive on that and see if it ever loses the GPU, whilst this is changing a few too many variables at once, it would at least indicate the HW itself is OK (Power/Temperature concerns etc).. Edited April 28, 20206 yr by Snubbers
April 28, 20206 yr Hey All, Trying to sort out an issue I'm experiencing after moving to some new hardware. I've got a P2000 I had been utilizing in my old system. I just recently upgraded the board/cpu/memory and am re-using some of the PCIe cards in the new configuration, one of which is the P2000. The new build is a 3950X in a ASRock Rack X470D4U Everything appears to be working correctly as far as the PCIe cards go, however I can't seem to get the P2000 to pull back in like it had been previously. Prior to swapping everything out I rolled back to 6.8.2 stock, and removed the additional configurations in the docker container that were pointing to the NVIDIA device. My system boots, and I can see the P2000 under Tools > System Devices I have tried reinstalling both 6.8.2 as well as 6.8.3 with the NVIDIA drivers and I see the same error for both: Running 'nvidia-smi' I see the same error in the console and the below error in the system log: The one thing I haven't tried is passing it to a VM yet, though I know this bypasses the need for the NVIDIA driver as it relates to the docker container so it may be a moot point to test. Reading back a bit it seems a few people have had similar issues but either haven't resolved it or I'm missing their fix. Anyone able to provide a little further insight into this I'd appreciate it Also attaching a diagnostic unraid-diagnostics-20200428-1346.zip
April 28, 20206 yr 21 hours ago, Solverz said: No it is not chosen in any vm templates in fact I don't have any vm templates at all. I also noticed when I remove the gpu and go into nividia unraid settings it states the driver could not be loaded. But when I have the gpu installed it just lists the nvidia driver version but no information about the gpu at all. You were the one that said you passed it through to a VM. I assume you removed the VM template then. Have you rebooted after you remove the VM template? If you haven't rebooted, do so and post the output of lspci -k. If you have rebooted, still post the output of of the above command.
April 28, 20206 yr On 4/28/2020 at 7:35 PM, saarg said: You were the one that said you passed it through to a VM. I assume you removed the VM template then. Have you rebooted after you remove the VM template? If you haven't rebooted, do so and post the output of lspci -k. If you have rebooted, still post the output of of the above command. Appologies for the confusion, all I did was click the + icon to add a new vm and then i was able to select the gpu from there, I did not actually create a vm. I just checked if the option to pass the GPU through to the vm was available. I have attached the results of lspci -k in the .txt file. Appreciate your help!!! Edited May 7, 20206 yr by Solverz
April 28, 20206 yr On 4/25/2020 at 6:18 AM, saarg said: You can't as we only use the latest driver at the time of building the new build. So even though the 340.108 version came out in Dec of 2019, you are only using the 400 series drivers? With your experience do you think there is another way for me to make these work? LINUX X64 (AMD64/EM64T) DISPLAY DRIVER Version:340.108 Release Date:2019.12.23 Operating System:Linux 64-bit Language:English (US) File Size:66.92 MB
April 28, 20206 yr On 4/25/2020 at 5:40 PM, beardymcgee said: Grid k2 cards aren't compatible with anything other then vmware, they dont have standard linux drivers, its a vmware esx only card Nvidia does have drivers for linux for these cards. I'm just having issues trying to get them added. LINUX X64 (AMD64/EM64T) DISPLAY DRIVER Version:340.108 Release Date:2019.12.23 Operating System:Linux 64-bit Language:English (US) File Size:66.92 MB
April 28, 20206 yr 43 minutes ago, Solverz said: Appologies for the confusion, all I did was click the + icon to add a new vm and then i was able to select the gpu from there, I did not actually create a vm. I just checked if the option to pass the GPU through to the vm was available. I have attached the results of lspci -k in the .txt file. Appreciate your help!!! Results.txt 7.2 kB · 0 downloads The correct modules are loaded, so it should work. Is the card recognized if you run the command nvidia-smi on the comman line? If it is, then it's just the command the plugin runs at boot to find the UUID of the card that fails for some reason. There is a command you can run to get the UUID that was posted by chbmb earlier in this thread you can try.
April 29, 20206 yr Hi all, So, I've recently upgraded my server. The RAM I put in turned out to be bad (failed Memtest miserably, oof) and all hell kinda broke loose with the server (cache drive wouldn't mount, got weird bugs when plugged into the monitor, network connection with web ui and telnet were super unstable, other stuff). Anyways, I reinstalled the old RAM and things are functioning as normal except my syslog reveals the following repeated bit of text: Apr 29 13:33:51 TheShire kernel: caller _nv000908rm+0x1bf/0x1f0 [nvidia] mapping multiple BARs Apr 29 13:33:53 TheShire kernel: resource sanity check: requesting [mem 0x000c0000-0x000fffff], which spans more than PCI Bus 0000:00 [mem 0x000c0000-0x000dffff window] Now, it was happening every 2 seconds when I had GPU stats installed. It stopped when I uninstalled. But I ran a "watch nvidia-smi" in terminal and it reappeared (bc of course it did). I've poked around the forums but haven't seen anything conclusive on: (a) what this bit of text means or (b) how to fix it. I'm hoping to find answers to both. Thanks in advance for your responses, syslog attached (from both today and yesterday when stuff was going haywire, just in case). p.s. Maybe better for a different thread, but here goes: given the snafu I had yesterday, would it be advisable to backup my stuff and do a clean install of unraid? I'd like to avoid that if possible. SysLog04292020.rtf theshire-diagnostics-20200428-2239.zip
April 30, 20206 yr On 4/28/2020 at 11:11 PM, saarg said: The correct modules are loaded, so it should work. Is the card recognized if you run the command nvidia-smi on the comman line? If it is, then it's just the command the plugin runs at boot to find the UUID of the card that fails for some reason. There is a command you can run to get the UUID that was posted by chbmb earlier in this thread you can try. Sorry for late response. I have just ran the command "nvidia-smi" and the results is "No devices were found". If the below command is what you mean by chbmb, then it gave the same result "No devices were found". nvidia-smi --query-gpu=gpu_name,gpu_bus_id,gpu_uuid --format=csv,noheader | sed -e s/00000000://g | sed 's/\,\ /\n/g'
April 30, 20206 yr 4 hours ago, Solverz said: Sorry for late response. I have just ran the command "nvidia-smi" and the results is "No devices were found". If the below command is what you mean by chbmb, then it gave the same result "No devices were found". nvidia-smi --query-gpu=gpu_name,gpu_bus_id,gpu_uuid --format=csv,noheader | sed -e s/00000000://g | sed 's/\,\ /\n/g' Then I don't know what is happening. I have a p400 also, but no issues at all. Do you have anything in /dev/dri? Have you tried using an older version of unraid o see if it works there?
April 30, 20206 yr 4 hours ago, saarg said: Then I don't know what is happening. I have a p400 also, but no issues at all. Do you have anything in /dev/dri? Have you tried using an older version of unraid o see if it works there? Yes I tried the previous 2 versions of unraid nvidia and same result. The below are in the /dev/dri/ directory. by-path/ card0 renderD128
April 30, 20206 yr 3 hours ago, Solverz said: Yes I tried the previous 2 versions of unraid nvidia and same result. The below are in the /dev/dri/ directory. by-path/ card0 renderD128 Are you using gui boot?
April 30, 20206 yr 2 minutes ago, saarg said: Are you using gui boot? I am not booting to gui mode no, it is running headless so I just let it boot to unraid without the gui
May 1, 20206 yr Try plugging in the HDMI on the GPU to a monitor or something. It may need to have some type of output for it to work. Try a Dummy HDMI adapter.
May 1, 20206 yr 6 hours ago, david279 said: Try plugging in the HDMI on the GPU to a monitor or something. It may need to have some type of output for it to work. Try a Dummy HDMI adapter. I don't have a dummy or monitor plugged in, so not shure if that is the case, but it's worth a try.
May 1, 20206 yr 10 hours ago, Solverz said: I am not booting to gui mode no, it is running headless so I just let it boot to unraid without the gui Can you post your diagnostics and I'll see if I find anything there.
May 1, 20206 yr I have a weird behaviour , I have a quadro P400 and a 1050ti in the system Nvidia Driver Version: 440.59 GPU 0 Model & Bus: Quadro P400 23:00.0 GPU 0 UUID: GPU-de8ab77e-8fff-db12-f93d-ebe991944a85 GPU 1 Model & Bus: GeForce GTX 1050 Ti 2D:00.0 GPU 1 UUID: GPU-e60379bf-191f-14ec-3841-d4dfd8e82ab8 All is working fine when passing through the 1050TI to the plex container, but when trying to pass the quadro, I get this /usr/bin/docker: Error response from daemon: OCI runtime create failed: container_linux.go:346: starting container process caused "process_linux.go:449: container init caused "process_linux.go:432: running prestart hook 0 caused \"error running hook: exit status 1, stdout: , stderr: nvidia-container-cli: device error: unknown device id: GPU-de8ab77e-8fff-db12-f93d-ebe991944a85\\n\""": unknown. The command failed.
May 1, 20206 yr 11 minutes ago, chocorem said: I have a weird behaviour , I have a quadro P400 and a 1050ti in the system Nvidia Driver Version: 440.59 GPU 0 Model & Bus: Quadro P400 23:00.0 GPU 0 UUID: GPU-de8ab77e-8fff-db12-f93d-ebe991944a85 GPU 1 Model & Bus: GeForce GTX 1050 Ti 2D:00.0 GPU 1 UUID: GPU-e60379bf-191f-14ec-3841-d4dfd8e82ab8 All is working fine when passing through the 1050TI to the plex container, but when trying to pass the quadro, I get this /usr/bin/docker: Error response from daemon: OCI runtime create failed: container_linux.go:346: starting container process caused "process_linux.go:449: container init caused "process_linux.go:432: running prestart hook 0 caused \"error running hook: exit status 1, stdout: , stderr: nvidia-container-cli: device error: unknown device id: GPU-de8ab77e-8fff-db12-f93d-ebe991944a85\\n\""": unknown. The command failed. It looks like you have some empty spaces at the end of the UUID.
Archived
This topic is now archived and is closed to further replies.