@Leoyzen First of all thanks for your work to provide custom kernels. 👍
Let me explain the issue i have. I upgraded from a TR4 1950x to a Threadripper 3960x on an Gigabyte Aorus Extreme TRX40 and kinda see the same issues people having with passing through the onboard audio. In my case it's the following device:
IOMMU group 42: [1022:1487] 23:00.4 Audio device: Advanced Micro Devices, Inc. [AMD] Starship/Matisse HD Audio Controller
root@UNRAID:~# lspci -v -s 23:00.4
23:00.4 Audio device: Advanced Micro Devices, Inc. [AMD] Starship/Matisse HD Audio Controller
Subsystem: Advanced Micro Devices, Inc. [AMD] Device d102
Flags: bus master, fast devsel, latency 0, IRQ 154
Memory at b1400000 (32-bit, non-prefetchable) [size=32K]
Capabilities: [48] Vendor Specific Information: Len=08 <?>
Capabilities: [50] Power Management version 3
Capabilities: [64] Express Endpoint, MSI 00
Capabilities: [a0] MSI: Enable- Count=1/1 Maskable- 64bit+
Capabilities: [100] Vendor Specific Information: ID=0001 Rev=1 Len=010 <?>
Capabilities: [150] Advanced Error Reporting
Capabilities: [2a0] Access Control Services
Capabilities: [370] Transaction Processing Hints
Kernel driver in use: vfio-pci
With the default 6.8.3 Kernel I have the same issue like other users with an 3rd gen Ryzen. Passing through the onboard audio controller to a VM and starting it up causing the server to freeze/lockup. No matter if I bind the device to vfio or not. Only way to get the server back to a working state is to force a shutdown.
Short snippet in the state when the server locks up. I can't pull the full diagnostic at this point, not via ssh nor via webui.
Apr 25 15:46:29 UNRAID kernel: vfio-pci 0000:23:00.4: not ready 1023ms after FLR; waiting
Apr 25 15:46:31 UNRAID kernel: vfio-pci 0000:23:00.4: not ready 2047ms after FLR; waiting
Apr 25 15:46:34 UNRAID kernel: vfio-pci 0000:23:00.4: not ready 4095ms after FLR; waiting
Apr 25 15:46:39 UNRAID kernel: vfio-pci 0000:23:00.4: not ready 8191ms after FLR; waiting
Apr 25 15:46:48 UNRAID kernel: vfio-pci 0000:23:00.4: not ready 16383ms after FLR; waiting
Apr 25 15:47:06 UNRAID kernel: vfio-pci 0000:23:00.4: not ready 32767ms after FLR; waiting
Apr 25 15:47:44 UNRAID kernel: vfio-pci 0000:23:00.4: not ready 65535ms after FLR; giving up
Apr 25 15:48:16 UNRAID nginx: 2020/04/25 15:48:16 [error] 11486#11486: *1680 upstream timed out (110: Connection timed out) while reading response header from upstream, client: 10.0.0.7, server: , request: "POST /plugins/dynamix.vm.manager/include/VMajax.php HTTP/2.0", upstream: "fastcgi://unix:/var/run/php5-fpm.sock", host: "unraid.local", referrer: "https://unraid.local/VMs"
Apr 25 15:48:43 UNRAID kernel: rcu: INFO: rcu_sched detected stalls on CPUs/tasks:
Apr 25 15:48:43 UNRAID kernel: rcu: 31-....: (59 ticks this GP) idle=a66/1/0x4000000000000000 softirq=131201/131201 fqs=14340
Apr 25 15:48:43 UNRAID kernel: rcu: (detected by 32, t=60002 jiffies, g=412949, q=631295)
Apr 25 15:48:43 UNRAID kernel: Sending NMI from CPU 32 to CPUs 31:
Apr 25 15:48:43 UNRAID kernel: NMI backtrace for cpu 31
Apr 25 15:48:43 UNRAID kernel: CPU: 31 PID: 47678 Comm: qemu-system-x86 Tainted: G O 4.19.107-Unraid #1
Apr 25 15:48:43 UNRAID kernel: Hardware name: Gigabyte Technology Co., Ltd. TRX40 AORUS XTREME/TRX40 AORUS XTREME, BIOS F4d 03/05/2020
Apr 25 15:48:43 UNRAID kernel: RIP: 0010:pci_mmcfg_read+0x98/0xa6
Apr 25 15:48:43 UNRAID kernel: Code: 83 fe 02 74 15 41 83 fe 04 74 1a 41 ff ce 75 1d 48 01 d8 8a 00 0f b6 c0 eb 10 48 01 d8 66 8b 00 0f b7 c0 eb 05 48 01 d8 8b 00 <89> 45 00 31 c0 5b 5d 41 5c 41 5d 41 5e c3 81 fa ff 00 00 00 41 56
Apr 25 15:48:43 UNRAID kernel: RSP: 0018:ffffc900122e3cc8 EFLAGS: 00000286
Apr 25 15:48:43 UNRAID kernel: RAX: 00000000ffffffff RBX: 0000000000000ffc RCX: 0000000000000ffc
Apr 25 15:48:43 UNRAID kernel: RDX: 000000000000007f RSI: 0000000000000023 RDI: ffffc90008000000
Apr 25 15:48:43 UNRAID kernel: RBP: ffffc900122e3cfc R08: 0000000000000004 R09: ffffc900122e3cfc
Apr 25 15:48:43 UNRAID kernel: R10: 0000000000000004 R11: 0000000000000084 R12: 0000000002304000
Apr 25 15:48:43 UNRAID kernel: R13: 0000000000000004 R14: 0000000000000004 R15: ffff888ebcb25b40
Apr 25 15:48:43 UNRAID kernel: FS: 00001524a07b4e00(0000) GS:ffff88902d5c0000(0000) knlGS:0000000000000000
Apr 25 15:48:43 UNRAID kernel: CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
Apr 25 15:48:43 UNRAID kernel: CR2: 000014f97fed62b0 CR3: 0000000f3cf9a000 CR4: 0000000000340ee0
Apr 25 15:48:43 UNRAID kernel: Call Trace:
Apr 25 15:48:43 UNRAID kernel: pci_bus_read_config_dword+0x44/0x65
Apr 25 15:48:43 UNRAID kernel: pci_find_next_ext_capability+0x9e/0xc9
Apr 25 15:48:43 UNRAID kernel: ? _raw_spin_unlock_irqrestore+0xc/0x12
Apr 25 15:48:43 UNRAID kernel: pci_restore_vc_state+0x20/0x5c
Apr 25 15:48:43 UNRAID kernel: pci_restore_state+0xd0/0x26e
Apr 25 15:48:43 UNRAID kernel: pci_dev_restore+0x18/0x34
Apr 25 15:48:43 UNRAID kernel: pci_try_reset_function+0x3f/0x4e
Apr 25 15:48:43 UNRAID kernel: vfio_pci_open+0x7e/0x3af
Apr 25 15:48:43 UNRAID kernel: vfio_group_fops_unl_ioctl+0x355/0x42e
Apr 25 15:48:43 UNRAID kernel: vfs_ioctl+0x19/0x26
Apr 25 15:48:43 UNRAID kernel: do_vfs_ioctl+0x533/0x55d
Apr 25 15:48:43 UNRAID kernel: ? __se_sys_newlstat+0x48/0x6b
Apr 25 15:48:43 UNRAID kernel: ksys_ioctl+0x37/0x56
Apr 25 15:48:43 UNRAID kernel: __x64_sys_ioctl+0x11/0x14
Apr 25 15:48:43 UNRAID kernel: do_syscall_64+0x57/0xf2
Apr 25 15:48:43 UNRAID kernel: entry_SYSCALL_64_after_hwframe+0x44/0xa9
Apr 25 15:48:43 UNRAID kernel: RIP: 0033:0x1524a1fa24b7
Apr 25 15:48:43 UNRAID kernel: Code: 00 00 90 48 8b 05 d9 29 0d 00 64 c7 00 26 00 00 00 48 c7 c0 ff ff ff ff c3 66 2e 0f 1f 84 00 00 00 00 00 b8 10 00 00 00 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 8b 0d a9 29 0d 00 f7 d8 64 89 01 48
Apr 25 15:48:43 UNRAID kernel: RSP: 002b:00007fff5878e818 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
Apr 25 15:48:43 UNRAID kernel: RAX: ffffffffffffffda RBX: 000015241dfc14e0 RCX: 00001524a1fa24b7
Apr 25 15:48:43 UNRAID kernel: RDX: 000015241c9a1520 RSI: 0000000000003b6a RDI: 000000000000001b
Apr 25 15:48:43 UNRAID kernel: RBP: 00001524181605c0 R08: 000015241c9a1520 R09: 00007fff5878c783
Apr 25 15:48:43 UNRAID kernel: R10: 0000000000000018 R11: 0000000000000246 R12: 00001524181605c0
Apr 25 15:48:43 UNRAID kernel: R13: 000015241c9a1520 R14: 00007fff5878fa30 R15: 000015241dfc0c00
At this point, the server becomes basically unusable.
If I use your custom Kernel "6.8.3-5.5.8-2" you posted 2 pages back and the kernel agument "pcie_no_flr"
append pcie_no_flr=1022:1487 vfio-pci.ids=1022:1487,1b21:2142 isolcpus=12-23,36-47 initrd=/bzroot
I'am able to passthrough the device to a Windows VM without freezing the whole server, BUT the device isn't shown in the device manager as new device. 1b21:2142 is one of the USB onboard controllers and passthrough is working fine. Drivers are installed in a bare metal Windows install on an NVME which I use in this VM. I can't trigger Windows to show it in the device list. No errors in the VM log except for some warning:
2020-04-25T15:24:26.974831Z qemu-system-x86_64: vfio: Cannot reset device 0000:23:00.4, depends on group 39 which is not owned.
2020-04-25T15:24:29.155826Z qemu-system-x86_64: vfio: Cannot reset device 0000:23:00.4, depends on group 39 which is not owned.
IOMMU group 39: [1022:1485] 23:00.0 Non-Essential Instrumentation [1300]: Advanced Micro Devices, Inc. [AMD] Starship/Matisse Reserved SPP
The "Starship/Matisse Reserved SPP" isn't bind to vfio or used for passthrough either. Do I have to bind and passthrough that device as well???
Any idea for an workaround to have the onboard audio device to show up in the Windows VM? I really wanna figure this out and maybe try to help @limetech to implement the fix for the onboard audio issue in the next builds hopefully. I'am willing to test to find a solution for this. From all I've read so far, a couple people have issues passing these devices to a VM on the latest AMD platforms. Maybe we can workout a solution thats worth to implement in the future Unraid builds. I've not ried the 6.9 beta build yet. It's kinda the next step for me to try, but I think this won't make any difference.
I'll report back soon. Thanks for the help ❤️
syslog_683_default_kernel.txt syslog_patched_5.5.8-2_kernel.txt Win10_VM_log_patched.txt 01_W10.xml IOMMU.txt
EDIT:
Tried the 6.9 beta 1, same issue as with unpatched 6.8.3 kernel. Same freezes, same errors. No passthrough of the onboard audio possible. After reverting back to 6.8.3 with the patched 5.5.8-2 Kernel I tried a couple combinations of binding and passing through the device in group 39 with no success. Either the "FLR; waiting" error occures and server freezes or VM starts fine without any extra audio device showing up in the device manager.