Hey all, I recently had one of my parity drives showing Parity device disabled.
To try to resolve this issue, I did these steps:
Ran a parity check with Write corrections to parity checked. This completed successfully, but the parity drive was still disabled.
Stopped the array, selected no device on the failed parity drive slot, restarted the array, stopped the array again, select the parity drive, restarted the array.
After doing that, both parity drives were down and disk 7 showed Device contents emulated as well as Unsupported or no file system.
Tried to debug this a bit with a helpful user on the forum, but unfortunately wasn't able to fix the issue.
The additional steps taken during debugging were:
Tools > New Config > Select to preserve all arrays and pools > Apply
Started the array and formatted disk 7
Clicked the red X on the two parity drives under unassigned devices to clear the drives, then formatted both drives.
Stopped the array, assigned the two parity drives, and restarted the array
I also reorganized drives and cables, to confirm that disk 7 is consistently failing and that it is not an issue with the cables or HBA. I connected a SATA cable from disk 7 directly to the motherboard instead of connecting it to the HBA.
Starting the array now seems very slow:
Oct 12 00:35:50 raphnas root: realtime =none extsz=4096 blocks=0, rtextents=0
Oct 12 00:35:50 raphnas emhttpd: mounting /mnt/disk9
Oct 12 00:35:50 raphnas emhttpd: shcmd (127): mkdir -p /mnt/disk9
Oct 12 00:35:50 raphnas emhttpd: shcmd (128): mount -t xfs -o noatime,nouuid /dev/md9p1 /mnt/disk9
Oct 12 00:35:50 raphnas kernel: XFS (md9p1): Mounting V5 Filesystem
Oct 12 00:35:51 raphnas kernel: ata6.00: exception Emask 0x50 SAct 0x200000 SErr 0x4890800 action 0xe frozen
Oct 12 00:35:51 raphnas kernel: ata6.00: irq_stat 0x0c400040, interface fatal error, connection status changed
Oct 12 00:35:51 raphnas kernel: ata6: SError: { HostInt PHYRdyChg 10B8B LinkSeq DevExch }
Oct 12 00:35:51 raphnas kernel: ata6.00: failed command: READ FPDMA QUEUED
Oct 12 00:35:51 raphnas kernel: ata6.00: cmd 60/00:a8:38:0d:01/01:00:80:04:00/40 tag 21 ncq dma 131072 in
Oct 12 00:35:51 raphnas kernel: res 40/00:00:38:0d:01/00:00:80:04:00/40 Emask 0x50 (ATA bus error)
Oct 12 00:35:51 raphnas kernel: ata6.00: status: { DRDY }
Oct 12 00:35:51 raphnas kernel: ata6: hard resetting link
Oct 12 00:35:52 raphnas kernel: ata6: SATA link down (SStatus 0 SControl 380)
Oct 12 00:35:58 raphnas kernel: ata6: hard resetting link
Oct 12 00:35:58 raphnas kernel: ata6: SATA link down (SStatus 0 SControl 380)
Oct 12 00:35:58 raphnas kernel: ata6: limiting SATA link speed to <unknown>
Oct 12 00:35:59 raphnas kernel: ata6: hard resetting link
Oct 12 00:36:04 raphnas kernel: ata6: link is slow to respond, please be patient (ready=0)
Oct 12 00:36:09 raphnas kernel: ata6: COMRESET failed (errno=-16)
Oct 12 00:36:09 raphnas kernel: ata6: hard resetting link
As it currently stands, when the array finally starts up, there are errors on disks 4 and 7.
A few more details about my setup in case it helps:
12x 20TB WD reds, 10x storage, 2x parity (all brand new drives, bought two months ago)
1x 1TB NVME
2x LSI SAS 9211-8i HBA (8x drives on one HBA, 4x drives on the other HBA)
Asus Prime H570-Plus mobo
Intel Celeron G5925 CPU
32GB DDR4 RAM
Corsair CX550 PSU
Attaching diagnostics for further debugging.
Please let me know if there is any further information I can provide to help identify the issue.
raphnas-diagnostics-20231012-0057.zip