September 2Sep 2 Long time lurker now reaching out for help.A few weeks ago I had one of my drives (drive 9) throw an error that took it offline. Unraid put it in emulated state. After running 3 preclear cycles on the drive, it passed. I then rebuilt the array back onto the flagged for failure drive.Since that time, I started digging into the logs curious for why I had the failure (something else caused the error?)Attached is system log of when I first had the failure (dated 8/25). I do have a full diagnostic on 8/25 but it is not "anonimized" Here (in the earlier system log) you will see that Disk 9 failed and disabled Disk 10 had interface problems (ATA bus errors - maybe bad cable?) I have had various CRC and Raw Read errors on numerous drives - although please note some CRC errors had occurred long in the past.Since then I have Run xfs file checks on all drives while in maintenance mode Run full extended smart checks on all disks (they all pass) Reseated SATA cables Replaced the Cables connected to the ASM1061 (Disk 5 and Disk 10) Replaced the Power Supply ( I had a 3-Rail 12Volt Power Supply, replaced with Single Rail - EVGA 550G3) Placed a fan on the LSI 9207-8i (I discovered it was running hot at 78 degrees C) [Disks 6,7,8,9,11,12,13,14 connected here]***TODAY***If I run a correcting parity sync - Unraid still informs me of a few dozen sync errorsIf I then run another non-correcting parity sync immediately after, Unraid is still complaining of sync errorsAt this point I am now trying to determine if drives may be failing. I have multiple drives with Raw Read errors, but NONE have uncorrectable, reallocated, or pending sectors. Hard drives are kinda expensive right now........ but Drive 8 and 9 seem to be increasing Raw Read errors (but nothing else)I started using Unraid Mover to clear Drive 8 to possibly remove from the array, maybe replace. During this operation, the log AGAIN complains of another "interface fatal error" I believe this was a link error on Drive 10 again if I am reading the log correctly.Main Question - what is causing the ongoing sync errors? Drive 8 or/and Drive 9 Raw Read ErrorsDrive 10 IO Error (cable has been replaced)Something Else?Would someone please help me look into my woes with fresh (and more experience) eyes. I seem to be running in circles and may possibly be having simultaneous problems. I have also attached a current diagnostics file.I appreciate your time.... syslog - 2026-08-25.txt towerbackup-diagnostics-20260901-2158.zip
September 2Sep 2 Community Expert The diagnostics suggest a connection-path problem affecting disk10 rather than clear evidence of multiple failing disks. Disk10 is connected to the second ASM1061 port and repeatedly logs SATA handshake/ATA bus errors, followed by link resets and a reduction from 6 to 3 Gb/s. The same error sequence is present in both supplied logs.Since you have already replaced the SATA cable and PSU, I would next move disk10 to a known-good Intel or LSI SATA connection, if possible. If no spare connection is available, power down and swap the ASM1061 connections for disk5 and disk10; that can help determine whether the errors remain with the controller port or follow disk10. I would avoid further correcting parity checks until this connection issue is resolved, and ensure any important data is backed up. The recent extended SMART tests for disks 8, 9, and 10 all passed, with no pending, reallocated, or offline-uncorrectable sectors, but the displayed WD raw-read values are not a good sign, especially if they keep going up. The older syslog also does not include the event that originally disabled disk9, so unfortunately we cannot determine its cause from that file.After moving disk10, exercise the array normally and post fresh diagnostics. If no further ATA errors appear, run one non-correcting parity check during otherwise quiet activity. Some errors during the 1st check can be normal even if the issue is resolved.If parity errors still recur after the 1st check with a clean storage log, the next step would be an extended Memtest86+ run to rule out RAM/platform instability.
September 5Sep 5 Author Well... I switched Disk 5 and Disk 10 on the ASM1061. Specifically switched the cables at the ports on the card.At the same time, I had also moved the card from a 16x PCIe slot to a dedicated 1x PCIe slot just in case that was an issue.I also had run a Memtest86+ last week (8 cycles) - no apparent memory problemsThe ATA bus errors / link resets followed from drive 10 to drive 5.The ASM1061 seems to possibly going bad? Or at least one port? Have you guys encountered one of these cards going bad?I also seem to have another drive (Disk 13) that is now showing pending and un-correctable sectors. Probably time to replace this one.What is the current recommended 2 or 4 port SATA expansion card? ASM1061/2 still seems to be recommended a bunch.Diagnostics Attached.Cheers,jsrtowerbackup-diagnostics-20260905-0000.zip Edited September 5Sep 5 by jsr
September 6Sep 6 Community Expert 18 hours ago, jsr said:The ATA bus errors / link resets followed from drive 10 to drive 5.The ASM1061 seems to possibly going bad?It's possible that the port is going bad.
11 hours ago11 hr Author OK. Follow Up.I have removed the ASM1061 - assuming it was a source of errors and replaced with a LSI 9211-8i (yes, maybe overkill, but gives future options) purchased from the venerable ArtOfServer.After erasing and performing multiple preclear passes on Drive 8 and Drive 13 - both passed (although concernedly they do have smart raw read errors)Now I have within the Log Files the following weird EXPERIMENTAL XFS drive flags:Oct 2 15:51:57 TowerBackup emhttpd: mounting /mnt/disk8 Oct 2 15:51:57 TowerBackup emhttpd: shcmd (659): mkdir -m 0666 -p /mnt/disk8 Oct 2 15:51:57 TowerBackup emhttpd: shcmd (660): mount -t xfs -o noatime,nodiscard,nouuid /dev/md8p1 /mnt/disk8 Oct 2 15:51:57 TowerBackup kernel: XFS (md8p1): EXPERIMENTAL exchange-range feature enabled. Use at your own risk! Oct 2 15:51:57 TowerBackup kernel: XFS (md8p1): EXPERIMENTAL parent pointer feature enabled. Use at your own risk! Oct 2 15:51:57 TowerBackup kernel: XFS (md8p1): Mounting V5 Filesystem e2890e67-66ef-4d56-b285-3c60b8732b99 Oct 2 15:51:57 TowerBackup kernel: XFS (md8p1): Ending clean mountDoes anyone know why I am getting these? I have already: used the "new configuration" tool within unraid, kept drive assignments, selected the Drives 8 and 13 clicked "erase" under each disk's setting formatted after bringing array onlineI have searched the forums and can't seem to find any firm clues. One possible clue may be that in the log file, when the drive is mounted the following flag is different than other disks:exchange=1Oct 2 15:51:57 TowerBackup emhttpd: shcmd (661): xfs_growfs /mnt/disk8 Oct 2 15:51:57 TowerBackup root: meta-data=/dev/md8p1 isize=512 agcount=4, agsize=244188659 blks Oct 2 15:51:57 TowerBackup root: = sectsz=512 attr=2, projid32bit=1 Oct 2 15:51:57 TowerBackup root: = crc=1 finobt=1, sparse=1, rmapbt=1 Oct 2 15:51:57 TowerBackup root: = reflink=1 bigtime=1 inobtcount=1 nrext64=1 Oct 2 15:51:57 TowerBackup root: = exchange=1 metadir=0 Oct 2 15:51:57 TowerBackup root: data = bsize=4096 blocks=976754633, imaxpct=5 Oct 2 15:51:57 TowerBackup root: = sunit=0 swidth=0 blks Oct 2 15:51:57 TowerBackup root: naming =version 2 bsize=4096 ascii-ci=0, ftype=1, parent=1 Oct 2 15:51:57 TowerBackup root: log =internal log bsize=4096 blocks=476930, version=2 Oct 2 15:51:57 TowerBackup root: = sectsz=512 sunit=0 blks, lazy-count=1 Oct 2 15:51:57 TowerBackup root: realtime =none extsz=4096 blocks=0, rtextents=0 Oct 2 15:51:57 TowerBackup root: = rgcount=0 rgsize=0 extents Oct 2 15:51:57 TowerBackup root: = zoned=0 start=0 reserved=0 Oct 2 15:51:58 TowerBackup emhttpd: creating volume: disk13 (xfs)Attached is a full system diagnostics. towerbackup-diagnostics-20261002-1600.zip
2 hours ago2 hr Community Expert That suggests disk8 was formatted outside Unraid with some experimental feature; post the output from:xfs_info /mnt/disk8
Join the conversation
You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.