Hyper-V 2 node cluster consistently restarts full sync

Software-based VM-centric and flash-friendly VM storage + free version
yaroslav (staff)
Staff
Posts: 4373
Joined: Mon Nov 18, 2019 11:11 am

Thu Dec 18, 2025 12:38 pm

Hi,

That should do it. Please keep me posted on how replication goes with 5%.
Electrum
Posts: 24
Joined: Tue Oct 08, 2024 2:22 pm

Thu Dec 18, 2025 7:28 pm

Hello,

Replication was set to 5%. I also let it attempt a full sync a few times. I'm still seeing the same issue.

Host 01 logs: https://mega.nz/file/iIM1AYzb#fmNW9e-Ct ... QCALHWymAw
Host 02 logs: https://mega.nz/file/OZFUECab#5IBNZz6Hw ... Ocd6oPXBqc
yaroslav (staff)
Staff
Posts: 4373
Joined: Mon Nov 18, 2019 11:11 am

Thu Dec 18, 2025 8:06 pm

There are storage hiccups. Here is what I see on host2.

For example
12/18 12:49:39.016942 21f4 IMG: ImageFile_IoCompleted: Warning(Time FileIO): request(0x0000023EDCEFD470) ssc(0x0000023EDE4B6E10) function(Execute SCSI Command) opCode(0x8A), timeFileIO = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019211 1918 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEF3B40) ssc(0x0000023EDE6F3A50) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019744 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE22DC30, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.019783 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE22D8C0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 8192, execution time (14547 msec) more than timeout (9000 msec)!
12/18 12:49:39.019941 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE785180, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 4096, execution time (9672 msec) more than timeout (9000 msec)!
12/18 12:49:39.019968 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE784730, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.019978 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE784AA0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.019989 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE784E10, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.019999 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE8D2EC0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.020009 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE8D2B50, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 8192, execution time (14547 msec) more than timeout (9000 msec)!
12/18 12:49:39.020024 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE8D3C80, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 8192, execution time (10359 msec) more than timeout (9000 msec)!
12/18 12:49:39.020035 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE8D3910, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 16384, execution time (13984 msec) more than timeout (9000 msec)!
12/18 12:49:39.020054 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE13D9D0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.020071 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE13D2F0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.020088 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE13C530, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.016940 1e4c IMG: ImageFile_IoCompleted: Warning(Time FileIO): request(0x0000023EDCEF5020) ssc(0x0000023EDE38C190) function(Execute SCSI Command) opCode(0x8A), timeFileIO = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019703 21f4 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEFD470) ssc(0x0000023EDE4B6E10) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.020260 21f4 Common: CStarWindStorageDevice::AsyncReadWriteCompleted: Underlying storage request(0x0000023EDF189E70, opcode 0x8A) execution time is 14765 ms.
12/18 12:49:39.019753 1918 Common: CStarWindStorageDevice::AsyncReadWriteCompleted: Underlying storage request(0x0000023EDF16B040, opcode 0x8A) execution time is 14765 ms.
12/18 12:49:39.019269 1ff8 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEF0FB0) ssc(0x0000023EDE86DA10) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019370 1ad0 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEF0E10) ssc(0x0000023EDE7681C0) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019398 23f8 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEF6130) ssc(0x0000023EDE76AD10) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'

and replication drop.
12/18 12:49:35.271667 13ec HA: HANode::setSyncStatus: Event: Device changed own sync status to 3 from 2. (target 'iqn.2008-08.com.starwindsoftware:home-hypv-h02-datastore02')

I can't help with anything here. You can try expanding the timeouts to compensate for the latencies, but I doubt it will help.
For each node:
1. Stop StarWind VSAN Service (beware of downtime).
2. Go to C:\Program Files\StarWind Software\StarWind\StarWind.cfg
3. Back it up.
4. Open and edit the following lines (i included values)
<StorPerfDegTimeLimitMs value="15000"/>
<iScsiGenCmdSendCmdTimeoutInSec value="18"/>
<iScsiPingCmdSendCmdTimeoutInSec value="14"/>
5. Save & start the service on that node.

If you have SSD (https://knowledgebase.starwindsoftware. ... dance/661/) or RAM (https://knowledgebase.starwindsoftware. ... -l1-cache/) cache on your devices, disable it.
Electrum
Posts: 24
Joined: Tue Oct 08, 2024 2:22 pm

Thu Dec 18, 2025 10:01 pm

yaroslav (staff) wrote:
Thu Dec 18, 2025 8:06 pm
There are storage hiccups. Here is what I see on host2.

For example
12/18 12:49:39.016942 21f4 IMG: ImageFile_IoCompleted: Warning(Time FileIO): request(0x0000023EDCEFD470) ssc(0x0000023EDE4B6E10) function(Execute SCSI Command) opCode(0x8A), timeFileIO = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019211 1918 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEF3B40) ssc(0x0000023EDE6F3A50) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019744 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE22DC30, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.019783 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE22D8C0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 8192, execution time (14547 msec) more than timeout (9000 msec)!
12/18 12:49:39.019941 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE785180, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 4096, execution time (9672 msec) more than timeout (9000 msec)!
12/18 12:49:39.019968 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE784730, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.019978 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE784AA0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.019989 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE784E10, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.019999 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE8D2EC0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.020009 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE8D2B50, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 8192, execution time (14547 msec) more than timeout (9000 msec)!
12/18 12:49:39.020024 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE8D3C80, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 8192, execution time (10359 msec) more than timeout (9000 msec)!
12/18 12:49:39.020035 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE8D3910, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 16384, execution time (13984 msec) more than timeout (9000 msec)!
12/18 12:49:39.020054 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE13D9D0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.020071 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE13D2F0, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.020088 1ecc Common: CStarWindStorageDevice::hangedTaskExist: Storage device hanged operation: location = E:\starwind\datastore02\datastore02.img, request = 0x0000023EDE13C530, cdb[0] = 0x8A, cdb[1] = 0x00, buf_size = 262144, execution time (14734 msec) more than timeout (9000 msec)!
12/18 12:49:39.016940 1e4c IMG: ImageFile_IoCompleted: Warning(Time FileIO): request(0x0000023EDCEF5020) ssc(0x0000023EDE38C190) function(Execute SCSI Command) opCode(0x8A), timeFileIO = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019703 21f4 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEFD470) ssc(0x0000023EDE4B6E10) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.020260 21f4 Common: CStarWindStorageDevice::AsyncReadWriteCompleted: Underlying storage request(0x0000023EDF189E70, opcode 0x8A) execution time is 14765 ms.
12/18 12:49:39.019753 1918 Common: CStarWindStorageDevice::AsyncReadWriteCompleted: Underlying storage request(0x0000023EDF16B040, opcode 0x8A) execution time is 14765 ms.
12/18 12:49:39.019269 1ff8 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEF0FB0) ssc(0x0000023EDE86DA10) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019370 1ad0 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEF0E10) ssc(0x0000023EDE7681C0) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'
12/18 12:49:39.019398 23f8 IMG: ImageFile_ScsiCompleteRequest: Warning(Time Request EXEC): request(0x0000023EDCEF6130) ssc(0x0000023EDE76AD10) function(Execute SCSI Command) opCode(0x8A), timeExecRequest = 14765 ms, g_cmdExecTimeWarningLimitInSec = 3 s. Device: 'D:\starwind\datastore01\datastore01.img'

and replication drop.
12/18 12:49:35.271667 13ec HA: HANode::setSyncStatus: Event: Device changed own sync status to 3 from 2. (target 'iqn.2008-08.com.starwindsoftware:home-hypv-h02-datastore02')

I can't help with anything here. You can try expanding the timeouts to compensate for the latencies, but I doubt it will help.
For each node:
1. Stop StarWind VSAN Service (beware of downtime).
2. Go to C:\Program Files\StarWind Software\StarWind\StarWind.cfg
3. Back it up.
4. Open and edit the following lines (i included values)
<StorPerfDegTimeLimitMs value="15000"/>
<iScsiGenCmdSendCmdTimeoutInSec value="18"/>
<iScsiPingCmdSendCmdTimeoutInSec value="14"/>
5. Save & start the service on that node.

If you have SSD (https://knowledgebase.starwindsoftware. ... dance/661/) or RAM (https://knowledgebase.starwindsoftware. ... -l1-cache/) cache on your devices, disable it.
I do see the configuration specified for cache, I'll disable it for all of them and try increasing the metrics you mentioned to see if it makes a difference.

The raid card on both of these is a PERC H730P Mini. Do you think the RAID card could be potentially what is causing these latency jumps? I've run several smartctl checks on the disks, they do come back clean. I was planning on replacing them regardless.
yaroslav (staff)
Staff
Posts: 4373
Joined: Mon Nov 18, 2019 11:11 am

Thu Dec 18, 2025 10:05 pm

If tou have SSD cache in place, it could be the culprit.
Electrum
Posts: 24
Joined: Tue Oct 08, 2024 2:22 pm

Thu Dec 18, 2025 10:13 pm

yaroslav (staff) wrote:
Thu Dec 18, 2025 10:05 pm
If tou have SSD cache in place, it could be the culprit.
Hello,

Apologies, I just checked and I don't see cache enabled on any of the disks (at least not in Starwind). I did increase the thresholds per your recommendation on both nodes. I also set the sync priority to 5% on both nodes, just to validate.
yaroslav (staff)
Staff
Posts: 4373
Joined: Mon Nov 18, 2019 11:11 am

Thu Dec 18, 2025 10:15 pm

Let's hope it helps.
Electrum
Posts: 24
Joined: Tue Oct 08, 2024 2:22 pm

Tue Dec 23, 2025 10:49 am

Wanted to provide an update in case anyone else runs into this issue. I noticed one of the HA starwind files on the second node didn't quite match (sector size was different). I checked they all matched, and everything else matched perfectly. I also adjusted the sync priority to 5%, although I suspect I can increase it now. Regardless, this resolved my issue. I have since reimaged and upgraded the nodes from 2022 to 2025, and everything continues to work as expected. Thanks again for your help!
yaroslav (staff)
Staff
Posts: 4373
Joined: Mon Nov 18, 2019 11:11 am

Tue Dec 23, 2025 12:31 pm

I am glad to read that your quest is over.
Merry Christmas and Happy New Year.
Post Reply