Sat Oct 10, 2026 6:02 pm
This actually occurs during a Cluster Aware Updating Run of the three nodes. I use the CAU.ps1 as a pre-update script to verify that the nodes are all in synchronous state before updating each node. Usually, node 2 (Starwind priority 1 node) will update first and all nodes are in synch when Cluster Aware Update starts so that node begins updating quickly. All VMs are drained from node 2 to either node 1 or node 3 and the update installs and afterward the computer will reboot. Once the node 2 computer reboots and there are no more updates on node 2, the VMs fail back to node 2, then generally node1, which is the priority 0 node in Starwind, will be the next to start to update. Most of the time, the CAU.ps1 script on node 1 will cause cluster aware updating to wait for only a few minutes for node 2 to get synchronized before node 1 begins the update process. Once all the nodes are in synchronization, the CAU.ps1 script finishes and then node 1 will drain its roles to node 2 and node 3 and begin updating. After node 1 updates, the computer reboots and the VMs fail back to node 1 and then node 3 (priority 2) will begin its update process. However, after node 1 reboots, the CAU.ps1 script will not let node 3 proceed because the starwind nodes are not in synchronization. That is when I check and see that all the disks on node 1 are having to perform a full resynchronization. As a result, the cluster aware update script has to wait for about 24 hours for node 1 to get back into synchronization before the updates can be applied to node 3. This is the problem that I am trying to resolve. This was never an observed issue on a 2-node cluster, it has only occurred since moving to a three-node cluster. This also may be the result of windows updates that cause multiple reboots during cumulative updates and servicing stack updates. Potentially, the fast synchronization starts on the first reboot and then if there is a second reboot before the fast resync finishes, it triggers a full resynchronization.
When you say "avoid restart mishandles", is there anything else that I can do to prevent that? I am verifying that the machines are synchronized using the CAU script before the update and reboot by Windows. When you say remove the write-back cache, do you mean to remove caching completely or change it from write-back to write-through?