Question 1
A platform engineer is troubleshooting a PowerStore 5000T appliance where node B is in service mode. The engineer observes a recurring alert: (0x7770101) Node B has rebooted unexpectedly. PowerStore Manager shows node A is healthy and serving all I/O. Analysis of the support bundle indicates a potential memory DIMM issue on node B. What is the most appropriate next step to resolve this issue while minimizing disruption?
Answer and explanation
Correct answer: C
The correct approach is to isolate the faulty node and perform a targeted CRU (Customer Replaceable Unit) replacement. Since node A is healthy and serving all I/O, the cluster is not down. The documented procedure for replacing a component like a DIMM in a single node involves safely powering down only that node while the peer node remains active. This minimizes disruption. A full cluster power-down is unnecessary and causes a complete outage. Forcing the node out of service mode is risky and does not address the underlying hardware fault. A full node replacement (FRU) is premature without first attempting the CRU replacement.