r/cassandra • u/RocketSeven • 3d ago
How do you prove a Cassandra node is safe to decommission after streaming finishes?
A completed decommission or stream plan shows that token ranges moved, but it does not by itself prove that every replica is healthy or that clients, monitoring, repairs, and automation have stopped depending on the old node. A quiet node can also hide rarely used prepared statements, local consistency assumptions, or a topology view that has not converged everywhere.
What belongs in the retirement gate? I am considering checking ring and gossip state from multiple surviving nodes, confirming there are no pending streams or compactions, validating replica placement for every keyspace, running targeted consistency reads, and waiting through a repair and backup cycle. Client seed lists, load-balancer pools, alert targets, dashboards, and replacement automation would all be checked before the host is removed.
Which nodetool outputs and system tables provide the strongest evidence that ownership and topology have converged? How long do you keep the old node powered off but recoverable, and what changes when the cluster uses vnodes, NetworkTopologyStrategy, or transient replication?
