Raft Consensus, Explained
After reading this you will know how a Raft cluster picks one leader, replicates a log, and refuses to split its brain during a network partition, and you will be able to trace every one of those steps in the visualizer.
What Raft actually does
Raft keeps a group of servers agreeing on a single ordered log of commands, even when some servers crash or the network drops messages. It was published in 2014 by Diego Ongaro and John Ousterhout in "In Search of an Understandable Consensus Algorithm." The goal in that title is the whole point: Raft solves the same problem as Paxos but factors it into three parts you can hold in your head at once. Those parts are leader election, log replication, and safety.
Here is the hook. Start a 5 node cluster in the visualizer and do nothing. Within a few hundred milliseconds one follower's countdown ring hits zero, it becomes a candidate, it collects votes from 3 of the 5 nodes, and it starts sending heartbeats. You just watched an election finish without any central coordinator. Now click that leader to kill it. After another timeout, a survivor times out and a new leader appears. The cluster healed itself with 4 nodes.
Every command a client sends flows through the leader, gets appended to its log, replicates to followers, and commits once a majority stores it. That majority rule is the single idea that makes everything else safe.
When to reach for Raft, and when not
Raft fits when several machines must agree on the same sequence of state changes and you cannot tolerate two of them disagreeing. Cluster coordinators use it: etcd (which backs Kubernetes), Consul, and CockroachDB all run Raft. If you need a replicated configuration store, a distributed lock, or a metadata log where a stale answer would corrupt data, this is the right shape.
Do not reach for it when a single writer plus backups already meets your durability target, or when you can tolerate eventual consistency. Consensus costs you a network round trip to a majority on every write. In a 5 node cluster spread across three data centers, that write cannot commit faster than the round trip to the third-closest node. If your reads dominate and staleness is fine, a simpler replicated cache serves you better.
Raft assumes a crash-recovery model, not a Byzantine one. It survives nodes that stop, restart, or lose messages. It does not defend against a node that lies about its log. If you need to tolerate malicious participants, you need a Byzantine fault tolerant protocol instead, which is a different and heavier design.
The majority rule and the math of quorums
The one formula you need governs how many nodes must agree and how many failures the cluster survives.
Here N is the total number of nodes and the quorum is the smallest set that counts as a majority. A cluster commits an entry, and elects a leader, only with quorum agreement. The number of node failures it survives is N - \text{quorum}.
Work the three sizes the tool offers. For N = 3, quorum is \lfloor 1.5 \rfloor + 1 = 2, so 2 nodes commit and 1 failure is tolerated. For N = 5, quorum is 3 and 2 failures are tolerated. For N = 7, quorum is 4 and 3 failures are tolerated.
Two majorities of the same cluster always share at least one node, because 2 \times (\lfloor N/2 \rfloor + 1) \gt N. That overlap is why two leaders in the same term cannot both commit conflicting entries: any entry committed by one majority is visible to the next, since some node sat in both. Odd sizes are conventional because going from 5 to 6 nodes raises quorum from 3 to 4 while still tolerating only 2 failures. You paid for a node and bought nothing.
Terms, timeouts, and how an election resolves
Raft's logical clock is the term, a number that only increases. Each election starts a new term. Every message carries its sender's term, and the rule is absolute: any node that sees a term higher than its own immediately steps down to follower and adopts that term. This is how a stale leader learns it has been replaced the instant it hears from the new one.
Split votes are the obvious risk. If two followers time out at the same instant, both become candidates in the same term, both request votes, and each node votes for at most one candidate per term. Nobody reaches quorum and the term ends with no leader. Raft's fix is randomized election timeouts. The visualizer draws each timeout as a shrinking ring, and the durations are drawn randomly (the paper suggests a range like 150 to 300 ms). Because the ranges differ, one node almost always fires first and wins before its peers even wake up.
Kill the leader in the visualizer, then pause immediately. Single-step through the re-election. You will see one follower's ring reach zero, the term increment by one, RequestVote arrows fan out, and votes return before any second candidate appears. That staggering is the whole trick.
A worked election in the default 5 node cluster
From five followers to one committed entry
Load the tool with its defaults: a 5 node cluster, no partition, all nodes alive. Here is the sequence you can reproduce.
- All 5 nodes start as followers in term 0 with empty logs. Each has a random countdown ring.
- Node with the shortest ring times out first. It becomes a candidate, increments to term 1, votes for itself (1 vote), and sends
RequestVoteto the other 4. - Each of the other 4 has not voted in term 1, so each grants its vote. The candidate now holds 5 votes, well past the quorum of 3. It becomes leader.
- The leader sends empty
AppendEntriesheartbeats. Every follower resets its ring on each heartbeat, so no one else ever times out. - Press client request. The leader appends entry
x=1at log index 1, term 1, and replicates it. As each follower acknowledges, the leader's acknowledgement count rises: 1 (itself), then 2, then 3. - At 3 acknowledgements it has a majority. The leader advances its commit index to 1 and the commit marker moves. The next heartbeat tells followers to commit too.
The entry committed the moment 3 of 5 nodes stored it, not when all 5 did. If two followers were slow or dead, the entry would still commit on the strength of the other three.
Partitions and why the minority commits nothing
Flip the partition switch and split the 5 nodes into a group of 3 and a group of 2. Watch what each side can and cannot do.
The majority side of 3 still meets quorum. If its leader survived on that side, it keeps committing. If the leader landed on the minority side, the majority side times out, elects a new leader in a higher term, and carries on. The minority side of 2 can never reach quorum, because 2 is less than the required 3. Its nodes time out, become candidates, and fail to collect enough votes over and over, so their term numbers climb while they elect nobody.
Now the safety punchline. Suppose the old leader is stuck on the minority side. A client hits it and it appends an entry to its own log. That entry can never commit, because it can never gather 3 acknowledgements. When the partition heals, the old leader hears the higher term from the majority side, steps down, and its uncommitted entries get overwritten to match the majority log. No client ever saw that entry as committed, so no split-brain occurred.
A common misreading: seeing a minority-side node still labeled leader and assuming the cluster now has two leaders that both accept writes. It does not. That minority leader is from an older term and cannot commit anything. Check the commit markers, not the leader badges. Only entries past the commit marker are real.
Try it: move the partition line
Common mistakes when reasoning about Raft
Confusing "replicated" with "committed" causes the most bugs. An entry sitting in a follower's log is not committed until the leader's commit index passes it. In the visualizer, uncommitted cells sit to the right of the commit marker and can still vanish.
Assuming all nodes must acknowledge is the second trap. Commit needs a majority, not unanimity. A 7 node cluster commits with 4 acknowledgements while 3 nodes are still catching up or dead.
Ignoring terms is the third. Two nodes can both call themselves leader at a given wall-clock instant, but never in the same term, and only the higher-term leader can commit. Always read the term number alongside the leader badge.
Related tools
If you want to see other distributed and algorithmic mechanisms behave live, several tools here pair well with this one. The Consistent Hashing Ring shows how a Raft-backed cluster might place keys so that adding a node moves only 1/N of them. The Tail Latency & Autoscaling Simulator explains why the majority round trip that Raft needs gets expensive as utilization rises. For the data structures a state machine keeps, try the Data Structure Visualizer, and for eviction policies in the caches that often front such a store, the Cache Replacement Simulator.
Frequently asked questions
Why are Raft clusters almost always an odd number of nodes?
Because an even size buys no extra fault tolerance. A 4 node cluster needs quorum 3 and tolerates 1 failure, exactly like 3 nodes needing quorum 2. You added a machine and still survive only one crash, while raising the number of nodes every write must reach.
Can two leaders exist at the same time?
Yes at the same wall-clock moment, no in the same term. A partitioned old leader can still wear the leader badge until it hears a higher term, but it cannot commit anything because it cannot reach a majority. Only one leader per term can ever commit.
What happens to entries a stale leader accepted during a partition?
They stay uncommitted. When the partition heals and the stale leader learns of a higher term, it steps down and overwrites those entries to match the current leader's log. No client ever observed them as committed, so nothing durable is lost.
How does Raft avoid endless split votes?
Randomized election timeouts. Each follower waits a random duration in a range (the paper uses roughly 150 to 300 ms), so one node almost always times out first and wins before another candidate appears. If a split vote does happen, the term ends with no leader and new random timeouts try again.
Is this visualizer a complete Raft implementation?
No, and deliberately so. It omits log snapshotting and cluster membership changes, and it models each RPC as a single message with uniform latency rather than batched streams. The election and replication rules it does show are faithful to the paper.