Raft leader election
A partitioned leader has no way to discover it has been deposed, and will go on believing it leads a cluster that has already moved on without it.
Consensus algorithms are usually explained through their happy path, which is the least interesting thing about them. The happy path is: one node is in charge, everyone else does what it says. Everything worth understanding happens when that arrangement breaks.
Raft was explicitly designed to be understandable, and it mostly succeeds.[1]paperIn Search of an Understandable Consensus Algorithm (Extended Version) But “understandable” is not the same as “obvious”, and there are a couple of behaviours below that surprise people who think they already know it.
Terms, timeouts, votes
Time is divided into terms, each with at most one leader. Every node holds an election timer. When it expires without hearing from a leader, the node becomes a candidate, increments its term, votes for itself, and asks everyone else for a vote.
A node grants its vote if the candidate’s term is at least its own and it has not already voted in that term. A candidate that collects votes from a majority becomes leader and starts sending heartbeats, which reset everyone else’s timers.
One rule does most of the safety work: any message carrying a higher term causes the receiver to step down and adopt that term. That is what stops a node stranded in an old term from doing damage when it reconnects.
Why two leaders in one term is impossible
Winning requires a majority of the cluster. Two candidates in the same term would both need a majority, and two majorities of the same set must overlap in at least one node. That node only gets one vote per term. So it cannot have voted for both.
This is Raft’s Election Safety property.[1, §5.2]paperIn Search of an Understandable Consensus Algorithm (Extended Version)§5.2
The part that surprises people
A leader that gets cut off from the majority does not step down.
There is no mechanism for it to do so. A leader only reverts to follower on hearing a higher term, and by definition that message cannot cross the partition. So the deposed leader sits in its corner of the network, still believing it is in charge, while the majority elects a successor and moves on.
Try it: let a leader establish itself, then hit Partition. Watch the minority side.
What saves you is not that the stale leader realises anything - it is that it cannot commit. Committing requires acknowledgements from a majority it can no longer reach. So it accepts writes it will never be able to complete, and when the partition heals it discovers a higher term and discards them.
Randomised timeouts are not a detail
If every node used the same election timeout, a leader failure would cause all of them to time out at the same instant, all become candidates in the same term, and all vote for themselves. Nobody gets a majority. They time out again, together, and repeat.
Raft breaks the symmetry by choosing each timeout randomly from a range.[1, §5.2]paperIn Search of an Understandable Consensus Algorithm (Extended Version)§5.2 Usually one node fires first and wins before the others wake up.
Set Timeout randomisation to 0 in the widget. The cluster churns through terms and spends much of its time with no leader at all - availability drops while safety holds perfectly, which is the tradeoff consensus always makes.
What it costs
Every election is a window with no leader, so no writes. Raft’s availability is therefore bounded by how fast it detects failure and elects a replacement - and detection cannot be faster than the election timeout without producing false positives that trigger unnecessary elections.
That gives you the tuning problem the thesis spends real time on: timeouts short enough to recover quickly, long enough that ordinary network hiccups do not depose a healthy leader.[2]paperConsensus: Bridging Theory and Practice Turn the Delay verb up past the election timeout and you can watch a perfectly healthy cluster tear itself apart re-electing.
The dial
The line - when someone asks
Raft elects one leader per term using randomised election timeouts: a node that hears nothing times out, increments the term, and asks everyone for a vote. A candidate needs a majority to win, which is why a minority partition can never elect anyone and why split brain cannot produce two leaders in the same term. The randomisation is the part people forget - without it, nodes campaign simultaneously, split the vote, and the cluster can churn through terms without ever electing anybody.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]
Secondary
- [3]
- [4]