Protobuf wire format
Field names never appear on the wire, which is why renaming a field is free and renumbering one silently corrupts every message you have ever stored.
Most engineers use protobuf for years without ever looking at a byte of it. That works right up until a schema change goes out and a service starts reading garbage, at which point the format stops being an implementation detail very abruptly.
It is worth twenty minutes. The whole encoding is four wire types and one clever integer representation.[1]specProtocol Buffers - Encoding
Tags: field number and type in one varint
Every field starts with a tag, which is a single varint holding two things:
tag = (field_number << 3) | wire_type
Three bits for the wire type, everything above for the field number. So field 1 with wire
type 0 encodes as 0x08, and a message {id: 150} is exactly three bytes: 08 96 01.
That layout is also why field numbers 1–15 are worth conserving. Five bits of field number plus three of wire type fit in one byte; field 16 needs two. Protobuf’s own advice is to spend the single-byte numbers on your hot fields.[1]specProtocol Buffers - Encoding
Varints: seven bits at a time
A varint stores seven bits of payload per byte, using the top bit as a continuation flag.[3]specDWARF Debugging Information Format - LEB128 encoding Small numbers are small: 1 is one byte, 300 is two.
The trap is negative numbers. A negative int32 is sign-extended to 64 bits before
encoding, so every negative int32 costs the full ten bytes. -1 encodes as
ff ff ff ff ff ff ff ff ff 01.
sint32 fixes this with zigzag, interleaving positives and negatives so that magnitude
rather than sign decides the size: 0 → 0, -1 → 1, 1 → 2, -2 → 3. Set score to a
negative value in the widget and toggle sint32 to watch ten bytes collapse into one.
Why old readers survive new writers
A reader that meets field 3 when its schema only knows fields 1 and 2 is not stuck. The wire type in the tag tells it how the value is framed - a varint it can scan to the terminator, a length-delimited value whose length comes next - so it knows how many bytes to skip without knowing what they mean.[1]specProtocol Buffers - Encoding
Press Writer adds a field in the widget. The reader shows the unknown field in amber, skips it correctly, and parses everything else.
The rules, and what breaks when you ignore them
Renaming is free. Names exist only in generated code; the wire carries numbers. Rename
id to user_id and not a single byte changes.
Renumbering is a breaking change even if the name is identical. Try Renumber a
field: the reader is looking for name at field 9, the writer is sending it at field 2,
and the result is a message where name is simply absent while a mysterious unknown field
sits in the payload.
Reusing a retired number is the dangerous one. Press Reuse a retired number. Field
2 used to be a string; the new schema says it is an int32. Nothing errors. The reader has
no way to know the number was ever used for something else. This is why protobuf gives you
reserved - and why using it is not optional bookkeeping.[2]docsProtocol Buffers - Updating a Message Type
The proto3 default-value trap
proto3 does not write fields that hold their default. A zero int32, an empty string and a
false bool all encode as nothing at all.
Set id to 0 and name to empty in the widget: the message becomes zero bytes.
The consequence is that a reader cannot distinguish “explicitly set to zero” from “never
set”. For a counter that is fine. For a settings update where 0 means “disable this”, it
is a bug that will reach production, because it only shows up when someone deliberately
sets the field to its default. proto3 later reintroduced optional to get presence
tracking back for exactly this reason.[2]docsProtocol Buffers - Updating a Message Type
The reference implementation is worth reading if you want the details of how skipping and framing are done in practice.[4]source codeprotobuf - wire_format_lite.h
The dial
The line - when someone asks
Protobuf encodes each field as a tag varint packing the field number and wire type, followed by the value, with length-delimited framing for anything variable-sized. That framing is what lets an old reader skip a field it does not recognise, which is the whole basis of schema evolution. Field numbers are the identity - names are compile-time only - so renaming is safe and renumbering or reusing a number is a data-corruption bug.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [2]
Secondary
- [3]
- [4]