Every concept has a breaking point.
The important, badly-explained parts of computer science - presented as systems you can push until they fail. Most engineers learn where these break from an outage. This is the cheaper way to find out.
Reading needs no account. Remembering is the hard part.
Every concept, every widget and every failure verb here is free and always will be - that is not a trial. What an account adds is the bookkeeping you would otherwise do by hand: call the outcome before the simulation runs, and each concept you get wrong comes back a few days later, then again just before you would have forgotten it. Your progress follows you to whatever machine you next sit down at.
Continue with GoogleFree. Nothing to cancel.
Your queue is where the ones you got wrong come back.
Each concept is scheduled for the day your recall is about to slip rather than a fixed interval. If nothing is due, the queue starts you on concepts you have not called yet.
Agents
7- Context compaction in agent loops
An agent that runs out of context does not crash - it summarises, forgets the one detail the task depended on, and continues with complete confidence.
- Error compounding
An agent that gets 95% of its steps right finishes a twenty-step task about a third of the time, and the number that fixes it is not the one everybody measures.
- Lost in the middle
Retrieval worked, the answer is sitting in the context, and the model still gets it wrong - because a fact in the middle of a long context is read far less reliably than the same fact at either end.
- Prompt caching economics
A cached prefix costs a tenth of an uncached one to read and a quarter more to write - so one badly placed token turns an 88% saving into paying 25% over the list price.
- Subagents and the coordination tax
Splitting work across subagents buys exactly one thing - a smaller context per agent - and you pay for it twice, in duplicated tokens and in the answers that needed something sitting in another agent's window.
- The agent loop and where it burns your budget
An agent that is stuck looks exactly like an agent that is working. It keeps calling tools, keeps producing turns, and spends the entire context window finding out it was wrong on the first one.
- Tool selection at scale
An agent does not fail to pick the right tool because it has too many tools. It fails because three of them look alike, and three lookalikes cost exactly what sixty unrelated tools would.
AI / LLM Internals
4- Continuous batching
A static batch of 16 leaves about a third of the GPU generating padding, and doubling the batch to fix it makes that share worse.
- KV cache and speculative decoding
Generating text is memory-bound, not compute-bound - the GPU spends most of its time waiting on weights, which is why you can get several tokens for the price of one.
- Reading an eval result
A 3-point win on a 200-item eval ships the worse system about one time in five, and the report will not say so.
- Tokenization and the price of a script
The same sentence costs several times more in one language than another, and roughly a third of that gap was decided by UTF-8 in 2003, before anyone trained a tokenizer.
Data Structures & Algorithms
4- Bloom filters
A structure that answers "have I seen this before?" in a few bits per item, by being allowed to occasionally lie in one direction only.
- Cuckoo filters
A membership filter you can actually delete from, bought by storing fingerprints instead of bits, and paid for with a table that fails loudly instead of degrading quietly.
- HyperLogLog
Counting distinct items normally costs memory proportional to how many there are; HyperLogLog counts billions in a couple of kilobytes by never storing a single one of them.
- Skip lists
A balanced search structure with no balancing code - it gets its shape from coin flips, and the coin flips are why it is easy to make concurrent.
Storage & Databases
3- B-trees and the cost of a split
Everyone can draw a B-tree. Almost nobody can say what one insert actually costs, or why an index built from sorted keys wastes a third more space than the same index built from shuffled ones.
- Consistent hashing with bounded loads
Consistent hashing solves the resharding problem and quietly creates a load-balancing one, because a ring carved at random is never carved evenly.
- LSM trees and compaction
Every write is fast because it only ever appends - and then the database spends the rest of its life rewriting that data over and over in the background, which is where all the real costs live.
Distributed Systems
3- Quorums and R+W>N
One inequality decides whether your replicated store can return stale data, and it has nothing to do with how fast replication is.
- Raft leader election
A partitioned leader has no way to discover it has been deposed, and will go on believing it leads a cluster that has already moved on without it.
- Vector clocks and causal order
Timestamps cannot tell you whether two events conflicted or merely happened at different times, and picking the larger one silently deletes half of every genuine conflict.
Performance & Systems
4- Backpressure and bufferbloat
A queue is not a solution to being overloaded; it is a place to store the evidence that you are, and a bigger one only means you find out later and hurt more.
- Hedged requests
Sending 5% of your requests twice cuts p99 by a quarter. The same setting, on a busy day, can make it forty times worse.
- Metastable failure
A server at 80% utilisation, briefly disturbed and then fully restored, can serve nobody at all for hours - because the retries meant to hide the disturbance are now the only thing holding it down.
- Tail latency amplification
A service where 99% of calls are fast becomes a service where most calls are slow, the moment one user request has to touch a hundred machines.
Messaging & Wire Formats
1- Protobuf wire format
Field names never appear on the wire, which is why renaming a field is free and renumbering one silently corrupts every message you have ever stored.
Famous Papers
1- Dynamo, and what it gives up
R+W>N is a proof, not a guideline. Dynamo keeps writing through a partition by making the set it was a proof about stop existing.
Security & Correctness
1- Rate limiting and the boundary burst
A fixed-window limiter set to 60 per minute will let 120 through in an instant, and the admitted-percentage graph you would use to check it reads exactly the same as one that works.