“Unique does not mean easy.
Ambiguous does not mean secret.”

What you really need to know
The previous artefacts in this series have included machines, working papers, and a moth. This one is sixty pages of mathematics (the artefact, not this article!) Nothing rotates. Nothing glows. Nobody has helpfully taped a theorem into a maintenance log. There is NO written examination for you to sit at the end, however tempted I was to add one. This section comprises a minimal TL;DR. But read the italics at the end of this section before abandoning ship!
Encryption is not meant to make a message look complicated. It is meant to stop an outsider learning its contents. The intended recipient can recover the message because they possess something the outsider does not: the appropriate secret key.
Claude Shannon made that promise precise. In his 1949 paper, perfect secrecy means that seeing the encrypted message does not change the outsider’s probabilities for what the original message might be. Not even unlimited computing power can extract evidence that the ciphertext does not contain. This does not erase knowledge obtained elsewhere.[1]
Perfect secrecy is possible, but its conditions are demanding. The standard one-time pad uses fresh, independent, secret randomness for every part of the message. For arbitrary binary messages, the pad must be as long as the data it protects, and none of it may be reused.[4]
Most everyday encryption makes a different promise. Rather than providing no information to an unlimited attacker, it aims to make useful extraction computationally infeasible for an attacker with realistic resources. That is not a failure to understand Shannon. It is a deliberate engineering trade-off.[4][10]
Neither promise means that the whole system is safe. Secret contents do not automatically mean an authentic sender, an unaltered message, hidden traffic patterns, or a secure device. Encryption cannot protect a message from someone who can read it before encryption or after decryption.
The question to carry away is simple: what can the observer learn, with what evidence, under which assumptions?
I encourage everyone to give the article a go. I’ve tried to make it as accessible as possible. The main argument runs through the prose and a simple worked example using the four points of the compass. The three optional orange SUPER-TECH call-outs provide more formal detail and some less familiar traps, but can be skipped without losing the main argument. If you can follow at least some of the article, then you will have grasped something genuinely fundamental.
Communications Secrecy
The subject is intensely practical. Two people want to communicate. Somebody else can listen. What would it mean for the listener to learn nothing useful from the intercepted message? And how could we distinguish that situation from one in which the listener simply has not worked out the answer yet?
Images in this article are generated by

EU AI Act Regulation 2024/1689
Claude Elwood Shannon’s Communication Theory of Secrecy Systems appeared in the Bell System Technical Journal in October 1949, volume 28, pages 656–715. The opening footnote traces its material to a confidential report, A Mathematical Theory of Cryptography, dated 1 September 1945. The public paper therefore emerged from wartime work.[1][3]
It is not the same paper as Shannon’s A Mathematical Theory of Communication, published in two instalments in 1948. The titles are easy to muddle; the questions are different. The 1948 work established a general mathematical treatment of communication, including the effects of noise and the statistical structure of messages. The 1949 paper brought that same way of thinking to the subject of secrecy.[1][2]
Shannon was working at Bell Laboratories, an institution concerned with making signals intelligible across real channels. Wartime projects also required some signals to remain unintelligible to the wrong people. That is a useful intellectual position: the engineer must understand both how information arrives and how it might escape.[3]
He did not invent encryption, the one-time pad, or the idea that an opponent might know a cipher’s mechanism. Gilbert Vernam had published a description of cipher-printing telegraph equipment in 1926, and Shannon explicitly cited that work. The achievement here was not to claim every earlier trick. It was to give secrecy a mathematical framework, distinguish different kinds of security, and identify what the strongest claims would cost.[1][5]
The paper’s structure reflects that ambition. First comes a model of secrecy systems. Then comes theoretical secrecy: what could an adversary establish without a computational limit? (That is, with unlimited computational resources available.) Finally comes practical secrecy: how much work would exploiting the available evidence require?[1]
Those last two questions are not interchangeable. Much of what follows is about refusing to confuse them.
First, separate the jobs
Before introducing entropy or equations, it helps to separate three things that can look similar on a screen.
Encoding changes the representation. Turning text into a public sequence of numbers, or into another standard format, may make it unreadable to someone unfamiliar with the convention. It does not make the convention secret. The intended reader and an informed outsider can both perform the same reversal.
Encryption introduces a secret-dependent distinction. The rules may be public, but a key selects the particular transformation. Someone with the key can reverse it efficiently; someone without the key should not be able to learn the protected information within the system’s stated security model.[4]
Concealment tries to hide the communication itself. An encrypted letter displayed on a noticeboard advertises that a secret message exists. A message hidden in an apparently innocent object may not. Shannon separated concealment from the class of secrecy systems he analysed in detail. His chosen problem assumed that the adversary could obtain the cryptogram, his term for the encrypted message.[1]
A fourth operation, hashing, often gatecrashes this conversation. A cryptographic hash produces a digest, not an encrypted message that a recipient can decrypt to recover the original. Hashes are useful components of many security systems, but they do a different job. For this party, hashing is not on the guest list, so it does not get in.[4]
These distinctions matter because an unfamiliar appearance proves very little. A document written in a language you cannot read has not thereby been encrypted. Nor does a row of intimidating symbols suddenly become secure.
| THE ENVELOPE ANALOGY, WITH ITS LIMITS An opaque envelope hides the words inside, but its address, size, arrival time, and sender may remain visible. Encryption often protects contents while leaving comparable surrounding facts exposed. This is an analogy about scope, not about cryptographic strength. A paper envelope can be opened or held against a light. Perfect secrecy is a mathematical statement about what the intercepted data reveals, not a claim that the wrapping is physically difficult to remove.[1][12] |
Give the listener a fair chance
Imagine a sender, a recipient, and an eavesdropper. The sender starts with the plaintext, meaning the original information. This could be a sentence, a photograph, a file, or a single decision. Encryption combines it with a key to produce ciphertext. Decryption uses the appropriate key to recover the plaintext.[4]
In Shannon’s shared-key model, the sender and recipient already possess the same secret key. The eavesdropper sees the ciphertext but not that key. The problem of getting the key safely to both ends has not evaporated; it is an assumption of the model.[1]
| Participant | What they have | What they can do |
| Sender | Plaintext, public encryption rules, secret key | Produce the ciphertext |
| Recipient | Ciphertext, public decryption rules, the same secret key | Recover the plaintext |
| Eavesdropper | Ciphertext, the public rules, and relevant background knowledge | Try to infer the plaintext or key |
The deliberately uncomfortable part is how much the eavesdropper is allowed to know. Shannon assumes knowledge of the system, including how keys are selected. A security argument that requires the attacker not to have read the manual is a fragile argument. Manuals travel. Employees leave. Equipment is captured. Software can be examined.[1]
This does not mean that every detail of an operational installation should be published. It means the cipher’s core security should not collapse merely because its design becomes known. Keeping the key secret is a different task from keeping the algorithm mysterious.
The listener also has background knowledge. They may know the language, the message format, or the small set of things the sender is likely to say. They may already know most of a standard document. A useful theory has to allow for that, rather than treating the adversary as somebody free of all knowledge before the interception.
For theoretical secrecy, Shannon goes further: do not rely on the listener running out of time, staff, or computing power. Ask what the evidence permits in principle. The distinction between that idealised adversary and a feasible one will become important later, but providing an upper bound is fundamental.[1][9]
Information is not the same thing as importance
In ordinary speech, information means useful facts. In Shannon’s mathematics, the central quantity is uncertainty among possible outcomes, described using probabilities. It is not a measure of wisdom, truth, beauty, or how urgently somebody ought to answer an email.[2]
A fair coin provides a simple example. Before an unseen toss is revealed, there are two equally likely possibilities. Learning which occurred resolves one bit of uncertainty. Two independent fair tosses have four equally likely outcomes and contain two bits of uncertainty. Three have eight outcomes and contain three bits.
A bit is therefore both a familiar binary digit and a unit for measuring information. The two uses are related, but they are not identical. A file can contain a million binary digits without containing a million bits of uncertainty.
Suppose everyone knows in advance that a machine will print one million zeroes. The resulting file is large, but revealing it settles no uncertainty about its contents. Now suppose the machine instead produces one million independent, fair random bits, unseen by the observer. That source supplies a million bits of uncertainty. Same file length, radically different informational situation.
Shannon’s entropy measures the average uncertainty of a source with a specified probability distribution. Equally likely choices maximise it for a fixed number of possibilities. When some outcomes are much more likely than others, the uncertainty is lower. The word does not mean that the file has become messy, or that its contents are important.[2]
A one-bit message can be consequential. APPROVED or REFUSED might settle an application on which someone’s future depends. A long, repetitive document may convey less new information in this technical sense. Shannon’s measure counts uncertainty resolved, not human consequences.
| A LONG SECRET IS NOT NECESSARILY AN UNPREDICTABLE SECRET Imagine a key selected from just four publicly known possibilities, each equally likely. It contains only two bits of selection uncertainty, even if every candidate is written as a thousand-character string. Conversely, a short string selected from a very large set by a good random process may be difficult to predict. Length, visual complexity, and unpredictability are different properties. The process that produced the secret, together with what the attacker knows, matters more than how impressive it looks.[7] |
There is another trap: a random-looking output does not prove secrecy. Imagine sending a genuinely random message completely unencrypted. The output passes straight through and can look statistically excellent. The observer nevertheless sees every bit of the message.
The relevant question is not simply whether ciphertext looks random. It is how the ciphertext and the protected message are related. Statistical tests can find some defects, but NIST explicitly cautions that such tests cannot substitute for cryptanalysis.[17]
Perfect secrecy means no new evidence
Here is Shannon’s central idea in ordinary language:
After seeing the ciphertext, the eavesdropper should have exactly the same probabilities for the possible messages as before seeing it.[1][4]
Suppose the message will name one of four directions: NORTH, EAST, SOUTH, or WEST. Before the interception, the listener believes NORTH is likely. Perhaps the sender usually goes north.
Perfect secrecy does not force the listener to forget that. It does not make the four directions equally likely. It ensures that the encrypted message supplies no evidence that changes those existing probabilities.
If the listener’s best estimate was NORTH with a 70% probability before interception, it remains 70% afterwards. Guessing NORTH and being right is not, by itself, proof that the encryption was broken. The listener could have made the same guess without the ciphertext.
Conversely, a system can fail the definition without exposing the complete message. Suppose the ciphertext reveals only whether the direction is north-or-south rather than east-or-west. The listener still cannot name the destination, but has learned something. Perfect secrecy has already been lost.
That is why this definition is stronger than ‘the attacker cannot read the whole thing’. Protecting the last unknown word is not much comfort when every other detail has escaped into the wild.
Shannon called the uncertainty remaining after an observation equivocation. In modern notation, it is conditional entropy: the uncertainty about one thing when another thing is known. We can speak separately about remaining uncertainty in the message and in the key. They need not be the same.[1][9]
| NOT ‘VERY DIFFICULT’: UNDETERMINED BY THE EVIDENCE A conventional puzzle may have one answer that is extremely difficult to find. With perfect secrecy, the intercepted ciphertext does not select the original message from the alternatives at all. Trying every key can list possible messages. It cannot turn that list into evidence for the real one when the system has preserved their original probabilities. Faster computers do not cure missing evidence. Additional evidence, such as a stolen key or a copy of the plaintext, is another matter. |
Worked example
We can build a perfectly secret example with four numbers. It is a teaching example, not a proposed communications product.
Agree publicly that NORTH is 0, EAST is 1, SOUTH is 2, and WEST is 3. Imagine those numbers arranged round a four-position dial.
Before the message is chosen, sender and recipient privately share a fresh random shift: 0, 1, 2, or 3, each with equal probability. This is the key. The key must be independent of the direction they later wish to send.
To encrypt, start at the direction’s number and move forwards by the secret shift. Wrap round to 0 after 3. Send the number where you finish. To decrypt, the recipient moves backwards by the same shift.
For example, SOUTH is 2. With secret shift 3, count forwards: 3, 0, 1. The transmitted ciphertext is 1. A recipient who knows the shift moves backwards three places from 1 and recovers 2: SOUTH.
Now give an eavesdropper that intercepted 1. What could it mean?
| Possible plaintext | Public number | Secret shift that would produce 1 | Intercepted number |
| NORTH | 0 | 1 | 1 |
| EAST | 1 | 0 | 1 |
| SOUTH | 2 | 3 | 1 |
| WEST | 3 | 2 | 1 |
Every direction can produce that same ciphertext. More importantly, for every direction, the necessary key occurs with the same probability: one in four. Observing 1 therefore provides no reason to favour one direction over another beyond what was already known.
The same reasoning works for intercepted 0, 2, or 3. For each possible direction, every possible ciphertext has a one-in-four chance. The direction does not change the ciphertext’s probability distribution. That is the property we need.[4]
The recipient’s position is different. They know which row is compatible with the actual secret shift. Their extra information removes the ambiguity. Encryption has not destroyed the message; it has made the key necessary to identify it.
Now spoil the key generation. Suppose shift 0 is selected 70% of the time, with each other shift selected 10% of the time. With equally likely directions, intercepted 1 now makes EAST much more likely than the alternatives. The operation is unchanged, but the security property has gone.
This is a small, precise version of the lesson from the cillies: an impressive range of available settings is not the same thing as an unpredictable choice among them.
The one-time pad: perfection. With a supplies problem
The four-direction system is a small additive one-time cipher. For longer messages, the same principle becomes the one-time pad: combine each piece of information with a fresh, independent random piece of key material, which the recipient also possesses secretly.[4]
In a binary version, the combining operation is usually exclusive OR, abbreviated XOR. No knowledge of electronics is required to follow it. A key bit of 0 leaves a message bit unchanged. A key bit of 1 flips it: 0 becomes 1, and 1 becomes 0. Apply the same key bit again and the original returns.
| Message bit | Secret key bit | Ciphertext bit |
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 0 |
If the key bit is independent and equally likely to be 0 or 1, either plaintext bit produces either ciphertext bit with equal probability. Repeat with fresh independent key bits, and the argument extends to a whole fixed-length message.
For any proposed plaintext of that length, exactly one pad would turn it into the observed ciphertext. All those pads were equally likely. Even a machine that tried every pad would find no ciphertext-based reason to prefer the real message over the alternatives.[4]
The requirements are strict. The pad must be uniformly random over its possible values, independent of the message, secret from the attacker, and long enough to cover the protected data. No portion may be used again to protect another message. A page from a novel is not such a pad. Nor is a long output from a predictable process.
‘One-time’ describes the lifetime of the key material, not how often the recipient opens the document. Keeping a pad safe for later decryption is different from using the same pad to encrypt a second plaintext.
Vernam’s 1926 paper belongs to the earlier engineering history of combining telegraph messages with key tapes. Shannon’s contribution was to formalise the security question and its conditions, not to make earlier inventors disappear.[1][5]
The inconvenient requirement is shared secret material at scale. A gigabyte of arbitrary binary data needs a gigabyte of fresh pad for this construction. Both ends must obtain matching copies securely, keep track of their use, and protect them. Randomness has to be generated, distributed, stored, and retired without convenient shortcuts.[4]
| WHY SEND A PAD INSTEAD OF SENDING THE MESSAGE? Because the pad can be delivered before the message exists. Two people might exchange random material while they can meet safely, then use it later to communicate over a channel everyone can hear. That is useful. It is not magic key distribution. The secret material was transported through an earlier trusted opportunity, and the later communication consumes it. |
A correctly constructed and operated one-time-pad system does not become mathematically decipherable merely because the attacker obtains a faster computer, an AI system, or a quantum computer. Those tools may help find stolen pads, exploit devices, or improve guesses from outside information. They do not alter the independence established by the model.
Can the key simply be smaller?
The size requirement is not merely a defect of the pad’s design. It reflects a limit on the secrecy being requested.
Return to the four directions. For each observed ciphertext, perfect secrecy must leave all four directions possible with their proper probabilities. But a recipient with a particular key must obtain one definite direction. If there were only two possible keys, an observed ciphertext could decrypt to at most two directions. At least two possibilities would have been ruled out by interception alone.
That would already be information. No ingenious rearrangement of those two keys can avoid the counting problem.
For a finite, correctly decryptable system with perfect secrecy across its message space, there must be at least as many possible keys as possible messages. If every n-bit message is possible, there are 2n possibilities to accommodate. The ordinary n-bit uniform pad meets that bound.[4]
This does not mean that every human meaning requires a key as long as the prose used to express it. If sender and recipient have agreed that only APPROVED or REFUSED can be sent, the decision can be encoded as one bit before encryption. A public codebook can reduce the representation without itself providing secrecy.
Similarly, lossless compression can reduce the amount of data that needs protection. But variable output lengths may reveal something, and the probabilities and format still matter. The theorem has not been beaten by making a file smaller; the object being encrypted has changed.[1][15]
| SUPER-TECH 1 | THE BOUND, WITH THE ASSUMPTIONS LEFT ATTACHED Let M denote the plaintext, C the ciphertext, and K the shared secret key. Perfect secrecy is independence of M and C, equivalently P(M = m | C = c) = P(M = m) for every message m and every observation c of non-zero probability. In entropy notation: H(M | C) = H(M), or I(M; C) = 0 Correct decryption gives H(M | C, K) = 0. Consequently: H(M) = I(M; K | C) ≤ H(K | C) ≤ H(K) Thus the secret key must supply at least the message’s entropy. For a uniformly distributed n-bit message this is n bits. The finite-space counting condition, at least as many possible keys as possible messages, is also necessary; the entropy inequality alone is not a construction or a sufficient security test.[1][4] There is a second subtlety. Shannon entropy is an average, not a guarantee against a good first guess. Consider a 256-bit key generator that returns all-zeroes with probability 50% and spreads the remaining probability uniformly over all other keys. Its Shannon entropy is almost 129 bits. Yet an attacker who first tries all-zeroes succeeds half the time. Its min-entropy, −log₂ of the most likely outcome’s probability, is only one bit. NIST’s entropy-source guidance uses this more conservative quantity. The numerical example is deliberately constructed; it is not an estimate of any real generator. Not a sane one, anyway.[7] |
Reuse the pad and the theorem exits stage left…
…pursued by a (likely Russian) bear.
Consider two messages protected with exactly the same binary pad. Each individual ciphertext may look reassuringly random. The pair is another matter.
XOR the two ciphertexts together and the repeated pad cancels. What remains is the XOR relationship between the two plaintexts. The attacker may not immediately recover either complete message, but the promised independence has disappeared. Known text in one can expose the corresponding text in the other.[4]
The four-direction version shows the same danger without binary notation. If two messages use the same secret shift, equal ciphertext numbers mean equal plaintext directions. Different ciphertext numbers reveal the relative displacement between the directions. One known direction reveals the shift and allows the other to be decoded.
That is already a substantial leak. A promise of secrecy does not survive merely because the attacker still has some work to do.
The VENONA material provides a historical example of the consequences of reused pad material. NSA histories describe duplicate key pages and the sustained analytical work needed to exploit messages encrypted in depth. This was not a refutation of perfect secrecy. It was exploitation of a system that had violated a condition necessary to support it.[8]
| A NONCE IS NOT A REPLACEMENT PAD Modern encryption often uses a public nonce, a value required to be fresh or unique within a specified context, alongside a reusable secret key. The nonce is not normally secret, and it is not a one-time pad. Its job depends on the construction. In ChaCha20, repeating the same key and nonce from the same counter position repeats the keystream, recreating the dangerous two-message relationship. Other constructions have their own rules and failure modes. ‘The value is public’ and ‘it does not matter if we repeat it’ are very different statements.[18] |
There is an important engineering distinction here. Reusing an independent pad bit is forbidden by the pad’s design. Reusing a modern cipher key may be expressly supported by its design, provided its nonce, usage-limit, and protocol requirements are followed. A slogan about never reusing any key would be both impractical and wrong.
Why ordinary language helps the attacker
Messages are rarely selected uniformly from every possible string. Language has structure. So do spreadsheets, image formats, protocol headers, and routine reports.
Consider a simple letter-substitution cipher: every A becomes one other letter, every B another, and so on. If a ciphertext word repeats its first letter in its third position, the plaintext does too. Even without knowing the alphabet substitution, the observer has learned a relationship within the word. More text can provide frequencies, repeated phrases, and increasingly strong constraints.[4]
This is redundancy in its information-theoretic sense: the predictability imposed by the source’s structure. It does not mean the prose is badly written or that the message contains unnecessary business information. A sentence can be concise and still obey spelling, grammar, and familiar patterns.[2][6]
A reader can often complete ‘Please close the d…’ because the surrounding text constrains what is likely. A cryptanalyst exploits comparable constraints, often with more patience and fewer assumptions about the sender’s literary ambitions.
In 1951 Shannon investigated the entropy of printed English by asking people to predict successive letters using preceding text. Human readers became measuring instruments for the language knowledge that simple letter counts miss. The estimates depended on how much context was available and on the kind of text; they were not timeless constants attached to the English language.[6]
There is a pleasing reversal here. The same predictability that helps a recipient repair a garbled message can help an adversary recognise a plausible decryption. Helpful redundancy and dangerous redundancy are not different substances. Their value depends on who can exploit them.
The result is not a recommendation to remove all structure from everything. Modern encryption should protect highly structured messages too. It is an explanation of why historical ciphers leaked more than their designers could see, and why a believable-looking output is not by itself proof of a successful decryption.
Unicity: enough evidence is not enough computing
With a short intercepted message, several keys may produce plausible plaintexts. As more material under the relevant key becomes available, language and format constraints can eliminate alternatives. Eventually there may be essentially one credible solution.
Shannon studied that transition through unicity distance: roughly, how much intercepted material is needed before ambiguity between plausible solutions largely disappears. It is an information question, not a stopwatch.[1]
A useful analogy is a crossword. Early clues may allow several completions. Enough intersecting clues can determine one solution, yet finding it may still require considerable work. The evidence can be sufficient before the solver becomes capable of using it.
Conversely, insufficient evidence can leave several valid completions no matter how long the solver works. A faster pencil does not invent the missing clue.
Shannon’s familiar estimate compares uncertainty in the key with the rate at which source redundancy supplies constraints. In a simplified random-cipher model it is often written as key entropy divided by redundancy per character. This is an estimate with assumptions, not a universal threshold at which all encryption fails. Hellman’s later analysis sharpened the model and its limitations.[1][9]
For scale, a uniformly chosen substitution alphabet over 26 letters has about 88.4 bits of key uncertainty: there are 26 choices for the first replacement, then 25 for the next, and so on, multiplied together. Mathematicians write that product as 26!. With an illustrative redundancy of three bits per character, the rough quotient is about 29.5 characters. Change the language model, key distribution, or cipher, and the result changes. It does not promise that every thirty-letter cryptogram has one easy answer.
Nor should that arithmetic be transferred casually to a modern cipher. A key or plaintext can be determined in principle by available data while remaining computationally inaccessible. Equally, a system can leave multiple possible complete plaintexts while revealing damaging partial information.
Unique does not mean easy. Ambiguous does not mean secret.
| SUPER-TECH 2 | SHANNON’S ‘IDEAL’ IS NOT THE SAME AS PERFECT The paper distinguishes perfect secrecy from source-dependent ideal systems. In particular, uncertainty about the key can remain unchanged even while a great deal about the message leaks.[1][9] Here is a deliberately stark example. Let M be n independent fair bits. Choose one independent fair secret bit K. Encrypt by leaving every message bit unchanged when K = 0, or complementing every bit when K = 1. After seeing C, an eavesdropper has exactly two equally likely plaintext candidates: C itself and its bitwise complement. The key remains wholly uncertain, however long the message. But: H(K | C) = 1; H(M | C) = 1; I(M; C) = n − 1 Only one bit of message uncertainty survives. For every pair of positions, the observer knows whether the plaintext bits agree. Knowing a single plaintext bit would reveal the key and all the remaining plaintext. This is an illustration of Shannon’s source-dependent strongly ideal category, not a secure construction. It makes the distinction unavoidable: failure to recover the key, or to name one unique whole message, is not proof that all useful information has remained secret. |
Make the evidence infeasible to exploit
If perfect secrecy consumes so much shared randomness, why does ordinary encrypted communication work with comparatively short keys?
Because its target is usually computational security, not perfect secrecy against unlimited computation. The aim is to prevent feasible adversaries from obtaining a useful advantage, under explicit assumptions about the algorithms, resources, and system.[4][10]
A well-designed cryptographic generator can expand a short secret seed into a long stream that a feasible observer should not be able to distinguish usefully from fresh randomness. But a deterministic expansion does not create unlimited independent entropy. Its outputs are restricted by the seed. Security rests on the difficulty of recognising or exploiting that restriction, not on pretending it does not exist.[4]
AES similarly uses a finite key to select a transformation. The standard specifies 128-bit blocks and keys of 128, 192, or 256 bits. Proper modes and protocols allow such building blocks to protect messages, while adding requirements beyond the bare block transformation. Those key lengths are not claims that arbitrarily long messages enjoy one-time-pad perfect secrecy.[11]
The intended advantage is asymmetry of work. The recipient with the key can recover the message efficiently. An attacker without it should face an infeasible problem. Known message formats need not ruin a modern cipher; resistance to known-plaintext and stronger attacks is part of what designers must establish.[4]
Shannon already separated theoretical ambiguity from the practical labour of solving a cipher. Later cryptography developed more precise computational security definitions. Goldwasser and Micali’s work on probabilistic encryption, published in journal form in 1984, formalised a powerful version of the aim: an efficient attacker should not gain useful information about the message merely by receiving its encryption, subject to the construction’s computational assumptions.[1][10]
That is the modern descendant of the question, not a claim that Shannon wrote today’s complete security definitions in 1949.
| WHY ‘NOT PERFECT’ DOES NOT MEAN ‘NOT WORTH USING’ A security claim can be extremely useful without defeating an unlimited adversary. The important distinction is between an honest, well-supported computational claim and a claim that hides its assumptions. ‘No practical attack is known’ is not a mathematical proof of perfect secrecy. Equally, an impossibility theorem for perfect secrecy with short keys is not a proof that practical encryption is useless. |
Confusion, diffusion, and some serious… pastry
Shannon’s paper is not solely a demonstration of an expensive ideal. It also asks how to make practical cryptanalysis difficult. Two enduring design ideas are confusion and diffusion.[1]
Diffusion spreads the consequences of the message’s structure across combinations of ciphertext elements. In a simple substitution, a common plaintext letter stays a common ciphertext letter. A better mixing process makes simple counts much less informative, so useful relationships become distributed across more complicated combinations.
Confusion makes the connection between observable ciphertext properties and the secret key difficult to exploit. A statistic may constrain the key without providing a conveniently solvable equation for one small part of it. The intended difficulty is for the analyst, not for the person reading the product manual.
The ideas are related but distinct: spread out the evidence, and make its relationship to the key hard to untangle. The design still needs to be reversible for the authorised recipient. Encryption is not a blender whose security rests on having destroyed the original; if that worked life would be too easy.
| THE PASTRY IS ACTUALLY BAKED INTO THE PAPER In section 25, Shannon invokes mixing transformations and cites Eberhard Hopf’s example of repeatedly rolling out and folding pastry dough. Alternating simple operations can distribute an initially concentrated region throughout the mixture.[1, p712] The analogy has limits: a finite digital permutation is not literally an infinitely divisible lump of dough. Shannon explicitly recognises the finite-versus-infinite distinction. The useful intuition is that repeated operations can create complicated global relationships from simple local steps. Modern AES uses repeated rounds involving substitution, rearrangement, linear mixing, and key addition. That is a concrete later example of the design family, not evidence that Shannon specified or predicted AES in advance.[11] |
More layers do not automatically mean more security. Two consecutive Caesar shifts are just one Caesar shift with a different total displacement. Likewise, composing ordinary substitution alphabets produces another substitution alphabet. An enormous-looking procedure can remain inside the same weak family.
Good composition has to change what the adversary can exploit. Complexity is not a security property merely because it was difficult to implement.
There is another important qualification: for a fixed key, a reversible transformation rearranges possible inputs; it does not manufacture new uncertainty about the input. The secrecy comes from what the observer does not know, and computational protection from the difficulty of exploiting what they do. Confusion and diffusion do not suspend those accounting rules.
Compression helps. Except when it doesn’t
If redundancy helps recognise a plausible message, removing redundant representation before encryption can reduce the amount of material exposed to that kind of analysis. Shannon discusses this connection. It is tempting to turn it into a universal instruction: compress, then encrypt, and security must improve.[1]
That conclusion is too quick. Compression also creates an observable result: the compressed length. Two equal-length inputs can compress to different sizes because one contains more predictable structure. If those sizes remain visible after encryption, they can reveal something about the plaintext.
John Kelsey’s 2002 work demonstrated how apparently small compression leaks can become useful, including when an attacker can influence neighbouring input. The information need not come from recovering an encryption key. It can come from measuring how the system responds to guesses.[15]
TLS 1.3 separates these issues in a practical protocol. It removes the older TLS-level compression mechanism, uses authenticated encryption, and permits record padding; it also explicitly warns that traffic lengths are not automatically hidden. Application-level behaviour still needs its own analysis.[12]
| SUPER-TECH 3 | REDUNDANCY’S LOCATION MATTERS A public deterministic encoding F applied after encryption cannot create information about M that was absent from C: M → C → F(C) is a data-processing chain. If F() is invertible on its outputs, the observer gets exactly the same information as from C, merely in another representation.[2] Thus a public error-correcting code applied to perfectly secret ciphertext does not, by itself, destroy content secrecy. It helps reconstruct an object that was already safe to expose. Encoding a ciphertext is not the same thing as exposing predictable relationships in the plaintext. However, a transform whose length or behaviour depends on secret plaintext, or an extra observable error response, can create another channel. The relevant observation is then not C alone. Include the lengths, responses, timing, and other available outputs in the model. Kelsey’s compression analysis supplies a concrete example.[15] This is why ‘redundancy is bad for secrecy’ is an inadequate design rule. Ask where it is introduced, what depends on the secret, and what the adversary can actually observe. |
What perfect secrecy does not promise
Suppose our one-time pad is ideal. The ciphertext reveals no information about the message. Does that make the communication safe?
Not by itself.
It does not authenticate the sender. A receiver needs a separate reason to accept the message as coming from the intended party. Possession of some bytes that decrypt is not a general proof of origin.
It does not prevent tampering. In the binary pad, flipping a ciphertext bit flips the corresponding decrypted plaintext bit. An attacker can change something without first learning what it was. In the four-direction example, adding one position to the ciphertext rotates the recipient’s recovered direction by one position. Confidentiality can survive while integrity fails.[4]
It does not prevent replay. A previously valid message may be copied and presented again unless the protocol includes suitable freshness checks. Yesterday’s authentic instruction need not be an instruction to perform the action again today.[4]
It does not ensure delivery. Someone can discard the ciphertext, cut the cable, or make the endpoint unavailable. A secrecy theorem does not keep the electricity on.
It does not necessarily hide metadata. Who communicates with whom, when, how often, and with what message sizes may remain visible. If the only possible plaintexts are differently sized, preserving their lengths may itself identify them. The elementary perfect-secrecy example fixes the message space and length; a system claiming more has to account for those observations too.[12]
It does not secure the endpoints. Malware, a stolen pad, an exposed screen, a plaintext log, or an unintended physical signal can provide information outside the ciphertext channel. That does not disprove the theorem. It demonstrates that the attacker was given more than the theorem assumed.[16]
| A FOOTNOTE WITH A MODERN-SOUNDING IDEA Shannon noticed that an interceptor might learn simply that a message had been sent. In footnote 9 on page 680, he proposes treating ‘no message’ as an encipherable blank and transmitting that when there is nothing substantive to send.[1] The idea anticipates the logic of cover traffic: activity need not announce that meaningful communication has occurred. It is not a complete anonymity system. Timing, lengths, routing, and observation of the endpoints would still need attention. NO MESSAGE IS STILL A MESSAGE: OPERATION RYAN Encryption protects what a message says. It does not necessarily conceal the fact that something is happening. That may be important in itself. During the Soviet Operation RYAN intelligence programme in the early 1980s, KGB officers were told to watch for signs that NATO might be preparing for nuclear war. Among the indicators were unusually late activity in government and military buildings. An instruction sent to the KGB’s London residency on 17 February 1983 explicitly required officers to establish the normal level of activity at important government institutions and headquarters. Among the indicators were the numbers of illuminated windows during and outside working hours, with particular attention to increases at night or activity on normally non-working days.[19] No secret document had been intercepted. No cipher had been broken. The pattern itself was information. An obvious defensive response would have been cover: routinely leave a varying selection of windows illuminated whether or not the offices were occupied, to make it harder for an external observer to distinguish normal activity from extraordinary activity. The digital equivalent is dummy traffic or traffic padding: sending communications which carry no useful message so that silence, timing and volume reveal less about the real ones. Merely choosing a random number of windows would not automatically solve the problem. If the statistical pattern of the decoys differed from genuine activity, a patient observer might eventually separate them. Modern traffic-flow protection has the same problem: random dummy packets can disguise periods of silence, but a sufficiently revealing traffic pattern may still leak information. A stronger defence against this particular indicator may be a deliberately shaped or even constant apparent level of activity.[20][21] Sometimes the secret is not in the message. The secret is that there was a message. |
The difference between secrecy and total security is not an academic nuisance. It explains why a correct cryptographic component can sit inside a failing service. The service has other promises to keep.
The assumptions are part of the result
A theorem is not weakened by having assumptions. It is weakened in use when somebody removes them from the glossy brochure, or otherwise hides or loses them.
The shared-key model assumes a secret available to the legitimate parties and unavailable to the listener. The perfect-secrecy example assumes a particular message space, fresh independent randomness, and an observation limited to the ciphertext. Correct decryption is required. These are not decorative clauses.
Later work explored different models. Diffie and Hellman’s 1976 public paper addressed communication without an already shared secret key and described a computational approach to key agreement. That changes the key-establishment problem, not the arithmetic of Shannon’s shared-secret bound. Authentication of the exchange remains a separate issue.[13]
Aaron Wyner’s 1975 wire-tap channel considered a different advantage: an eavesdropper whose observation passes through an additional noisy channel. Under suitable channel conditions, reliable communication with information-theoretic secrecy becomes possible as code length grows, without the same pre-shared-pad arrangement. The extra resource is a difference between what the recipient and the eavesdropper can observe. Again, the model changes; a theorem is not contradicted.[14]
The practical lesson is not that one of these models is the only respectable one. It is that a security claim must say which advantage is being used: secret material, computational difficulty, a physical observation advantage, or some combination.
Clear assumptions, clear rules, and a clear understanding of what security claims remain valid within them. And what happens to those claims if the rules or assumptions cease to be valid.
That is also the bridge to the next artefact, TEMPEST. A cipher may be perfect against someone who sees its intended output while the equipment leaks plaintext through another physical path. A declassified NSA history describes precisely why compromising emanations matter to cryptographic equipment. The attack can bypass the mathematical channel rather than defeat its encryption; physics can trump maths.[16]
The proof may be correct. The diagram may be missing a wire.
What to ask when someone says ‘encrypted’
For a reader who never intends to design a cipher, the most useful outcome is a better set of questions.
What is protected: the contents, the identities, the timing, or only one of those?
Who holds the keys, and can the service provider also read the plaintext?
How are keys generated, distributed, and protected?
Is the message authenticated as well as encrypted?
What happens when randomness repeats, a device is stolen, or the communication is interrupted?
What observations are outside the security claim?
Those questions do not require memorising an entropy formula. They require recognising that ‘encrypted’ names a mechanism, not a complete account of who can learn what.
For technical readers, the corresponding discipline is to state the adversary, the observations, the correctness requirement, the key and message distributions, and the security property. Distinguish information-theoretic independence from computational indistinguishability.
Do not replace an argument with a randomness plot, a key length, or the assertion that nobody has complained yet.
And do not mistake a theorem’s silence about a threat for a promise that the threat does not exist.
What the artefact really is
The physical artefact is the 1949 paper. It is less photogenic than a cipher machine, although considerably easier to carry.
The more important artefact is actually the change in question.
Instead of asking whether a message looks scrambled, ask what an observer can infer. Instead of counting settings and declaring victory, ask how the key was chosen. Instead of saying that a cipher is unbreakable, distinguish absent evidence from inaccessible evidence. Instead of protecting an algorithm in isolation, identify everything the adversary can see.
Shannon did not make secrecy simple. He made the claims separable enough to examine. He made it possible to say much more clearly which claims stand, and under what circumstances.
The one-time pad shows that perfect content secrecy is possible. Its demanding assumptions show why operating it is a different achievement. Practical cryptography shows why weaker, carefully defined claims can still support extraordinarily useful systems. The attacks around the edges show why no theorem can protect a channel it was never asked to model.
The message must remain recoverable to someone.
The question is what makes that person different from everyone listening.
Shannon made us account for the difference.

00010111 ETB
References & Source notes
The historical argument centres on Shannon’s 1949 paper. Later sources are identified as later work, rather than being folded into claims about what he had already proved. Bracketed references in the article link to the numbered entries below.
The four-direction cipher, biased-key calculation, complement-only example, and illustrative unicity calculation are worked teaching examples prepared for this article. They are not proposed encryption products, nor reported measurements of existing systems.
Edition check: the predecessor report is dated 1 September 1945 in the original printed footnote on page 656. A commonly circulated re-typeset copy says 1946; the original-page facsimile is the authority used here.
Historical content and cited standards were reviewed for this draft on 16 & 17 September 2026. A historical publication is not presented as current implementation guidance.
[1] Claude E. Shannon. Communication Theory of Secrecy Systems
Bell System Technical Journal 28(4), October 1949, pp. 656–715.
The principal artefact. Original pagination: perfect secrecy, pp. 679–682; ideal systems, sections 17–18; practical secrecy, Part III; pastry and mixing, section 25, p. 712. The opening footnote dates the predecessor report to 1 September 1945.
Publisher record | Original-page facsimile | DOI
[2] Claude E. Shannon. A Mathematical Theory of Communication
Bell System Technical Journal 27, July and October 1948, pp. 379–423 and 623–656.
The earlier communication-theory paper: uncertainty, entropy, conditional entropy, source structure, and noisy channels. The article uses the resulting information identities in its own worked examples.
[3] Nokia Bell Labs. Claude Shannon
Institutional historical profile.
Biographical and Bell Laboratories context. Used for that context, not as a substitute for the 1949 paper itself.
[4] Dan Boneh and Victor Shoup. A Graduate Course in Applied Cryptography
Version 0.6, January 2023.
Chapters 2–3 for perfect and computational security, the one-time pad, key-space bounds, and stream ciphers; chapters 6 and 9 for authentication and authenticated encryption. Modern notation and explanations are distinguished from Shannon’s historical terminology.
[5] Gilbert S. Vernam. Cipher Printing Telegraph Systems for Secret Wire and Radio Telegraphic Communications
Journal of the American Institute of Electrical Engineers 45(2), February 1926, pp. 109–115.
Primary engineering account of telegraph encryption using key material. It predates Shannon’s mathematical treatment and is cited in the 1949 paper.
[6] Claude E. Shannon. Prediction and Entropy of Printed English
Bell System Technical Journal 30(1), January 1951, pp. 50–64.
The letter-prediction experiments and the importance of context when estimating the redundancy of a language source. No single numerical estimate is treated here as a universal constant for English.
[7] NIST. Recommendation for the Entropy Sources Used for Random Bit Generation
Special Publication 800–90B, January 2018.
Section 2.1 and Appendix D for min-entropy and its relation to the most likely output. The biased 256-bit generator in the article is an original numerical illustration, not a reported defective product.
Publication record | Full publication
[8] National Security Agency. VENONA: An Overview
Cryptologic Almanac 50th Anniversary Series, declassified historical account.
Duplicate one-time-pad material and the work needed to exploit it. The article uses the documented cryptographic failure, without treating the programme as a simple one-step decryption story.
[9] Martin E. Hellman. An Extension of the Shannon Theory Approach to Cryptography
IEEE Transactions on Information Theory IT-23(3), May 1977, pp. 289–294.
The random-cipher model, unicity, spurious solutions, and the limits of inferring security merely from remaining ambiguity.
[10] Shafi Goldwasser and Silvio Micali. Probabilistic Encryption
Journal of Computer and System Sciences 28(2), 1984, pp. 270–299.
A later formalisation of computationally protecting partial information. This is a development after Shannon, not a definition retrospectively attributed to him.
[11] NIST. Advanced Encryption Standard (AES)
FIPS 197, originally 2001; editorial update 9 May 2023.
AES block and key sizes and its round transformations. Used as a concrete later example, not as a claim that Shannon designed AES.
Standard record | Full standard
[12] Eric Rescorla. The Transport Layer Security (TLS) Protocol Version 1.3
RFC 8446, August 2018.
Section 1 for security goals; removal of TLS-level compression; section 5.4 for record padding; Appendix E.3 for traffic analysis. TLS records do not automatically conceal all lengths or application behaviour.
[13] Whitfield Diffie and Martin E. Hellman. New Directions in Cryptography
IEEE Transactions on Information Theory IT-22(6), November 1976, pp. 644–654.
The public proposal for computational key agreement and public-key cryptography. Changing how a key is established does not refute a theorem about a different model.
[14] Aaron D. Wyner. The Wire-Tap Channel
Bell System Technical Journal 54(8), October 1975, pp. 1355–1387.
A noisy-channel observation advantage and asymptotic information-theoretic secrecy. This differs from perfect secrecy in the elementary finite, pre-shared-key model.
[15] John Kelsey. Compression and Information Leakage of Plaintext
Fast Software Encryption 2002, LNCS 2365, pp. 263–276.
Plaintext information revealed by compressed length, including the additional leverage from attacker-influenced input. Used to qualify a simplistic compress-then-encrypt rule.
[16] National Security Agency. TEMPEST: A Signal Problem
Cryptologic Spectrum 2(3), Summer 1972; subsequently declassified.
Public historical evidence of compromising emanations from cryptographic equipment. The bridge to artefact 024 uses this public source, not non-public operational guidance.
[17] NIST. A Statistical Test Suite for Random and Pseudorandom Number Generators for Cryptographic Applications
Special Publication 800–22, Revision 1a, April 2010.
The explicit warning that statistical testing cannot substitute for cryptanalysis. A satisfactory-looking output distribution is not by itself a confidentiality argument.
[18] Yoav Nir and Adam Langley. ChaCha20 and Poly1305 for IETF Protocols
RFC 8439, June 2018.
Sections 2.3 and 4 for the stream construction and nonce requirements. Repeating key, nonce, and counter position repeats keystream; this is not a universal description of every cipher’s nonce behaviour.
[19] KGB Headquarters Moscow to the London KGB Residency. Permanent Operational Assignment to Uncover NATO Preparations for a Nuclear Missile Attack on the USSR
17 February 1983, Top Secret; subsequently published by Oleg Gordievsky and Christopher Andrew and reproduced by the National Security Archive.
Primary Operation RYAN instruction requiring the London residency to establish normal patterns of activity at important government institutions and headquarters, including numbers of illuminated windows during and outside working hours, and to report significant departures from those patterns. Benjamin B. Fischer’s CIA historical study A Cold War Conundrum: The 1983 Soviet War Scare provides the wider Operation RYAN and war-scare context.
National Security Archive: KGB Operation RYAN instruction
CIA historical study: A Cold War Conundrum: The 1983 Soviet War Scare
[20] R. Shirey. Internet Security Glossary, Version 2
RFC 4949, August 2007.
Defines “communications cover” as concealing or altering characteristic communications patterns so that they do not reveal useful information to an adversary.
RFC 4949: Internet Security Glossary, Version 2
[21] Stephen Kent. IP Encapsulating Security Payload (ESP)
RFC 4303, December 2005, sections 2.6–2.7.
Concrete protocol example of traffic-flow confidentiality: dummy packets may be inserted, including during otherwise silent periods, and traffic may be padded or shaped to conceal observable communication patterns.
RFC 4303: IP Encapsulating Security Payload (ESP)
