About Sophie

Cybersecurity, assurance, technology, identity, and life.

“I dare do all that may become a woman. Who dares do more is none.”

New Series: Sophie Baskerville’s History of Cybersecurity

sophie @ baskerville.net ©️1977-2026 Sophie Baskerville

023 | Shannon’s Communication Theory of Secrecy Systems


“Unique does not mean easy.
Ambiguous does not mean secret.”

Collage in purple and sepia illustrating Claude Shannon’s work on secrecy. On the left, a thoughtful man rests his chin on his hand beside books and a paper titled “Communication Theory of Secrecy Systems”, dated 1949. Across the centre, a diagram shows a sender encrypting “Meet tomorrow at noon” with a secret key, transmitting patterned ciphertext, and a recipient using a matching key to recover the message. Beneath the transmission, a shadowy eavesdropper is shown with question marks. Mathematical formulae, logic circuits, and curling strips of printed tape surround the scene, with an illuminated Bell Telephone Laboratories building and a tower in the background. This is a symbolic illustration, not an archival photograph.

Communications Secrecy

The subject is intensely practical. Two people want to communicate. Somebody else can listen. What would it mean for the listener to learn nothing useful from the intercepted message? And how could we distinguish that situation from one in which the listener simply has not worked out the answer yet?

Images in this article are generated by

EU AI Marker icon

EU AI Act Regulation 2024/1689

Claude Elwood Shannon’s Communication Theory of Secrecy Systems appeared in the Bell System Technical Journal in October 1949, volume 28, pages 656–715. The opening footnote traces its material to a confidential report, A Mathematical Theory of Cryptography, dated 1 September 1945. The public paper therefore emerged from wartime work.[1][3]

It is not the same paper as Shannon’s A Mathematical Theory of Communication, published in two instalments in 1948. The titles are easy to muddle; the questions are different. The 1948 work established a general mathematical treatment of communication, including the effects of noise and the statistical structure of messages. The 1949 paper brought that same way of thinking to the subject of secrecy.[1][2]

Shannon was working at Bell Laboratories, an institution concerned with making signals intelligible across real channels. Wartime projects also required some signals to remain unintelligible to the wrong people. That is a useful intellectual position: the engineer must understand both how information arrives and how it might escape.[3]

He did not invent encryption, the one-time pad, or the idea that an opponent might know a cipher’s mechanism. Gilbert Vernam had published a description of cipher-printing telegraph equipment in 1926, and Shannon explicitly cited that work. The achievement here was not to claim every earlier trick. It was to give secrecy a mathematical framework, distinguish different kinds of security, and identify what the strongest claims would cost.[1][5]

The paper’s structure reflects that ambition. First comes a model of secrecy systems. Then comes theoretical secrecy: what could an adversary establish without a computational limit? (That is, with unlimited computational resources available.) Finally comes practical secrecy: how much work would exploiting the available evidence require?[1]

Those last two questions are not interchangeable. Much of what follows is about refusing to confuse them.

First, separate the jobs

Before introducing entropy or equations, it helps to separate three things that can look similar on a screen.

Encoding changes the representation. Turning text into a public sequence of numbers, or into another standard format, may make it unreadable to someone unfamiliar with the convention. It does not make the convention secret. The intended reader and an informed outsider can both perform the same reversal.

Encryption introduces a secret-dependent distinction. The rules may be public, but a key selects the particular transformation. Someone with the key can reverse it efficiently; someone without the key should not be able to learn the protected information within the system’s stated security model.[4]

Concealment tries to hide the communication itself. An encrypted letter displayed on a noticeboard advertises that a secret message exists. A message hidden in an apparently innocent object may not. Shannon separated concealment from the class of secrecy systems he analysed in detail. His chosen problem assumed that the adversary could obtain the cryptogram, his term for the encrypted message.[1]

A fourth operation, hashing, often gatecrashes this conversation. A cryptographic hash produces a digest, not an encrypted message that a recipient can decrypt to recover the original. Hashes are useful components of many security systems, but they do a different job. For this party, hashing is not on the guest list, so it does not get in.[4]

These distinctions matter because an unfamiliar appearance proves very little. A document written in a language you cannot read has not thereby been encrypted. Nor does a row of intimidating symbols suddenly become secure.

Give the listener a fair chance

Imagine a sender, a recipient, and an eavesdropper. The sender starts with the plaintext, meaning the original information. This could be a sentence, a photograph, a file, or a single decision. Encryption combines it with a key to produce ciphertext. Decryption uses the appropriate key to recover the plaintext.[4]

In Shannon’s shared-key model, the sender and recipient already possess the same secret key. The eavesdropper sees the ciphertext but not that key. The problem of getting the key safely to both ends has not evaporated; it is an assumption of the model.[1]

The deliberately uncomfortable part is how much the eavesdropper is allowed to know. Shannon assumes knowledge of the system, including how keys are selected. A security argument that requires the attacker not to have read the manual is a fragile argument. Manuals travel. Employees leave. Equipment is captured. Software can be examined.[1]

This does not mean that every detail of an operational installation should be published. It means the cipher’s core security should not collapse merely because its design becomes known. Keeping the key secret is a different task from keeping the algorithm mysterious.

The listener also has background knowledge. They may know the language, the message format, or the small set of things the sender is likely to say. They may already know most of a standard document. A useful theory has to allow for that, rather than treating the adversary as somebody free of all knowledge before the interception.

For theoretical secrecy, Shannon goes further: do not rely on the listener running out of time, staff, or computing power. Ask what the evidence permits in principle. The distinction between that idealised adversary and a feasible one will become important later, but providing an upper bound is fundamental.[1][9]

Information is not the same thing as importance

In ordinary speech, information means useful facts. In Shannon’s mathematics, the central quantity is uncertainty among possible outcomes, described using probabilities. It is not a measure of wisdom, truth, beauty, or how urgently somebody ought to answer an email.[2]

A fair coin provides a simple example. Before an unseen toss is revealed, there are two equally likely possibilities. Learning which occurred resolves one bit of uncertainty. Two independent fair tosses have four equally likely outcomes and contain two bits of uncertainty. Three have eight outcomes and contain three bits.

A bit is therefore both a familiar binary digit and a unit for measuring information. The two uses are related, but they are not identical. A file can contain a million binary digits without containing a million bits of uncertainty.

Suppose everyone knows in advance that a machine will print one million zeroes. The resulting file is large, but revealing it settles no uncertainty about its contents. Now suppose the machine instead produces one million independent, fair random bits, unseen by the observer. That source supplies a million bits of uncertainty. Same file length, radically different informational situation.

Shannon’s entropy measures the average uncertainty of a source with a specified probability distribution. Equally likely choices maximise it for a fixed number of possibilities. When some outcomes are much more likely than others, the uncertainty is lower. The word does not mean that the file has become messy, or that its contents are important.[2]

A one-bit message can be consequential. APPROVED or REFUSED might settle an application on which someone’s future depends. A long, repetitive document may convey less new information in this technical sense. Shannon’s measure counts uncertainty resolved, not human consequences.

There is another trap: a random-looking output does not prove secrecy. Imagine sending a genuinely random message completely unencrypted. The output passes straight through and can look statistically excellent. The observer nevertheless sees every bit of the message.

The relevant question is not simply whether ciphertext looks random. It is how the ciphertext and the protected message are related. Statistical tests can find some defects, but NIST explicitly cautions that such tests cannot substitute for cryptanalysis.[17]

Perfect secrecy means no new evidence

Here is Shannon’s central idea in ordinary language:

After seeing the ciphertext, the eavesdropper should have exactly the same probabilities for the possible messages as before seeing it.[1][4]

Suppose the message will name one of four directions: NORTH, EAST, SOUTH, or WEST. Before the interception, the listener believes NORTH is likely. Perhaps the sender usually goes north.

Perfect secrecy does not force the listener to forget that. It does not make the four directions equally likely. It ensures that the encrypted message supplies no evidence that changes those existing probabilities.

If the listener’s best estimate was NORTH with a 70% probability before interception, it remains 70% afterwards. Guessing NORTH and being right is not, by itself, proof that the encryption was broken. The listener could have made the same guess without the ciphertext.

Conversely, a system can fail the definition without exposing the complete message. Suppose the ciphertext reveals only whether the direction is north-or-south rather than east-or-west. The listener still cannot name the destination, but has learned something. Perfect secrecy has already been lost.

That is why this definition is stronger than ‘the attacker cannot read the whole thing’. Protecting the last unknown word is not much comfort when every other detail has escaped into the wild.

Shannon called the uncertainty remaining after an observation equivocation. In modern notation, it is conditional entropy: the uncertainty about one thing when another thing is known. We can speak separately about remaining uncertainty in the message and in the key. They need not be the same.[1][9]

Worked example

We can build a perfectly secret example with four numbers. It is a teaching example, not a proposed communications product.

Agree publicly that NORTH is 0, EAST is 1, SOUTH is 2, and WEST is 3. Imagine those numbers arranged round a four-position dial.

Before the message is chosen, sender and recipient privately share a fresh random shift: 0, 1, 2, or 3, each with equal probability. This is the key. The key must be independent of the direction they later wish to send.

To encrypt, start at the direction’s number and move forwards by the secret shift. Wrap round to 0 after 3. Send the number where you finish. To decrypt, the recipient moves backwards by the same shift.

For example, SOUTH is 2. With secret shift 3, count forwards: 3, 0, 1. The transmitted ciphertext is 1. A recipient who knows the shift moves backwards three places from 1 and recovers 2: SOUTH.

Now give an eavesdropper that intercepted 1. What could it mean?

Possible plaintextPublic numberSecret shift that
would produce 1
Intercepted
number
NORTH011
EAST101
SOUTH231
WEST321

Every direction can produce that same ciphertext. More importantly, for every direction, the necessary key occurs with the same probability: one in four. Observing 1 therefore provides no reason to favour one direction over another beyond what was already known.

The same reasoning works for intercepted 0, 2, or 3. For each possible direction, every possible ciphertext has a one-in-four chance. The direction does not change the ciphertext’s probability distribution. That is the property we need.[4]

The recipient’s position is different. They know which row is compatible with the actual secret shift. Their extra information removes the ambiguity. Encryption has not destroyed the message; it has made the key necessary to identify it.

Now spoil the key generation. Suppose shift 0 is selected 70% of the time, with each other shift selected 10% of the time. With equally likely directions, intercepted 1 now makes EAST much more likely than the alternatives. The operation is unchanged, but the security property has gone.

This is a small, precise version of the lesson from the cillies: an impressive range of available settings is not the same thing as an unpredictable choice among them.

The one-time pad: perfection. With a supplies problem

The four-direction system is a small additive one-time cipher. For longer messages, the same principle becomes the one-time pad: combine each piece of information with a fresh, independent random piece of key material, which the recipient also possesses secretly.[4]

In a binary version, the combining operation is usually exclusive OR, abbreviated XOR. No knowledge of electronics is required to follow it. A key bit of 0 leaves a message bit unchanged. A key bit of 1 flips it: 0 becomes 1, and 1 becomes 0. Apply the same key bit again and the original returns.

Message bitSecret key bitCiphertext bit
000
011
101
110

If the key bit is independent and equally likely to be 0 or 1, either plaintext bit produces either ciphertext bit with equal probability. Repeat with fresh independent key bits, and the argument extends to a whole fixed-length message.

For any proposed plaintext of that length, exactly one pad would turn it into the observed ciphertext. All those pads were equally likely. Even a machine that tried every pad would find no ciphertext-based reason to prefer the real message over the alternatives.[4]

The requirements are strict. The pad must be uniformly random over its possible values, independent of the message, secret from the attacker, and long enough to cover the protected data. No portion may be used again to protect another message. A page from a novel is not such a pad. Nor is a long output from a predictable process.

‘One-time’ describes the lifetime of the key material, not how often the recipient opens the document. Keeping a pad safe for later decryption is different from using the same pad to encrypt a second plaintext.

Vernam’s 1926 paper belongs to the earlier engineering history of combining telegraph messages with key tapes. Shannon’s contribution was to formalise the security question and its conditions, not to make earlier inventors disappear.[1][5]

The inconvenient requirement is shared secret material at scale. A gigabyte of arbitrary binary data needs a gigabyte of fresh pad for this construction. Both ends must obtain matching copies securely, keep track of their use, and protect them. Randomness has to be generated, distributed, stored, and retired without convenient shortcuts.[4]

A correctly constructed and operated one-time-pad system does not become mathematically decipherable merely because the attacker obtains a faster computer, an AI system, or a quantum computer. Those tools may help find stolen pads, exploit devices, or improve guesses from outside information. They do not alter the independence established by the model.

Can the key simply be smaller?

The size requirement is not merely a defect of the pad’s design. It reflects a limit on the secrecy being requested.

Return to the four directions. For each observed ciphertext, perfect secrecy must leave all four directions possible with their proper probabilities. But a recipient with a particular key must obtain one definite direction. If there were only two possible keys, an observed ciphertext could decrypt to at most two directions. At least two possibilities would have been ruled out by interception alone.

That would already be information. No ingenious rearrangement of those two keys can avoid the counting problem.

For a finite, correctly decryptable system with perfect secrecy across its message space, there must be at least as many possible keys as possible messages. If every n-bit message is possible, there are 2n possibilities to accommodate. The ordinary n-bit uniform pad meets that bound.[4]

This does not mean that every human meaning requires a key as long as the prose used to express it. If sender and recipient have agreed that only APPROVED or REFUSED can be sent, the decision can be encoded as one bit before encryption. A public codebook can reduce the representation without itself providing secrecy.

Similarly, lossless compression can reduce the amount of data that needs protection. But variable output lengths may reveal something, and the probabilities and format still matter. The theorem has not been beaten by making a file smaller; the object being encrypted has changed.[1][15]

Reuse the pad and the theorem exits stage left…

…pursued by a (likely Russian) bear.

Consider two messages protected with exactly the same binary pad. Each individual ciphertext may look reassuringly random. The pair is another matter.

XOR the two ciphertexts together and the repeated pad cancels. What remains is the XOR relationship between the two plaintexts. The attacker may not immediately recover either complete message, but the promised independence has disappeared. Known text in one can expose the corresponding text in the other.[4]

The four-direction version shows the same danger without binary notation. If two messages use the same secret shift, equal ciphertext numbers mean equal plaintext directions. Different ciphertext numbers reveal the relative displacement between the directions. One known direction reveals the shift and allows the other to be decoded.

That is already a substantial leak. A promise of secrecy does not survive merely because the attacker still has some work to do.

The VENONA material provides a historical example of the consequences of reused pad material. NSA histories describe duplicate key pages and the sustained analytical work needed to exploit messages encrypted in depth. This was not a refutation of perfect secrecy. It was exploitation of a system that had violated a condition necessary to support it.[8]

There is an important engineering distinction here. Reusing an independent pad bit is forbidden by the pad’s design. Reusing a modern cipher key may be expressly supported by its design, provided its nonce, usage-limit, and protocol requirements are followed. A slogan about never reusing any key would be both impractical and wrong.

Why ordinary language helps the attacker

Messages are rarely selected uniformly from every possible string. Language has structure. So do spreadsheets, image formats, protocol headers, and routine reports.

Consider a simple letter-substitution cipher: every A becomes one other letter, every B another, and so on. If a ciphertext word repeats its first letter in its third position, the plaintext does too. Even without knowing the alphabet substitution, the observer has learned a relationship within the word. More text can provide frequencies, repeated phrases, and increasingly strong constraints.[4]

This is redundancy in its information-theoretic sense: the predictability imposed by the source’s structure. It does not mean the prose is badly written or that the message contains unnecessary business information. A sentence can be concise and still obey spelling, grammar, and familiar patterns.[2][6]

A reader can often complete ‘Please close the d…’ because the surrounding text constrains what is likely. A cryptanalyst exploits comparable constraints, often with more patience and fewer assumptions about the sender’s literary ambitions.

In 1951 Shannon investigated the entropy of printed English by asking people to predict successive letters using preceding text. Human readers became measuring instruments for the language knowledge that simple letter counts miss. The estimates depended on how much context was available and on the kind of text; they were not timeless constants attached to the English language.[6]

There is a pleasing reversal here. The same predictability that helps a recipient repair a garbled message can help an adversary recognise a plausible decryption. Helpful redundancy and dangerous redundancy are not different substances. Their value depends on who can exploit them.

The result is not a recommendation to remove all structure from everything. Modern encryption should protect highly structured messages too. It is an explanation of why historical ciphers leaked more than their designers could see, and why a believable-looking output is not by itself proof of a successful decryption.

Unicity: enough evidence is not enough computing

With a short intercepted message, several keys may produce plausible plaintexts. As more material under the relevant key becomes available, language and format constraints can eliminate alternatives. Eventually there may be essentially one credible solution.

Shannon studied that transition through unicity distance: roughly, how much intercepted material is needed before ambiguity between plausible solutions largely disappears. It is an information question, not a stopwatch.[1]

A useful analogy is a crossword. Early clues may allow several completions. Enough intersecting clues can determine one solution, yet finding it may still require considerable work. The evidence can be sufficient before the solver becomes capable of using it.

Conversely, insufficient evidence can leave several valid completions no matter how long the solver works. A faster pencil does not invent the missing clue.

Shannon’s familiar estimate compares uncertainty in the key with the rate at which source redundancy supplies constraints. In a simplified random-cipher model it is often written as key entropy divided by redundancy per character. This is an estimate with assumptions, not a universal threshold at which all encryption fails. Hellman’s later analysis sharpened the model and its limitations.[1][9]

For scale, a uniformly chosen substitution alphabet over 26 letters has about 88.4 bits of key uncertainty: there are 26 choices for the first replacement, then 25 for the next, and so on, multiplied together. Mathematicians write that product as 26!. With an illustrative redundancy of three bits per character, the rough quotient is about 29.5 characters. Change the language model, key distribution, or cipher, and the result changes. It does not promise that every thirty-letter cryptogram has one easy answer.

Nor should that arithmetic be transferred casually to a modern cipher. A key or plaintext can be determined in principle by available data while remaining computationally inaccessible. Equally, a system can leave multiple possible complete plaintexts while revealing damaging partial information.

Unique does not mean easy. Ambiguous does not mean secret.

Make the evidence infeasible to exploit

If perfect secrecy consumes so much shared randomness, why does ordinary encrypted communication work with comparatively short keys?

Because its target is usually computational security, not perfect secrecy against unlimited computation. The aim is to prevent feasible adversaries from obtaining a useful advantage, under explicit assumptions about the algorithms, resources, and system.[4][10]

A well-designed cryptographic generator can expand a short secret seed into a long stream that a feasible observer should not be able to distinguish usefully from fresh randomness. But a deterministic expansion does not create unlimited independent entropy. Its outputs are restricted by the seed. Security rests on the difficulty of recognising or exploiting that restriction, not on pretending it does not exist.[4]

AES similarly uses a finite key to select a transformation. The standard specifies 128-bit blocks and keys of 128, 192, or 256 bits. Proper modes and protocols allow such building blocks to protect messages, while adding requirements beyond the bare block transformation. Those key lengths are not claims that arbitrarily long messages enjoy one-time-pad perfect secrecy.[11]

The intended advantage is asymmetry of work. The recipient with the key can recover the message efficiently. An attacker without it should face an infeasible problem. Known message formats need not ruin a modern cipher; resistance to known-plaintext and stronger attacks is part of what designers must establish.[4]

Shannon already separated theoretical ambiguity from the practical labour of solving a cipher. Later cryptography developed more precise computational security definitions. Goldwasser and Micali’s work on probabilistic encryption, published in journal form in 1984, formalised a powerful version of the aim: an efficient attacker should not gain useful information about the message merely by receiving its encryption, subject to the construction’s computational assumptions.[1][10]

That is the modern descendant of the question, not a claim that Shannon wrote today’s complete security definitions in 1949.

Confusion, diffusion, and some serious… pastry

Shannon’s paper is not solely a demonstration of an expensive ideal. It also asks how to make practical cryptanalysis difficult. Two enduring design ideas are confusion and diffusion.[1]

Diffusion spreads the consequences of the message’s structure across combinations of ciphertext elements. In a simple substitution, a common plaintext letter stays a common ciphertext letter. A better mixing process makes simple counts much less informative, so useful relationships become distributed across more complicated combinations.

Confusion makes the connection between observable ciphertext properties and the secret key difficult to exploit. A statistic may constrain the key without providing a conveniently solvable equation for one small part of it. The intended difficulty is for the analyst, not for the person reading the product manual.

The ideas are related but distinct: spread out the evidence, and make its relationship to the key hard to untangle. The design still needs to be reversible for the authorised recipient. Encryption is not a blender whose security rests on having destroyed the original; if that worked life would be too easy.

More layers do not automatically mean more security. Two consecutive Caesar shifts are just one Caesar shift with a different total displacement. Likewise, composing ordinary substitution alphabets produces another substitution alphabet. An enormous-looking procedure can remain inside the same weak family.

Good composition has to change what the adversary can exploit. Complexity is not a security property merely because it was difficult to implement.

There is another important qualification: for a fixed key, a reversible transformation rearranges possible inputs; it does not manufacture new uncertainty about the input. The secrecy comes from what the observer does not know, and computational protection from the difficulty of exploiting what they do. Confusion and diffusion do not suspend those accounting rules.

Compression helps. Except when it doesn’t

If redundancy helps recognise a plausible message, removing redundant representation before encryption can reduce the amount of material exposed to that kind of analysis. Shannon discusses this connection. It is tempting to turn it into a universal instruction: compress, then encrypt, and security must improve.[1]

That conclusion is too quick. Compression also creates an observable result: the compressed length. Two equal-length inputs can compress to different sizes because one contains more predictable structure. If those sizes remain visible after encryption, they can reveal something about the plaintext.

John Kelsey’s 2002 work demonstrated how apparently small compression leaks can become useful, including when an attacker can influence neighbouring input. The information need not come from recovering an encryption key. It can come from measuring how the system responds to guesses.[15]

TLS 1.3 separates these issues in a practical protocol. It removes the older TLS-level compression mechanism, uses authenticated encryption, and permits record padding; it also explicitly warns that traffic lengths are not automatically hidden. Application-level behaviour still needs its own analysis.[12]

What perfect secrecy does not promise

Suppose our one-time pad is ideal. The ciphertext reveals no information about the message. Does that make the communication safe?

Not by itself.

It does not authenticate the sender. A receiver needs a separate reason to accept the message as coming from the intended party. Possession of some bytes that decrypt is not a general proof of origin.

It does not prevent tampering. In the binary pad, flipping a ciphertext bit flips the corresponding decrypted plaintext bit. An attacker can change something without first learning what it was. In the four-direction example, adding one position to the ciphertext rotates the recipient’s recovered direction by one position. Confidentiality can survive while integrity fails.[4]

It does not prevent replay. A previously valid message may be copied and presented again unless the protocol includes suitable freshness checks. Yesterday’s authentic instruction need not be an instruction to perform the action again today.[4]

It does not ensure delivery. Someone can discard the ciphertext, cut the cable, or make the endpoint unavailable. A secrecy theorem does not keep the electricity on.

It does not necessarily hide metadata. Who communicates with whom, when, how often, and with what message sizes may remain visible. If the only possible plaintexts are differently sized, preserving their lengths may itself identify them. The elementary perfect-secrecy example fixes the message space and length; a system claiming more has to account for those observations too.[12]

It does not secure the endpoints. Malware, a stolen pad, an exposed screen, a plaintext log, or an unintended physical signal can provide information outside the ciphertext channel. That does not disprove the theorem. It demonstrates that the attacker was given more than the theorem assumed.[16]

The difference between secrecy and total security is not an academic nuisance. It explains why a correct cryptographic component can sit inside a failing service. The service has other promises to keep.

The assumptions are part of the result

A theorem is not weakened by having assumptions. It is weakened in use when somebody removes them from the glossy brochure, or otherwise hides or loses them.

The shared-key model assumes a secret available to the legitimate parties and unavailable to the listener. The perfect-secrecy example assumes a particular message space, fresh independent randomness, and an observation limited to the ciphertext. Correct decryption is required. These are not decorative clauses.

Later work explored different models. Diffie and Hellman’s 1976 public paper addressed communication without an already shared secret key and described a computational approach to key agreement. That changes the key-establishment problem, not the arithmetic of Shannon’s shared-secret bound. Authentication of the exchange remains a separate issue.[13]

Aaron Wyner’s 1975 wire-tap channel considered a different advantage: an eavesdropper whose observation passes through an additional noisy channel. Under suitable channel conditions, reliable communication with information-theoretic secrecy becomes possible as code length grows, without the same pre-shared-pad arrangement. The extra resource is a difference between what the recipient and the eavesdropper can observe. Again, the model changes; a theorem is not contradicted.[14]

The practical lesson is not that one of these models is the only respectable one. It is that a security claim must say which advantage is being used: secret material, computational difficulty, a physical observation advantage, or some combination.

Clear assumptions, clear rules, and a clear understanding of what security claims remain valid within them. And what happens to those claims if the rules or assumptions cease to be valid.

That is also the bridge to the next artefact, TEMPEST. A cipher may be perfect against someone who sees its intended output while the equipment leaks plaintext through another physical path. A declassified NSA history describes precisely why compromising emanations matter to cryptographic equipment. The attack can bypass the mathematical channel rather than defeat its encryption; physics can trump maths.[16]

The proof may be correct. The diagram may be missing a wire.

What to ask when someone says ‘encrypted’

For a reader who never intends to design a cipher, the most useful outcome is a better set of questions.

What is protected: the contents, the identities, the timing, or only one of those?

Who holds the keys, and can the service provider also read the plaintext?

How are keys generated, distributed, and protected?

Is the message authenticated as well as encrypted?

What happens when randomness repeats, a device is stolen, or the communication is interrupted?

What observations are outside the security claim?

Those questions do not require memorising an entropy formula. They require recognising that ‘encrypted’ names a mechanism, not a complete account of who can learn what.

For technical readers, the corresponding discipline is to state the adversary, the observations, the correctness requirement, the key and message distributions, and the security property. Distinguish information-theoretic independence from computational indistinguishability.

Do not replace an argument with a randomness plot, a key length, or the assertion that nobody has complained yet.

And do not mistake a theorem’s silence about a threat for a promise that the threat does not exist.

What the artefact really is

The physical artefact is the 1949 paper. It is less photogenic than a cipher machine, although considerably easier to carry.

The more important artefact is actually the change in question.

Instead of asking whether a message looks scrambled, ask what an observer can infer. Instead of counting settings and declaring victory, ask how the key was chosen. Instead of saying that a cipher is unbreakable, distinguish absent evidence from inaccessible evidence. Instead of protecting an algorithm in isolation, identify everything the adversary can see.

Shannon did not make secrecy simple. He made the claims separable enough to examine. He made it possible to say much more clearly which claims stand, and under what circumstances.

The one-time pad shows that perfect content secrecy is possible. Its demanding assumptions show why operating it is a different achievement. Practical cryptography shows why weaker, carefully defined claims can still support extraordinarily useful systems. The attacks around the edges show why no theorem can protect a channel it was never asked to model.

The message must remain recoverable to someone.

The question is what makes that person different from everyone listening.

Shannon made us account for the difference.

Purple Signature of Sophie Ada Mathison Violet Baskerville
References & Source notes

The historical argument centres on Shannon’s 1949 paper. Later sources are identified as later work, rather than being folded into claims about what he had already proved. Bracketed references in the article link to the numbered entries below.

The four-direction cipher, biased-key calculation, complement-only example, and illustrative unicity calculation are worked teaching examples prepared for this article. They are not proposed encryption products, nor reported measurements of existing systems.

Edition check: the predecessor report is dated 1 September 1945 in the original printed footnote on page 656. A commonly circulated re-typeset copy says 1946; the original-page facsimile is the authority used here.

Historical content and cited standards were reviewed for this draft on 16 & 17 September 2026. A historical publication is not presented as current implementation guidance.

[1] Claude E. Shannon. Communication Theory of Secrecy Systems

Bell System Technical Journal 28(4), October 1949, pp. 656–715.

The principal artefact. Original pagination: perfect secrecy, pp. 679–682; ideal systems, sections 17–18; practical secrecy, Part III; pastry and mixing, section 25, p. 712. The opening footnote dates the predecessor report to 1 September 1945.

Publisher record  |  Original-page facsimile  |  DOI

[2] Claude E. Shannon. A Mathematical Theory of Communication

Bell System Technical Journal 27, July and October 1948, pp. 379–423 and 623–656.

The earlier communication-theory paper: uncertainty, entropy, conditional entropy, source structure, and noisy channels. The article uses the resulting information identities in its own worked examples.

Full paper

[3] Nokia Bell Labs. Claude Shannon

Institutional historical profile.

Biographical and Bell Laboratories context. Used for that context, not as a substitute for the 1949 paper itself.

Institutional profile

[4] Dan Boneh and Victor Shoup. A Graduate Course in Applied Cryptography

Version 0.6, January 2023.

Chapters 2–3 for perfect and computational security, the one-time pad, key-space bounds, and stream ciphers; chapters 6 and 9 for authentication and authenticated encryption. Modern notation and explanations are distinguished from Shannon’s historical terminology.

Full textbook

[5] Gilbert S. Vernam. Cipher Printing Telegraph Systems for Secret Wire and Radio Telegraphic Communications

Journal of the American Institute of Electrical Engineers 45(2), February 1926, pp. 109–115.

Primary engineering account of telegraph encryption using key material. It predates Shannon’s mathematical treatment and is cited in the 1949 paper.

Original paper

[6] Claude E. Shannon. Prediction and Entropy of Printed English

Bell System Technical Journal 30(1), January 1951, pp. 50–64.

The letter-prediction experiments and the importance of context when estimating the redundancy of a language source. No single numerical estimate is treated here as a universal constant for English.

Full paper

[7] NIST. Recommendation for the Entropy Sources Used for Random Bit Generation

Special Publication 800–90B, January 2018.

Section 2.1 and Appendix D for min-entropy and its relation to the most likely output. The biased 256-bit generator in the article is an original numerical illustration, not a reported defective product.

Publication record  |  Full publication

[8] National Security Agency. VENONA: An Overview

Cryptologic Almanac 50th Anniversary Series, declassified historical account.

Duplicate one-time-pad material and the work needed to exploit it. The article uses the documented cryptographic failure, without treating the programme as a simple one-step decryption story.

NSA historical account

[9] Martin E. Hellman. An Extension of the Shannon Theory Approach to Cryptography

IEEE Transactions on Information Theory IT-23(3), May 1977, pp. 289–294.

The random-cipher model, unicity, spurious solutions, and the limits of inferring security merely from remaining ambiguity.

Author-hosted paper

[10] Shafi Goldwasser and Silvio Micali. Probabilistic Encryption

Journal of Computer and System Sciences 28(2), 1984, pp. 270–299.

A later formalisation of computationally protecting partial information. This is a development after Shannon, not a definition retrospectively attributed to him.

Full paper

[11] NIST. Advanced Encryption Standard (AES)

FIPS 197, originally 2001; editorial update 9 May 2023.

AES block and key sizes and its round transformations. Used as a concrete later example, not as a claim that Shannon designed AES.

Standard record  |  Full standard

[12] Eric Rescorla. The Transport Layer Security (TLS) Protocol Version 1.3

RFC 8446, August 2018.

Section 1 for security goals; removal of TLS-level compression; section 5.4 for record padding; Appendix E.3 for traffic analysis. TLS records do not automatically conceal all lengths or application behaviour.

RFC 8446

[13] Whitfield Diffie and Martin E. Hellman. New Directions in Cryptography

IEEE Transactions on Information Theory IT-22(6), November 1976, pp. 644–654.

The public proposal for computational key agreement and public-key cryptography. Changing how a key is established does not refute a theorem about a different model.

Author-hosted paper

[14] Aaron D. Wyner. The Wire-Tap Channel

Bell System Technical Journal 54(8), October 1975, pp. 1355–1387.

A noisy-channel observation advantage and asymptotic information-theoretic secrecy. This differs from perfect secrecy in the elementary finite, pre-shared-key model.

Publisher record  |  Full paper

[15] John Kelsey. Compression and Information Leakage of Plaintext

Fast Software Encryption 2002, LNCS 2365, pp. 263–276.

Plaintext information revealed by compressed length, including the additional leverage from attacker-influenced input. Used to qualify a simplistic compress-then-encrypt rule.

Full paper

[16] National Security Agency. TEMPEST: A Signal Problem

Cryptologic Spectrum 2(3), Summer 1972; subsequently declassified.

Public historical evidence of compromising emanations from cryptographic equipment. The bridge to artefact 024 uses this public source, not non-public operational guidance.

Declassified article

[17] NIST. A Statistical Test Suite for Random and Pseudorandom Number Generators for Cryptographic Applications

Special Publication 800–22, Revision 1a, April 2010.

The explicit warning that statistical testing cannot substitute for cryptanalysis. A satisfactory-looking output distribution is not by itself a confidentiality argument.

Publication record

[18] Yoav Nir and Adam Langley. ChaCha20 and Poly1305 for IETF Protocols

RFC 8439, June 2018.

Sections 2.3 and 4 for the stream construction and nonce requirements. Repeating key, nonce, and counter position repeats keystream; this is not a universal description of every cipher’s nonce behaviour.

RFC 8439

[19] KGB Headquarters Moscow to the London KGB Residency. Permanent Operational Assignment to Uncover NATO Preparations for a Nuclear Missile Attack on the USSR

17 February 1983, Top Secret; subsequently published by Oleg Gordievsky and Christopher Andrew and reproduced by the National Security Archive.

Primary Operation RYAN instruction requiring the London residency to establish normal patterns of activity at important government institutions and headquarters, including numbers of illuminated windows during and outside working hours, and to report significant departures from those patterns. Benjamin B. Fischer’s CIA historical study A Cold War Conundrum: The 1983 Soviet War Scare provides the wider Operation RYAN and war-scare context.

National Security Archive: KGB Operation RYAN instruction
CIA historical study: A Cold War Conundrum: The 1983 Soviet War Scare

[20] R. Shirey. Internet Security Glossary, Version 2

RFC 4949, August 2007.

Defines “communications cover” as concealing or altering characteristic communications patterns so that they do not reveal useful information to an adversary.

RFC 4949: Internet Security Glossary, Version 2

[21] Stephen Kent. IP Encapsulating Security Payload (ESP)

RFC 4303, December 2005, sections 2.6–2.7.

Concrete protocol example of traffic-flow confidentiality: dummy packets may be inserted, including during otherwise silent periods, and traffic may be padded or shaped to conceal observable communication patterns.

RFC 4303: IP Encapsulating Security Payload (ESP)