Optional: Conditional Information and Data Processing

On this page

This follows Entropy and Information Measures. Read it after understanding joint entropy, conditional entropy, and mutual information; it is not a prerequisite for Chapter 2. All variables are finite and discrete, and all information is in bits.

Two apparently conflicting ideas meet here: adding a condition can strengthen dependence, while further processing cannot create information from nothing. The distinction concerns which information becomes available and which variables we compare.

Conditional mutual information: compare within contexts, then average

How much do and share once is known? Fix , compute mutual information within that conditional distribution, then average over contexts:

Equivalently,

The first term is uncertainty with the context already known; the second is what remains after also observing . Conditional mutual information is nonnegative, but it has no fixed ordering relative to the unconditional .

Preparing the visual
Given Z, dependence can appear or disappear

Keep four X,Y combinations fixed and filter by Z, renormalizing probabilities. XOR turns independence into perfect conditional dependence; a shared source removes dependence once known.

XOR: independent overall, perfectly dependent within groups

Take independent fair bits and set . Without , all four pairs have probability , so .

Given , only and remain. Given , only and remain. Within either group, alone still has 1-bit entropy, but determines . Each group's mutual information is 1 bit, as is their average.

This does not violate the average reduction of entropy under conditioning: , while . Extra information changes what can be predicted jointly with ; it does not alter the original samples.

A shared source: context can remove additional dependence

Instead let be fair and set . Overall, and are identical and . Once is given, both are determined, so .

These examples show conditional mutual information both above and below ordinary mutual information. Adding conditions does not always weaken dependence. Nor can their difference be treated as a necessarily nonnegative third overlap in a Venn diagram.

Data processing: the bound requires using only the observation

Suppose is a Markov chain:

Producing uses only ; once is known, it receives no additional information from . The data processing inequality states

Information need not strictly decrease: an invertible transformation can preserve all of it. The result also does not mean a processed representation is less useful. Processing can make existing information easier for a restricted predictor to use; task accuracy and total mutual information are different quantities.

Preparing the visual
Where does information lost in processing go?

Four equally likely X values are reversibly relabeled as Y. Keep Y, retain parity, or merge everything. Lines show indistinguishable original states; a strip separates retained and lost information.

Choose an observed and try to reconstruct . Reversible processing leaves one candidate, with success 1; parity leaves two equiprobable candidates, with best success ; a constant leaves four, with success . Relabeling the result as leaves the candidate set unchanged: reading only the result cannot separate merged originals. These success rates belong to this uniform deterministic-grouping example; mutual information alone does not give such a general formula.

Here reversibly relabels and preserves all 2 bits. Keeping only Y's parity merges the four original states into two groups, leaving 1 bit. Mapping every value to a constant leaves zero. The lost information is the distinction between states that can no longer be told apart.

The gap is exactly a conditional mutual information

Expand the same in two orders:

The Markov condition gives , hence

The right-hand side asks: after observing processed output , how much more can the old observation tell us about ? This is the information processing discarded, shown in gold. Equality holds exactly when that additional information is zero. Invertibility is sufficient for equality, but not necessary.

Why doesn't XOR violate data processing?

In the XOR example, is computed from both and . It is not generated only from , so that Markov chain and its bound do not apply.

Likewise, if a processing step also accesses original data, relevant retrieved material, or another source of information about , include those inputs in the observation before checking the Markov relation. Omitting additional inputs can make a valid computation appear to violate the inequality.

QuestionCorrect comparison
How much dependence remains with extra context? can exceed or fall below
How much survives processing only an existing observation?Under ,
How much information was lost?Under the same condition, the gap is

For representation learning, see the Information Bottleneck. Densities and differential entropy are introduced separately with Gaussian Channels.

References