---
title: 'Optional: Conditional Information and Data Processing'
url: https://doc.liz6.com/en/theory/01-information-theory/07-conditional-information-and-data-processing
locale: en
area: theory
tags:
- Theory
- Information theory
date: 2026-09-08
modified: 2026-09-08
description: 'Optional advanced reading: use conditional groups and merged states to understand why conditional mutual information can rise or fall, and why data processing requires the Markov condition.'
---

# Optional: Conditional Information and Data Processing

This follows [Entropy and Information Measures](./01-entropy-and-information-measures.md). Read it after understanding joint entropy, conditional entropy, and mutual information; it is not a prerequisite for Chapter 2. All variables are finite and discrete, and all information is in bits.

Two apparently conflicting ideas meet here: **adding a condition can strengthen dependence, while further processing cannot create information from nothing.** The distinction concerns which information becomes available and which variables we compare.

## Conditional mutual information: compare within contexts, then average

How much do $X$ and $Y$ share once $Z$ is known? Fix $Z=z$, compute mutual information within that conditional distribution, then average over contexts:

$$I(X;Y\mid Z)=\sum_zp(z)I(X;Y\mid Z=z).$$

Equivalently,

$$I(X;Y\mid Z)=H(Y\mid Z)-H(Y\mid X,Z).$$

The first term is uncertainty with the context already known; the second is what remains after also observing $X$. Conditional mutual information is nonnegative, but **it has no fixed ordering relative to the unconditional $I(X;Y)$**.

**Given Z, dependence can appear or disappear**

Keep four X,Y combinations fixed and filter by Z, renormalizing probabilities. XOR turns independence into perfect conditional dependence; a shared source removes dependence once known.


### XOR: independent overall, perfectly dependent within groups

Take independent fair bits $X,Y$ and set $Z=X\oplus Y$. Without $Z$, all four pairs have probability $1/4$, so $I(X;Y)=0$.

Given $Z=0$, only $(0,0)$ and $(1,1)$ remain. Given $Z=1$, only $(0,1)$ and $(1,0)$ remain. Within either group, $Y$ alone still has 1-bit entropy, but $X$ determines $Y$. Each group's mutual information is 1 bit, as is their average.

This does not violate the average reduction of entropy under conditioning: $H(Y\mid Z)=H(Y)=1$, while $H(Y\mid X,Z)=0$. Extra information $Z$ changes what can be predicted jointly with $X$; it does not alter the original samples.

### A shared source: context can remove additional dependence

Instead let $Z$ be fair and set $X=Y=Z$. Overall, $X$ and $Y$ are identical and $I(X;Y)=1$. Once $Z$ is given, both are determined, so $I(X;Y\mid Z)=0$.

These examples show conditional mutual information both above and below ordinary mutual information. Adding conditions does not always weaken dependence. Nor can their difference be treated as a necessarily nonnegative third overlap in a Venn diagram.

## Data processing: the bound requires using only the observation

Suppose $X\to Y\to Z$ is a Markov chain:

$$p(x,y,z)=p(x,y)p(z\mid y).$$

Producing $Z$ uses only $Y$; once $Y$ is known, it receives no additional information from $X$. The **data processing inequality** states

$$I(X;Z)\le I(X;Y).$$

Information need not strictly decrease: an invertible transformation can preserve all of it. The result also does not mean a processed representation is less useful. Processing can make existing information easier for a restricted predictor to use; task accuracy and total mutual information are different quantities.

**Where does information lost in processing go?**

Four equally likely X values are reversibly relabeled as Y. Keep Y, retain parity, or merge everything. Lines show indistinguishable original states; a strip separates retained and lost information.


Choose an observed $Z$ and try to reconstruct $X$. Reversible processing leaves one candidate, with success 1; parity leaves two equiprobable candidates, with best success $1/2$; a constant leaves four, with success $1/4$. Relabeling the result as $W=3-Z$ leaves the candidate set unchanged: reading only the result cannot separate merged originals. These success rates belong to this uniform deterministic-grouping example; mutual information alone does not give such a general formula.

Here $Y=(X+1)\bmod4$ reversibly relabels $X$ and preserves all 2 bits. Keeping only Y's parity merges the four original states into two groups, leaving 1 bit. Mapping every value to a constant leaves zero. The lost information is the distinction between states that can no longer be told apart.

### The gap is exactly a conditional mutual information

Expand the same $I(X;Y,Z)$ in two orders:

$$\begin{aligned}
I(X;Y,Z)&=I(X;Y)+I(X;Z\mid Y),\\
&=I(X;Z)+I(X;Y\mid Z).
\end{aligned}$$

The Markov condition gives $I(X;Z\mid Y)=0$, hence

$$I(X;Y)-I(X;Z)=I(X;Y\mid Z)\ge0.$$

The right-hand side asks: **after observing processed output $Z$, how much more can the old observation $Y$ tell us about $X$?** This is the information processing discarded, shown in gold. Equality holds exactly when that additional information is zero. Invertibility is sufficient for equality, but not necessary.

## Why doesn't XOR violate data processing?

In the XOR example, $Z$ is computed from both $X$ and $Y$. It is not generated only from $Y$, so that Markov chain and its bound do not apply.

Likewise, if a processing step also accesses original data, relevant retrieved material, or another source of information about $X$, include those inputs in the observation before checking the Markov relation. Omitting additional inputs can make a valid computation appear to violate the inequality.

| Question | Correct comparison |
|---|---|
| How much dependence remains with extra context? | $I(X;Y\mid Z)$ can exceed or fall below $I(X;Y)$ |
| How much survives processing only an existing observation? | Under $X\to Y\to Z$, $I(X;Z)\le I(X;Y)$ |
| How much information was lost? | Under the same condition, the gap is $I(X;Y\mid Z)$ |

For representation learning, see [the Information Bottleneck](./05-rate-distortion-and-information-bottleneck.md). Densities and differential entropy are introduced separately with [Gaussian Channels](./04-channel-capacity-and-coding.md).

## References

- [MIT 6.441, Chapter 2: Mutual Information](https://ocw.mit.edu/courses/6-441-information-theory-spring-2016/184197ca5d5418da2415d37e929860b9_MIT6_441S16_chapter_2.pdf).
- Cover & Thomas, *Elements of Information Theory*, Chapter 2.
