Quantum states and channels

Density matrices for mixed states, quantum channels and their representations, general measurements, purifications and fidelity.

Density matrices

So far, we have described quantum systems with state vectors: ∣ψ⟩\lvert\psi\rangle for a qubit, and longer vectors for larger registers. But a state vector describes one definite state. What if we only have a probability distribution over states?

This happens in three common situations:

  • Randomness. A source might produce ∣0⟩\lvert 0\rangle or ∣+⟩\lvert +\rangle at random. This is not a superposition of the two—it is a classical random choice between them.
  • Noise. A real quantum device does not always produce the intended state. After a noisy operation, the system may be in different states with different probabilities.
  • Part of a compound system. If two qubits are entangled, the pair has a state vector, but an individual qubit generally does not have one of its own.

These look like different problems, but they have the same solution: instead of describing one state, we need an object that can describe a mixture of possible states.

That object is the density matrix.

It includes ordinary state vectors as a special case, while also describing classical randomness, noise, and the parts of entangled systems. It gives us one framework for describing quantum states, applying operations, combining systems, and predicting measurements.

Definition of density matrices

Suppose that XX is a system and Σ\Sigma is its classical state set. A density matrix describing a state of XX is a matrix with complex-number entries whose rows and columns correspond to the elements of Σ\Sigma. We typically write density matrices as ρ\rho, σ\sigma, or ξ\xi.

The purpose of ρ\rho is to let us calculate probabilities of measurement outcomes. If ∣x⟩\lvert x\rangle is a classical state from Σ\Sigma, then the probability of obtaining that state when measuring in the classical basis is given by the corresponding diagonal entry Pr⁡(x)=⟨x∣ρ∣x⟩=ρx,x\Pr(x)=\langle x\rvert\rho\lvert x\rangle=\rho_{x,x}. Thus, the diagonal entries of ρ\rho give the probabilities of the classical states.

But we are not limited to measurements in the classical basis. For any unit vector ∣ψ⟩\lvert\psi\rangle, we can consider a measurement that asks whether the system is in the state ∣ψ⟩\lvert\psi\rangle. The probability of a “yes” outcome is

Pr⁡(ψ)=⟨ψ∣ρ∣ψ⟩.\Pr(\psi)=\langle\psi\rvert\rho\lvert\psi\rangle.

This is the same rule we already know for state vectors. If the system is in the state ∣ϕ⟩\lvert\phi\rangle, its density matrix is ρ=∣ϕ⟩⟨ϕ∣\rho=\lvert\phi\rangle\langle\phi\rvert. Substituting this into the measurement rule gives ⟨ψ∣ρ∣ψ⟩=⟨ψ∣ϕ⟩⟨ϕ∣ψ⟩=∣⟨ψ∣ϕ⟩∣2\langle\psi\rvert\rho\lvert\psi\rangle=\langle\psi|\phi\rangle\langle\phi|\psi\rangle=\lvert\langle\psi|\phi\rangle\rvert^{2}, which is exactly the usual measurement probability for the state ∣ϕ⟩\lvert\phi\rangle.

So ⟨ψ∣ρ∣ψ⟩\langle\psi\rvert\rho\lvert\psi\rangle is simply the density-matrix version of the familiar probability formula. The bra ⟨ψ∣\langle\psi\rvert and ket ∣ψ⟩\lvert\psi\rangle select the part of ρ\rho relevant to the question “is the state ∣ψ⟩\lvert\psi\rangle?” and reduce it to the single number that gives the probability of a “yes” outcome.

What is the probability that a measurement finds the system in the state ∣ψ⟩\lvert\psi\rangle?

1100
⟨ψ∣\langle\psi\rvert the question
12\tfrac1214\tfrac1414\tfrac1412\tfrac12
ρ\rho the state
1100
∣ψ⟩\lvert\psi\rangle the same question
=
12\tfrac12
Pr⁡(ψ)\Pr(\psi) the chance of a “yes”
⟨0∣ρ∣0⟩=ρ0,0=12\langle0\rvert\rho\lvert0\rangle=\rho_{0,0}=\tfrac12

The off-diagonal entries encode coherence between the corresponding classical states. They affect measurements in superposition bases, as the demo shows.

Together, the entries must produce valid measurement probabilities: the probabilities of all possible outcomes must add up to one, and none can be negative. These requirements give us two mathematical conditions on ρ\rho: it must have trace one, and it must be positive semidefinite.

1. Unit trace

The trace of a square matrix is the sum of its diagonal entries.

Tr(A0,0A0,1⋯A0,n−1A1,0A1,1⋯A1,n−1⋮⋮⋱⋮An−1,0An−1,1⋯An−1,n−1)=A0,0+A1,1+⋯+An−1,n−1\mathrm{Tr}\begin{pmatrix}\textcolor{#7c3aed}{A_{0,0}}&A_{0,1}&\cdots&A_{0,n-1}\\A_{1,0}&\textcolor{#7c3aed}{A_{1,1}}&\cdots&A_{1,n-1}\\\vdots&\vdots&\textcolor{#7c3aed}{\ddots}&\vdots\\A_{n-1,0}&A_{n-1,1}&\cdots&\textcolor{#7c3aed}{A_{n-1,n-1}}\end{pmatrix}=\textcolor{#7c3aed}{A_{0,0}}+\textcolor{#7c3aed}{A_{1,1}}+\cdots+\textcolor{#7c3aed}{A_{n-1,n-1}}

It reads only the diagonal and ignores all off-diagonal entries. The trace is also a linear function, meaning that Tr(αA+βB)=α Tr(A)+β Tr(B)\mathrm{Tr}(\alpha A+\beta B)=\alpha\,\mathrm{Tr}(A)+\beta\,\mathrm{Tr}(B).

For a density matrix, the diagonal entries are the probabilities of the classical states. Therefore, Tr(ρ)=1\mathrm{Tr}(\rho)=1 says exactly that those probabilities add up to one: Pr⁡(0)+Pr⁡(1)+⋯+Pr⁡(n−1)=1\Pr(0)+\Pr(1)+\cdots+\Pr(n-1)=1.

2. Positive semidefinite

Unit trace guarantees that the probabilities on the diagonal add up to one, but those are only the probabilities for measurements in the classical basis, and a quantum system can be measured in other directions too. For every unit vector ∣ψ⟩\lvert\psi\rangle the quantity ⟨ψ∣ρ∣ψ⟩\langle\psi\rvert\rho\lvert\psi\rangle is a measurement probability, so it must never come out negative.

⟨ψ∣ρ∣ψ⟩≥0for every ∣ψ⟩.\langle\psi\rvert\rho\lvert\psi\rangle\geq 0\quad\text{for every }\lvert\psi\rangle.

A matrix satisfying this is called positive semidefinite, written ρ≥0\rho\geq 0.

Notice that positive semidefinite does not mean that every entry of ρ\rho is nonnegative: the off-diagonal entries can be negative or complex, as the shading above already showed. What must be nonnegative is the number ⟨ψ∣ρ∣ψ⟩\langle\psi\rvert\rho\lvert\psi\rangle for every possible ∣ψ⟩\lvert\psi\rangle. There are several equivalent ways to recognize this property, describing the same requirement from different viewpoints:

  • ⟨ψ∣ρ∣ψ⟩≥0\langle\psi|\rho|\psi\rangle\geq 0 for every complex vector ∣ψ⟩\lvert\psi\rangle.This is the physical formulation: every measurement probability must be nonnegative.
  • ρ\rho is Hermitian, meaning that it equals its ρ=ρ†\rho=\rho^\dagger, and all its eigenvalues are nonnegative.The second formulation is usually the most useful for calculations.An eigenvalue λ\lambda of ρ\rho is a number for which there exists a nonzero vector ∣v⟩\lvert v\rangle satisfying ρ∣v⟩=λ∣v⟩\rho\lvert v\rangle=\lambda\lvert v\rangle. If we choose ∣v⟩\lvert v\rangle to be normalized, multiplying the equation on the left by ⟨v∣\langle v\rvert gives ⟨v∣ρ∣v⟩=λ\langle v\rvert\rho\lvert v\rangle=\lambda.Eigenvalues are not just some of the possible values of ⟨ψ∣ρ∣ψ⟩\langle\psi\rvert\rho\lvert\psi\rangle: for a Hermitian matrix, they determine its extreme values. The smallest eigenvalue is the minimum possible value of ⟨ψ∣ρ∣ψ⟩\langle\psi\rvert\rho\lvert\psi\rangle, and the largest eigenvalue is the maximum possible value. Therefore, ⟨ψ∣ρ∣ψ⟩\langle\psi\rvert\rho\lvert\psi\rangle can never be negative exactly when the smallest eigenvalue is nonnegative, which is equivalent to all eigenvalues being nonnegative.Hermiticity also guarantees that the eigenvalues are real. For a Hermitian matrix, ⟨v∣ρ∣v⟩\langle v\rvert\rho\lvert v\rangle is always real, and for a normalized eigenvector ⟨v∣ρ∣v⟩=λ\langle v\rvert\rho\lvert v\rangle=\lambda. Therefore, every eigenvalue λ\lambda is real.
  • There exists a matrix MM such that ρ=M†M\rho=M^\dagger M.The third formulation makes nonnegativity especially transparent: ⟨ψ∣ρ∣ψ⟩=⟨ψ∣M†M∣ψ⟩=∥M∣ψ⟩∥2≥0\langle\psi\rvert\rho\lvert\psi\rangle=\langle\psi\rvert M^\dagger M\lvert\psi\rangle=\bigl\lVert M\lvert\psi\rangle\bigr\rVert^{2}\geq 0, because a squared length cannot be negative.

These are not three separate conditions. They are three equivalent ways of expressing the same requirement: ρ\rho must never predict a negative probability.

Testing positivity for a qubit

For a 2×22\times2 density matrix, positivity can be checked particularly simply using its eigenvalues. First, however, we must check that ρ\rho is Hermitian, meaning ρ=ρ†\rho=\rho^\dagger. This guarantees that its eigenvalues are real.

The eigenvalues are the roots of the characteristic equation det⁡(ρ−λI)=0\det(\rho-\lambda I)=0, which for a 2×22\times2 matrix becomes λ2−Tr(ρ) λ+det⁡(ρ)=0\lambda^{2}-\mathrm{Tr}(\rho)\,\lambda+\det(\rho)=0. Therefore, the two eigenvalues satisfy λ1+λ2=Tr(ρ)=1\lambda_{1}+\lambda_{2}=\mathrm{Tr}(\rho)=1 and λ1λ2=det⁡(ρ)\lambda_{1}\lambda_{2}=\det(\rho).

Because the eigenvalues are real and their sum is positive, they cannot both be negative. Therefore, positivity can fail only when one eigenvalue is negative and the other is positive, which happens exactly when their product is negative. Hence, the eigenvalues are both nonnegative exactly when their product is nonnegative, so for a 2×22\times2 density matrix ρ≥0  ⟺  det⁡(ρ)≥0\rho\geq0\iff\det(\rho)\geq0, provided that ρ\rho is Hermitian and has trace 11.

Applying the two conditions of density matrices

Density matrices
Look-alikes
A=(1000)A=\begin{pmatrix}1&0\\0&0\end{pmatrix}|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
10
∣0⟩|0\rangle
∣1⟩|1\rangle
−∣0⟩-|0\rangle
−∣1⟩-|1\rangle
λ+=1.00\lambda_{+}=1.00
λ−=0.00\lambda_{-}=0.00
30°
⟨ψ∣A∣ψ⟩\langle\psi|A|\psi\rangle0.75
  1. Tr(A)=1+0=1\mathrm{Tr}(A)=1+0=1Holds
  2. M=(1000),  M†=(1000),  M†M=AM=\left(\begin{smallmatrix}1&0\\0&0\end{smallmatrix}\right),\ \ M^\dagger=\left(\begin{smallmatrix}1&0\\0&0\end{smallmatrix}\right),\ \ M^\dagger M=AHolds
    λ±=12±12=1, 0\lambda_{\pm}=\tfrac12\pm\tfrac12=1,\,0Holds
    min⁡θ ⟨ψ∣A∣ψ⟩=0.00\min_\theta\,\langle\psi|A|\psi\rangle=0.00Holds

Constructing positive semidefinite matrices

The third formulation gives a direct way to construct a positive semidefinite matrix. Start with any matrix MM and form A=M†MA=M^\dagger M. By construction, AA is positive semidefinite, whatever MM is. This does not ensure trace one, so to obtain a density matrix, normalize AA by its trace.

Size of M
M=M=
M†M=(17234+169i−6−71i34−169i26113−69i−6+71i13+69i231)M^\dagger M=\begin{pmatrix}172&34+169i&-6-71i\\34-169i&261&13-69i\\-6+71i&13+69i&231\end{pmatrix}
Tr(M†M)=664\mathrm{Tr}(M^\dagger M)=664
ρ=M†M664\rho=\dfrac{M^\dagger M}{664}

Connection to state vectors

Every state vector is already a density matrix in disguise. A quantum state vector ∣ψ⟩\lvert\psi\rangle is a column vector of Euclidean norm one, and the density matrix describing that same state is the column multiplied by its own conjugate transpose.

ρ=∣ψ⟩⟨ψ∣.\rho=\lvert\psi\rangle\langle\psi\rvert.

States represented by density matrices of this form are called pure states. Written out with the amplitudes of ∣ψ⟩\lvert\psi\rangle, the product puts αjαk‾\alpha_{j}\overline{\alpha_{k}} in row jj and column kk.

∣ψ⟩=(α0α1⋮αn−1)⟹∣ψ⟩⟨ψ∣=(∣α0∣2α0α1‾⋯α0αn−1‾α1α0‾∣α1∣2⋯α1αn−1‾⋮⋮⋱⋮αn−1α0‾αn−1α1‾⋯∣αn−1∣2)\lvert\psi\rangle=\begin{pmatrix}\alpha_{0}\\\alpha_{1}\\\vdots\\\alpha_{n-1}\end{pmatrix}\quad\Longrightarrow\quad\lvert\psi\rangle\langle\psi\rvert=\begin{pmatrix}\lvert\alpha_{0}\rvert^{2}&\alpha_{0}\overline{\alpha_{1}}&\cdots&\alpha_{0}\overline{\alpha_{n-1}}\\\alpha_{1}\overline{\alpha_{0}}&\lvert\alpha_{1}\rvert^{2}&\cdots&\alpha_{1}\overline{\alpha_{n-1}}\\\vdots&\vdots&\ddots&\vdots\\\alpha_{n-1}\overline{\alpha_{0}}&\alpha_{n-1}\overline{\alpha_{1}}&\cdots&\lvert\alpha_{n-1}\rvert^{2}\end{pmatrix}

The diagonal carries the squared amplitudes ∣αj∣2\lvert\alpha_{j}\rvert^{2}, which are exactly the probabilities of the classical states, and the off-diagonal entries carry the relative phases between them.

∣0⟩=(10)\lvert0\rangle=\begin{pmatrix}1\\0\end{pmatrix}
∣0⟩⟨0∣=(10)(10)=(1000)\lvert0\rangle\langle0\rvert=\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}=\begin{pmatrix}1&0\\0&0\end{pmatrix}

A pure state also satisfies both conditions of the definition automatically. Its trace is ∣α0∣2+⋯+∣αn−1∣2=1\lvert\alpha_{0}\rvert^{2}+\cdots+\lvert\alpha_{n-1}\rvert^{2}=1, which is what the norm of ∣ψ⟩\lvert\psi\rangle says, and it is positive semidefinite by the third formulation, taking M=⟨ψ∣M=\langle\psi\rvert.

Global phase disappears

Remember that a state vector carries a , but this phase has no physical meaning: multiplying ∣ψ⟩\lvert\psi\rangle by eiθe^{i\theta} does not change the physical state. Thus, ∣ψ⟩\lvert\psi\rangle and ∣ϕ⟩=eiθ∣ψ⟩\lvert\phi\rangle=e^{i\theta}\lvert\psi\rangle represent the same state.

For a pure state, the density matrix is ρ=∣ψ⟩⟨ψ∣\rho=\lvert\psi\rangle\langle\psi\rvert. If we use ∣ϕ⟩\lvert\phi\rangle instead, the global phase cancels.

∣ϕ⟩⟨ϕ∣=(eiθ∣ψ⟩)(eiθ∣ψ⟩)†=ei(θ−θ)∣ψ⟩⟨ψ∣=∣ψ⟩⟨ψ∣.\lvert\phi\rangle\langle\phi\rvert=\bigl(e^{i\theta}\lvert\psi\rangle\bigr)\bigl(e^{i\theta}\lvert\psi\rangle\bigr)^{\dagger}=e^{i(\theta-\theta)}\lvert\psi\rangle\langle\psi\rvert=\lvert\psi\rangle\langle\psi\rvert.

So the global phase carried by state vectors is simply absent from their density matrices. Two state vectors give the same density matrix exactly when they differ only by a global phase.

Not every state is pure

The density matrices that can be written as ∣ψ⟩⟨ψ∣\lvert\psi\rangle\langle\psi\rvert describe exactly the states that can already be described by a state vector. But not every density matrix has this form. The others capture randomness, noise, and subsystems of entangled systems—things that a state vector alone cannot describe.

Probabilistic mixtures

A source prepares a qubit in the state ∣0⟩\lvert0\rangle half of the time and in the state ∣+⟩\lvert+\rangle the other half, then hands it over without saying which one it prepared. Nothing about the qubit is undecided, but our description of it is: we hold one of two definite states and we do not know which.

No state vector says that. Writing ∣0⟩\lvert0\rangle or ∣+⟩\lvert+\rangle claims knowledge we do not have, and a superposition of the two is a third definite state, prepared by nobody. What we can still do is predict measurements, because we know the recipe the source followed.

Ask any measurement question. It is answered half the time by a qubit in the state ∣0⟩\lvert0\rangle and half the time by one in the state ∣+⟩\lvert+\rangle, so its probability is the average of the two answers. In general, for a source preparing ρ\rho with probability pp and σ\sigma with probability 1−p1-p, every measurement obeys

Pr⁡(ψ)=p ⟨ψ∣ρ∣ψ⟩+(1−p) ⟨ψ∣σ∣ψ⟩=⟨ψ∣(pρ+(1−p)σ)∣ψ⟩,\Pr(\psi)=p\,\langle\psi\rvert\rho\lvert\psi\rangle+(1-p)\,\langle\psi\rvert\sigma\lvert\psi\rangle=\langle\psi\rvert\bigl(p\rho+(1-p)\sigma\bigr)\lvert\psi\rangle,

where the second equality is just linearity: the sandwich ⟨ψ∣ ⋅ ∣ψ⟩\langle\psi\rvert\,\cdot\,\lvert\psi\rangle passes through the weighted sum. Averaging the probabilities and averaging the matrices give the same predictions, so the single matrix pρ+(1−p)σp\rho+(1-p)\sigma already describes the whole preparation.

The same argument runs with any number of choices. If a system is prepared in state ρk\rho_{k} with probability pkp_{k}, the resulting state is the weighted sum of the ρk\rho_{k}, and when the preparations are state vectors ∣ψk⟩\lvert\psi_{k}\rangle, each contributes the density matrix it makes on its own.

∑k=0m−1pkρkand∑k=0m−1pk∣ψk⟩⟨ψk∣.\sum_{k=0}^{m-1}p_{k}\rho_{k}\qquad\text{and}\qquad\sum_{k=0}^{m-1}p_{k}\lvert\psi_{k}\rangle\langle\psi_{k}\rvert.

A weighted sum with nonnegative weights adding to one is a convex combination, so the key property is this: convex combinations of density matrices represent probabilistic mixtures of quantum states. The result is always a density matrix again. Its trace is the average of traces, which is one, and ⟨ψ∣ρ∣ψ⟩\langle\psi\rvert\rho\lvert\psi\rangle is an average of nonnegative numbers, so it cannot be negative.

Prepared state
|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
12∣0⟩⟨0∣+12∣+⟩⟨+∣=(34141414)\tfrac{1}{2}\lvert0\rangle\langle0\rvert+\tfrac{1}{2}\lvert+\rangle\langle+\rvert=\begin{pmatrix}\tfrac{3}{4}&\tfrac{1}{4}\\\tfrac{1}{4}&\tfrac{1}{4}\end{pmatrix}
1/2

Classical states are density matrices

Take a classical state kk of Σ\Sigma. Its vector representation is ∣k⟩\lvert k\rangle, and its density matrix is ∣k⟩⟨k∣\lvert k\rangle\langle k\rvert: a matrix with a single 11 on the diagonal.

If a source prepares state kk with probability pkp_{k}, the resulting density matrix is

ρ=∑k=0n−1pk∣k⟩⟨k∣=(p00⋯00p1⋱⋮⋮⋱⋱00⋯0pn−1).\rho=\sum_{k=0}^{n-1}p_{k}\lvert k\rangle\langle k\rvert=\begin{pmatrix}p_{0}&0&\cdots&0\\0&p_{1}&\ddots&\vdots\\\vdots&\ddots&\ddots&0\\0&\cdots&0&p_{n-1}\end{pmatrix}.

Thus a classical probability distribution is exactly a diagonal density matrix. Classical probability sits inside the density-matrix formalism as the diagonal case, and the off-diagonal entries are what allow density matrices to represent genuinely quantum states.

The completely mixed state

The uniform classical distribution has a special name. If all nn classical states are equally likely, pk=1/np_{k}=1/n, the sum collapses to ρ=1nI\rho=\tfrac{1}{n}I. For a qubit prepared as ∣0⟩\lvert0\rangle or ∣1⟩\lvert1\rangle by a fair coin flip,

12∣0⟩⟨0∣+12∣1⟩⟨1∣=12I.\tfrac12\lvert0\rangle\langle0\rvert+\tfrac12\lvert1\rangle\langle1\rvert=\tfrac12 I.

This is the completely mixed state. It represents complete uncertainty about the qubit: every measurement, in every basis, gives its two outcomes with equal probability. On the Bloch sphere, it is the centre—the point furthest from every pure state.

The preparation procedure need not use ∣0⟩\lvert0\rangle and ∣1⟩\lvert1\rangle. Flipping a fair coin between ∣+⟩\lvert+\rangle and ∣−⟩\lvert-\rangle gives the same density matrix, 12∣+⟩⟨+∣+12∣−⟩⟨−∣=12I\tfrac12\lvert+\rangle\langle+\rvert+\tfrac12\lvert-\rangle\langle-\rvert=\tfrac12 I.

The two procedures therefore produce exactly the same physical state. Since the density matrix determines every measurement probability, no experiment on the qubit can distinguish how it was prepared.

Mixing is not superposition

A probabilistic mixture of ∣0⟩\lvert0\rangle and ∣1⟩\lvert1\rangle is not the same as the superposition ∣+⟩\lvert+\rangle. The mixture has density matrix 12I\tfrac12 I, whereas the superposition has

∣+⟩⟨+∣=(12121212)≠12I.\lvert+\rangle\langle+\rvert=\begin{pmatrix}\tfrac12&\tfrac12\\\tfrac12&\tfrac12\end{pmatrix}\neq\tfrac12 I.

Measuring both in the classical basis gives the same fifty-fifty outcomes, so that measurement alone cannot distinguish them. But measuring in the ∣+⟩,∣−⟩\lvert+\rangle,\lvert-\rangle basis does: ∣+⟩\lvert+\rangle gives ∣+⟩\lvert+\rangle with certainty, while the completely mixed state gives each outcome with probability 1/21/2.

The difference is in the off-diagonal entries. A superposition has them, and the classical mixture does not.

One matrix, many preparations

The two preparations of 12I\tfrac12 I above are not a special coincidence. In general, a mixed density matrix can be written as a convex combination of pure states in many different ways. These different decompositions correspond to different preparation procedures, but they all represent the same physical state.

The density matrix does not record which procedure was used. It records exactly what can affect measurement outcomes—and nothing about the preparation history beyond that.

The spectral theorem for density matrices

The states that every normal matrix has an orthonormal basis of eigenvectors. A density matrix is Hermitian, and hence normal, so the theorem applies. Moreover, its eigenvalues are real, and positive semidefiniteness ensures that they are nonnegative. Thus, for an n×nn\times n positive semidefinite matrix PP, there is an orthonormal basis {∣ψ0⟩,…,∣ψn−1⟩}\{\lvert\psi_{0}\rangle,\ldots,\lvert\psi_{n-1}\rangle\} and nonnegative real numbers λ0,…,λn−1\lambda_{0},\ldots,\lambda_{n-1} such that

P=∑k=0n−1λk∣ψk⟩⟨ψk∣.P=\sum_{k=0}^{n-1}\lambda_{k}\lvert\psi_{k}\rangle\langle\psi_{k}\rvert.

The decomposition above has an important interpretation. Each projector ∣ψk⟩⟨ψk∣\lvert\psi_{k}\rangle\langle\psi_{k}\rvert describes a pure state, so the density matrix is expressed as a weighted combination of pure states. We would therefore like to interpret the coefficients λk\lambda_{k} as probabilities. To do so, we only need to check that they are nonnegative and sum to one.

We already know that λk≥0\lambda_{k}\geq 0 because ρ\rho is positive semidefinite. It remains to check their sum. Each projector ∣ψk⟩⟨ψk∣\lvert\psi_{k}\rangle\langle\psi_{k}\rvert has trace one, so taking the trace of the spectral decomposition gives

Tr(ρ)=∑kλk Tr(∣ψk⟩⟨ψk∣)=∑kλk.\mathrm{Tr}(\rho)=\sum_{k}\lambda_{k}\,\mathrm{Tr}\bigl(\lvert\psi_{k}\rangle\langle\psi_{k}\rvert\bigr)=\sum_{k}\lambda_{k}.

Since a density matrix has trace one, ∑kλk=1\sum_{k}\lambda_{k}=1. Thus the eigenvalues are nonnegative numbers summing to one, so they form a probability vector. Writing pk=λkp_{k}=\lambda_{k}, any n×nn\times n density matrix can therefore be written in terms of an orthonormal basis {∣ψ0⟩,…,∣ψn−1⟩}\{\lvert\psi_{0}\rangle,\ldots,\lvert\psi_{n-1}\rangle\} and a probability vector (p0,…,pn−1)(p_{0},\ldots,p_{n-1}). This is the spectral decomposition of ρ\rho:

ρ=∑k=0n−1pk∣ψk⟩⟨ψk∣.\rho=\sum_{k=0}^{n-1}p_{k}\lvert\psi_{k}\rangle\langle\psi_{k}\rvert.

Every density matrix is therefore a probabilistic mixture of orthogonal pure states, with its eigenvalues as the probabilities. A mixed state can have many different probabilistic preparations, but the spectral decomposition gives one that is determined by the density matrix itself: its eigenstates and their corresponding eigenvalues.

|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
12∣0⟩⟨0∣+12∣+⟩⟨+∣=(34141414)\tfrac12\lvert0\rangle\langle0\rvert+\tfrac12\lvert+\rangle\langle+\rvert=\begin{pmatrix}\tfrac34&\tfrac14\\\tfrac14&\tfrac14\end{pmatrix}

Orthogonal qubit states sit at opposite points of the Bloch sphere, so the spectral decomposition is the segment through the centre: the diameter that the point lies on. Any other preparation joins two states that are not orthogonal, along a chord that misses the centre.

Reading purity off the eigenvalues

The decomposition also settles when a state is pure. If some pk=1p_{k}=1, all the other probabilities are zero, so the sum collapses to a single term ∣ψk⟩⟨ψk∣\lvert\psi_{k}\rangle\langle\psi_{k}\rvert, which is a pure state. On the Bloch sphere, this is a point on the surface. Otherwise, at least two probabilities are nonzero, so the state is a genuine mixture of orthogonal pure states. On the Bloch sphere, this lies inside the sphere.

The Bloch sphere

For a single qubit, a density matrix contains just three independent real numbers. Using them as coordinates turns each state into a point in the Bloch ball: pure states lie on its surface, the Bloch sphere, and mixed states lie inside. The picture lets us read the same state geometrically, while the matrix still tells us its measurement probabilities.

θφ|0⟩|1⟩|+⟩|−⟩|+i⟩|−i⟩
θ = 60°φ = 45°

Two angles locate a pure state: θ moves from pole to pole, and φ turns around the vertical axis.

60°
45°

The angle θ sets the sizes of the two amplitudes, while φ sets their relative phase:

∣ψ⟩=cos⁡θ2∣0⟩+eiφsin⁡θ2∣1⟩\lvert\psi\rangle=\cos\frac{\textcolor{#d97706}{\theta}}2\lvert0\rangle+e^{i\textcolor{#0d9488}{\varphi}}\sin\frac{\textcolor{#d97706}{\theta}}2\lvert1\rangle

For the selected angles:

∣ψ⟩=cos⁡(60∘2)∣0⟩+ei 45∘sin⁡(60∘2)∣1⟩≈0.87∣0⟩+0.50ei 45∘∣1⟩\begin{aligned}\lvert\psi\rangle&=\cos\left(\frac{60^\circ}{2}\right)\lvert0\rangle+e^{i\,45^\circ}\sin\left(\frac{60^\circ}{2}\right)\lvert1\rangle\\[4pt]&\approx0.87\lvert0\rangle+0.50e^{i\,45^\circ}\lvert1\rangle\end{aligned}
Why this works

Start with the density matrix of a single qubit. It is a 2×22\times2 matrix, but its four entries cannot vary independently. Because it is Hermitian, the diagonal entries are real and the two off-diagonal entries are complex conjugates. Its trace is one, so choosing the first diagonal entry also fixes the second. We are therefore left with exactly three real numbers: one diagonal value, and the real and imaginary parts of an off-diagonal entry.

Call these three numbers aa, bb, and cc, with the signs arranged as follows:

ρ=(ab−icb+ic1−a).\rho=\begin{pmatrix}a&b-ic\\b+ic&1-a\end{pmatrix}.

We now turn these three numbers into three coordinates. Shift and rescale them so that equal diagonal entries lie at height zero and the boundary of the allowed region will have radius one:

r=(rx,ry,rz)=(2b, 2c, 2a−1).r=(r_x,r_y,r_z)=(2b,\,2c,\,2a-1).

We can recover every entry of ρ\rho from these three coordinates, so the point rr contains exactly the same information as the density matrix. The vertical coordinate rzr_z records the difference between the two diagonal probabilities. The other two coordinates record the real and imaginary parts of the off-diagonal entry, including the phase information that the diagonal entries alone cannot describe.

Solving for aa, bb, and cc gives

ρ=12(1+rzrx−iryrx+iry1−rz)=I+rxσx+ryσy+rzσz2.\rho=\dfrac12\begin{pmatrix}1+r_z&r_x-ir_y\\r_x+ir_y&1-r_z\end{pmatrix}=\dfrac{I+r_x\sigma_x+r_y\sigma_y+r_z\sigma_z}{2}.

The expression on the right writes the same matrix in terms of the identity and the three Pauli matrices:

I=(1001),σx=(0110),σy=(0−ii0),σz=(100−1).I=\begin{pmatrix}1&0\\0&1\end{pmatrix},\qquad\sigma_x=\begin{pmatrix}0&1\\1&0\end{pmatrix},\qquad\sigma_y=\begin{pmatrix}0&-i\\i&0\end{pmatrix},\qquad\sigma_z=\begin{pmatrix}1&0\\0&-1\end{pmatrix}.

Each matrix supplies one of the three patterns needed to describe ρ\rho: σx\sigma_x changes the real off-diagonal part, σy\sigma_y changes the imaginary part, and σz\sigma_z changes the difference between the diagonal entries. Their coefficients are precisely the three coordinates of the point rr.

At this stage, we have mapped every qubit density matrix to a point in ordinary three-dimensional space. The remaining question is: which points are actually allowed?

The answer comes from the final condition on a density matrix: it must be positive semidefinite. The two eigenvalues of ρ\rho depend only on the distance ∥r∥\lVert r\rVert from the origin:

λ±=1±∥r∥2.\lambda_\pm=\dfrac{1\pm\lVert r\rVert}{2}.

Both eigenvalues must be nonnegative. If a point lies farther than one unit from the origin, then λ−\lambda_- becomes negative, so that point cannot represent a quantum state. Conversely, every point with ∥r∥≤1\lVert r\rVert\le1 gives a positive semidefinite density matrix. The allowed states therefore fill exactly the unit ball.

The boundary has ∥r∥=1\lVert r\rVert=1, so its eigenvalues are one and zero. A density matrix with these eigenvalues has rank one and describes a pure state. This is why the surface of the ball is called the Bloch sphere. Points inside the sphere represent mixed states.

At the centre, all three coordinates vanish, leaving ρ=I/2\rho=I/2: the completely mixed state. It gives equal probabilities for the two outcomes of every orthonormal measurement basis.

So far, we have described the entire Bloch ball using three coordinates. For a pure state, however, the point lies on the unit sphere, so only two angles are needed to locate it. These are the controls in the first view: θ\theta measures the angle down from the ∣0⟩\lvert0\rangle pole, while φ\varphi measures the turn around the vertical axis. The ranges θ∈[0,π]\theta\in[0,\pi] and φ∈[0,2π)\varphi\in[0,2\pi) cover the whole sphere, and the corresponding state vector is

∣ψ⟩=cos⁡θ2∣0⟩+eiφsin⁡θ2∣1⟩.\lvert\psi\rangle=\cos\frac\theta2\lvert0\rangle+e^{i\varphi}\sin\frac\theta2\lvert1\rangle.

Its density matrix is the outer product ∣ψ⟩⟨ψ∣\lvert\psi\rangle\langle\psi\rvert:

∣ψ⟩⟨ψ∣=(cos⁡2θ2e−iφcos⁡θ2sin⁡θ2eiφcos⁡θ2sin⁡θ2sin⁡2θ2).\lvert\psi\rangle\langle\psi\rvert=\begin{pmatrix}\cos^2\frac\theta2&e^{-i\varphi}\cos\frac\theta2\sin\frac\theta2\\e^{i\varphi}\cos\frac\theta2\sin\frac\theta2&\sin^2\frac\theta2\end{pmatrix}.

Euler’s formula eiφ=cos⁡φ+isin⁡φe^{i\varphi}=\cos\varphi+i\sin\varphi and the half-angle identities

cos⁡2θ2=1+cos⁡θ2,sin⁡2θ2=1−cos⁡θ2,cos⁡θ2sin⁡θ2=sin⁡θ2\cos^2\frac\theta2=\frac{1+\cos\theta}2,\qquad\sin^2\frac\theta2=\frac{1-\cos\theta}2,\qquad\cos\frac\theta2\sin\frac\theta2=\frac{\sin\theta}2

rewrite these entries in the Pauli form:

∣ψ⟩⟨ψ∣=I+sin⁡θcos⁡φ σx+sin⁡θsin⁡φ σy+cos⁡θ σz2.\lvert\psi\rangle\langle\psi\rvert=\frac{I+\sin\theta\cos\varphi\,\sigma_x+\sin\theta\sin\varphi\,\sigma_y+\cos\theta\,\sigma_z}2.

Reading off the coefficients gives the coordinates of the point:

r=(sin⁡θcos⁡φ, sin⁡θsin⁡φ, cos⁡θ).r=(\sin\theta\cos\varphi,\,\sin\theta\sin\varphi,\,\cos\theta).

This is exactly the unit vector pointing in the direction specified by the two angles. The two descriptions therefore capture the same state from two complementary viewpoints: the angles specify where the point is on the sphere, while the coordinates give the coefficients of the Pauli matrices in its density matrix. A global phase does not change the outer product ∣ψ⟩⟨ψ∣\lvert\psi\rangle\langle\psi\rvert, so it does not change the point on the Bloch sphere either.

This complete description by a single three-dimensional ball is special to a single qubit. An nn-qubit density matrix has 4n−14^n-1 independent real parameters — already fifteen for two qubits. The allowed states therefore form a convex body in a 4n−14^n-1-dimensional space, not a three-dimensional ball.

We can still draw the reduced state of each individual qubit as a point in its own Bloch ball, but those separate pictures do not capture everything about the joint state. In particular, they cannot represent all the correlations between the qubits.

Density matrices of multiple systems

Every density matrix so far has described one system, but nothing in the definition says how many systems that is. It asks for a square matrix whose rows and columns are indexed by the classical states of the system in question, with trace one and positive semidefinite. The only thing a system contributes to that is its set of classical states, so moving to several systems is a question of what to index by, not a new definition.

We answer it the way we did for state vectors. A pair (X,Y)(\mathsf{X},\mathsf{Y}) is treated as one compound system, whose classical states are the pairs (a,b)(a,b) in the Cartesian product Σ×Γ\Sigma\times\Gamma. That product is the new index set, rows and columns are labelled by its elements, and the two conditions are unchanged.

For two qubits that means a 4×44\times4 matrix whose rows and columns are labelled ∣00⟩,∣01⟩,∣10⟩,∣11⟩\lvert00\rangle,\lvert01\rangle,\lvert10\rangle,\lvert11\rangle in lexicographic order, exactly as for the state vectors. Everything we saw for a single system carries over unchanged: a pure state is still ∣ψ⟩⟨ψ∣\lvert\psi\rangle\langle\psi\rvert, mixtures are still convex combinations, and the spectral decomposition still applies.

The four Bell states are pure, so each one is the outer product of its own state vector. Writing them out makes it clear how their differences appear in the density matrix.

∣ϕ+⟩⟨ϕ+∣=(12001200000000120012)\lvert\phi^{+}\rangle\langle\phi^{+}\rvert=\begin{pmatrix}\textcolor{#7c3aed}{\tfrac12}&0&0&\textcolor{#d97706}{\tfrac12}\\0&\textcolor{#7c3aed}{0}&0&0\\0&0&\textcolor{#7c3aed}{0}&0\\\textcolor{#d97706}{\tfrac12}&0&0&\textcolor{#7c3aed}{\tfrac12}\end{pmatrix}
∣ϕ−⟩⟨ϕ−∣=(1200−1200000000−120012)\lvert\phi^{-}\rangle\langle\phi^{-}\rvert=\begin{pmatrix}\textcolor{#7c3aed}{\tfrac12}&0&0&\textcolor{#d97706}{-\tfrac12}\\0&\textcolor{#7c3aed}{0}&0&0\\0&0&\textcolor{#7c3aed}{0}&0\\\textcolor{#d97706}{-\tfrac12}&0&0&\textcolor{#7c3aed}{\tfrac12}\end{pmatrix}
∣ψ+⟩⟨ψ+∣=(00000121200121200000)\lvert\psi^{+}\rangle\langle\psi^{+}\rvert=\begin{pmatrix}\textcolor{#7c3aed}{0}&0&0&0\\0&\textcolor{#7c3aed}{\tfrac12}&\textcolor{#d97706}{\tfrac12}&0\\0&\textcolor{#d97706}{\tfrac12}&\textcolor{#7c3aed}{\tfrac12}&0\\0&0&0&\textcolor{#7c3aed}{0}\end{pmatrix}
∣ψ−⟩⟨ψ−∣=(0000012−1200−121200000)\lvert\psi^{-}\rangle\langle\psi^{-}\rvert=\begin{pmatrix}\textcolor{#7c3aed}{0}&0&0&0\\0&\textcolor{#7c3aed}{\tfrac12}&\textcolor{#d97706}{-\tfrac12}&0\\0&\textcolor{#d97706}{-\tfrac12}&\textcolor{#7c3aed}{\tfrac12}&0\\0&0&0&\textcolor{#7c3aed}{0}\end{pmatrix}

The two ϕ\phi states have the same diagonal and differ only in the signs of the corner entries, and the same is true of the two ψ\psi states. Since the diagonal holds the probabilities of a standard-basis measurement, those measurements cannot distinguish the members of a pair. The signs that do distinguish them sit off the diagonal, in the entries that encode the coherence between the basis states. This is the distinction between a superposition and a mixture.

Independence is a tensor product

If X\mathsf{X} is prepared in the state ρ\rho and, independently, Y\mathsf{Y} is prepared in the state σ\sigma, then the pair is in the state ρ⊗σ\rho\otimes\sigma. States of this form are called product states.

The same rule follows naturally from the definition of a density matrix. If ∣ψ⟩\lvert\psi\rangle and ∣π⟩\lvert\pi\rangle are prepared independently, the pair is ∣ψ⟩⊗∣π⟩\lvert\psi\rangle\otimes\lvert\pi\rangle, whose density matrix is

(∣ψ⟩⊗∣π⟩)(⟨ψ∣⊗⟨π∣)=(∣ψ⟩⟨ψ∣)⊗(∣π⟩⟨π∣)\bigl(\lvert\psi\rangle\otimes\lvert\pi\rangle\bigr)\bigl(\langle\psi\rvert\otimes\langle\pi\rvert\bigr)=\bigl(\lvert\psi\rangle\langle\psi\rvert\bigr)\otimes\bigl(\lvert\pi\rangle\langle\pi\rvert\bigr)

For instance, suppose one qubit is prepared in the state ∣0⟩\lvert0\rangle and another comes from a completely noisy source. The first factor is pure and the second is completely mixed, so their joint state is

∣0⟩⟨0∣⊗12I=∣0⟩⟨0∣⊗(12∣0⟩⟨0∣+12∣1⟩⟨1∣)=(120000120000000000)\lvert0\rangle\langle0\rvert\otimes\tfrac12 I=\lvert0\rangle\langle0\rvert\otimes\bigl(\tfrac12\lvert0\rangle\langle0\rvert+\tfrac12\lvert1\rangle\langle1\rvert\bigr)=\begin{pmatrix}\tfrac12&0&0&0\\0&\tfrac12&0&0\\0&0&0&0\\0&0&0&0\end{pmatrix}

Correlation between systems

A density matrix that cannot be written as a product state contains correlation between the two systems: learning something about one of them tells us something about the other. Correlation is not by itself a quantum phenomenon, and the simplest example is entirely classical. Alice and Bob share a uniform random bit, each holding a copy of it.

12 ∣0⟩⟨0∣⊗∣0⟩⟨0∣+12 ∣1⟩⟨1∣⊗∣1⟩⟨1∣=(120000000000000012)\tfrac12\,\lvert0\rangle\langle0\rvert\otimes\lvert0\rangle\langle0\rvert+\tfrac12\,\lvert1\rangle\langle1\rvert\otimes\lvert1\rangle\langle1\rvert=\begin{pmatrix}\tfrac12&0&0&0\\0&0&0&0\\0&0&0&0\\0&0&0&\tfrac12\end{pmatrix}

This is a mixture of two product states, but it is not itself a product state.

A product state is a joint state of the form ρ⊗σ\rho\otimes\sigma, where ρ\rho describes the first system and σ\sigma describes the second. It represents two systems whose joint probabilities factor into independent probabilities for the two systems. The diagonal of a product state contains the joint probabilities of the two standard-basis measurements. For a product ρ⊗σ\rho\otimes\sigma of two qubit states, those entries are ρ00σ00, ρ00σ11, ρ11σ00, ρ11σ11\rho_{00}\sigma_{00},\ \rho_{00}\sigma_{11},\ \rho_{11}\sigma_{00},\ \rho_{11}\sigma_{11}.

Now look at the diagonal of the state above. The first entry is 12\tfrac12, so both ρ00\rho_{00} and σ00\sigma_{00} must be nonzero. The third entry is zero, so ρ11=0\rho_{11}=0. But then the fourth entry, ρ11σ11\rho_{11}\sigma_{11}, must also be zero, whereas here it is 12\tfrac12. This is a contradiction.

The argument only used the diagonal, so it rules out a product state for every density matrix with this diagonal. One such density matrix is ∣ϕ+⟩⟨ϕ+∣\lvert\phi^{+}\rangle\langle\phi^{+}\rvert. We already know that this state is not a product state, but now we can see why from the diagonal alone: its joint probabilities cannot be written as independent probabilities for the two qubits. The corner entries are not needed at all.

Recording which state was prepared

A classical label lets us keep a record of which state was prepared instead of averaging the possibilities into a single mixed state. Suppose a source chooses kk with probability pkp_{k} and prepares the state ρk\rho_{k}. The probabilities satisfy pk≥0p_{k}\geq0 and ∑k=0m−1pk=1\sum_{k=0}^{m-1}p_{k}=1, and all the density matrices ρk\rho_{k} have the same dimensions. If we record the value of kk in a classical register alongside the system, we obtain an ensemble: a collection of possible states together with the probabilities with which they are prepared.

∑k=0m−1pk ∣k⟩⟨k∣⊗ρk\sum_{k=0}^{m-1}p_{k}\,\lvert k\rangle\langle k\rvert\otimes\rho_{k}

Here ∣k⟩⟨k∣\lvert k\rangle\langle k\rvert records the classical value kk, while ρk\rho_{k} is the state prepared when that value is chosen. The label therefore tells us which state was prepared, and the two are correlated: knowing kk tells us exactly which ρk\rho_{k} to expect.

If the label is discarded, only the system remains, and its state becomes the mixture ∑kpkρk\sum_{k}p_{k}\rho_{k}.

Separable states vs entanglement

A state is separable if it can be written as a mixture of product states:

ρ=∑k=0m−1pk ρk⊗σk\rho=\sum_{k=0}^{m-1}p_{k}\,\rho_{k}\otimes\sigma_{k}

This has a simple preparation interpretation: choose kk with probability pkp_{k}, then prepare ρk\rho_{k} on one side and σk\sigma_{k} on the other. The two systems can therefore be correlated, but all their correlations come from a shared classical random choice.

A state that cannot be written in this form is entangled. For pure states, this reduces to the familiar distinction: a pure state is separable exactly when it is a product state. The difference matters for mixed states, where a mixture of product states can be correlated without being entangled.

The definition is simple to state but difficult to apply. To prove that a state is separable, it is enough to find one decomposition of this form. To prove that it is entangled, every such decomposition must be ruled out.

Reduced states and the partial trace

A system can also be part of a larger one. Alice and Bob share an e-bit: Alice holds A\mathsf{A}, Bob holds B\mathsf{B}, and the pair is in the state

∣ϕ+⟩=12∣00⟩+12∣11⟩\lvert\phi^{+}\rangle=\frac{1}{\sqrt2}\lvert00\rangle+\frac{1}{\sqrt2}\lvert11\rangle

Alice can hold her qubit, measure it, and operate on it without ever touching Bob’s, so she needs a description of her qubit on its own—a state that gives the probabilities of all measurements she can perform. A state vector describes the pair as a whole, but it cannot describe Alice’s qubit alone: because the pair is entangled, the joint state cannot be written as a tensor product of a state for A\mathsf{A} and a state for B\mathsf{B}. We therefore need a different kind of description for one part of an entangled system.

Suppose Bob measures his qubit in the standard basis. We already know what that does to an e-bit.

OutcomeProbabilityResulting state of A\mathsf{A}
0012\tfrac12∣0⟩\lvert0\rangle
1112\tfrac12∣1⟩\lvert1\rangle

If Alice does not learn which outcome Bob obtained, her qubit is a probabilistic mixture of the two:

12∣0⟩⟨0∣+12∣1⟩⟨1∣=12I\tfrac12\lvert0\rangle\langle0\rvert+\tfrac12\lvert1\rangle\langle1\rvert=\tfrac12 I

Alice’s qubit is therefore in the completely mixed state. But this is not merely the state Alice would have after Bob happened to measure. Bob need not measure at all. Whatever Bob does to his qubit—or whether he does anything—cannot change the probabilities of Alice’s measurements. Otherwise, Alice could learn what Bob chose to do from her own measurement outcomes, allowing them to signal instantaneously at a distance.

So 12I\tfrac12 I is not a description of what happens after Bob measures. It is the description of Alice’s qubit itself. This is the reduced state of A\mathsf{A}. Imagining a measurement on B\mathsf{B} was only a way to derive it.

The reduced state in general

The e-bit example suggests a general strategy. To describe A\mathsf{A} alone, imagine measuring B\mathsf{B} in some basis, find the state that A\mathsf{A} would have for each possible outcome, and then average over those outcomes. We now carry out that construction for an arbitrary pure state ∣ψ⟩\lvert\psi\rangle of a pair (A,B)(\mathsf{A},\mathsf{B}).

Choose a basis {∣b⟩}\{\lvert b\rangle\} for B\mathsf{B}, with classical state set Γ\Gamma. Grouping the terms of ∣ψ⟩\lvert\psi\rangle according to the state of B\mathsf{B} gives

∣ψ⟩=∑b∈Γ∣ϕb⟩⊗∣b⟩∣ϕb⟩=(IA⊗⟨b∣)∣ψ⟩\begin{aligned}\lvert\psi\rangle&=\sum_{b\in\Gamma}\lvert\phi_{b}\rangle\otimes\lvert b\rangle\\[6pt]\lvert\phi_{b}\rangle&=\bigl(I_{\mathsf{A}}\otimes\langle b\rvert\bigr)\lvert\psi\rangle\end{aligned}

The vectors ∣ϕb⟩\lvert\phi_{b}\rangle are determined by the chosen basis and need not be normalised. This is useful because their squared lengths give the probabilities of the corresponding outcomes: measuring B\mathsf{B} in this basis gives bb with probability ∥∣ϕb⟩∥2\lVert\lvert\phi_{b}\rangle\rVert^{2}. When that probability is nonzero, the resulting state of A\mathsf{A} is the normalised vector ∣ϕb⟩/∥∣ϕb⟩∥\lvert\phi_{b}\rangle/\lVert\lvert\phi_{b}\rangle\rVert.

To describe A\mathsf{A} without keeping track of which outcome occurred, we average these states using their probabilities. The normalisation factors cancel:

∑b : ∥∣ϕb⟩∥>0∥∣ϕb⟩∥2⋅∣ϕb⟩⟨ϕb∣∥∣ϕb⟩∥2=∑b∈Γ∣ϕb⟩⟨ϕb∣\sum_{b\,:\,\lVert\lvert\phi_{b}\rangle\rVert>0}\lVert\lvert\phi_{b}\rangle\rVert^{2}\cdot\frac{\lvert\phi_{b}\rangle\langle\phi_{b}\rvert}{\lVert\lvert\phi_{b}\rangle\rVert^{2}}=\sum_{b\in\Gamma}\lvert\phi_{b}\rangle\langle\phi_{b}\rvert

For an outcome with zero probability, ∣ϕb⟩=0\lvert\phi_{b}\rangle=0, so its outer product contributes nothing. We can therefore include every basis state in the sum.

Substituting the definition of ∣ϕb⟩\lvert\phi_{b}\rangle gives the reduced state directly in terms of the density matrix of the pair:

ρA=∑b∈Γ(IA⊗⟨b∣)∣ψ⟩⟨ψ∣(IA⊗∣b⟩)\rho_{\mathsf{A}}=\sum_{b\in\Gamma}\bigl(I_{\mathsf{A}}\otimes\langle b\rvert\bigr)\lvert\psi\rangle\langle\psi\rvert\bigl(I_{\mathsf{A}}\otimes\lvert b\rangle\bigr)

At this point, ∣ψ⟩⟨ψ∣\lvert\psi\rangle\langle\psi\rvert is the only thing that identifies the pair as being in a pure state. The same expression makes sense for an arbitrary density matrix ρ\rho, so this gives the general definition of the reduced state of A\mathsf{A}:

ρA=∑b∈Γ(IA⊗⟨b∣)ρ(IA⊗∣b⟩)\rho_{\mathsf{A}}=\sum_{b\in\Gamma}\bigl(I_{\mathsf{A}}\otimes\langle b\rvert\bigr)\rho\bigl(I_{\mathsf{A}}\otimes\lvert b\rangle\bigr)

Exchanging the roles of the two systems gives the reduced state of B\mathsf{B}:

ρB=∑a∈Σ(⟨a∣⊗IB)ρ(∣a⟩⊗IB)\rho_{\mathsf{B}}=\sum_{a\in\Sigma}\bigl(\langle a\rvert\otimes I_{\mathsf{B}}\bigr)\rho\bigl(\lvert a\rangle\otimes I_{\mathsf{B}}\bigr)

The partial trace

The formula for the reduced state has a name. To obtain the state of A\mathsf{A} alone, we discard B\mathsf{B}. This operation is called the partial trace over B\mathsf{B} and is written TrB\mathrm{Tr}_{\mathsf{B}}. Similarly, TrA\mathrm{Tr}_{\mathsf{A}} traces out A\mathsf{A} and leaves the state of B\mathsf{B}.

The reduced states are therefore:

ρA=∑b∈Γ(IA⊗⟨b∣)ρ(IA⊗∣b⟩)=TrB(ρ)ρB=∑a∈Σ(⟨a∣⊗IB)ρ(∣a⟩⊗IB)=TrA(ρ)\begin{aligned}\rho_{\mathsf{A}}&=\sum_{b\in\Gamma}\bigl(I_{\mathsf{A}}\otimes\langle b\rvert\bigr)\rho\bigl(I_{\mathsf{A}}\otimes\lvert b\rangle\bigr)=\mathrm{Tr}_{\mathsf{B}}(\rho)\\[6pt]\rho_{\mathsf{B}}&=\sum_{a\in\Sigma}\bigl(\langle a\rvert\otimes I_{\mathsf{B}}\bigr)\rho\bigl(\lvert a\rangle\otimes I_{\mathsf{B}}\bigr)=\mathrm{Tr}_{\mathsf{A}}(\rho)\end{aligned}

The name comes from what the operation does to tensor products. For square matrices MM and NN, the partial trace takes the ordinary trace of the factor being discarded and leaves the other factor unchanged:

TrA(M⊗N)=Tr(M) NTrB(M⊗N)=Tr(N) M\begin{aligned}\mathrm{Tr}_{\mathsf{A}}(M\otimes N)&=\mathrm{Tr}(M)\,N\\[6pt]\mathrm{Tr}_{\mathsf{B}}(M\otimes N)&=\mathrm{Tr}(N)\,M\end{aligned}

It is useful to see what the partial trace does to the entries of a density matrix. Write a two-qubit density matrix as four 2×22\times2 blocks, one for each pair of basis states of A\mathsf{A}:

ρ=(C00C01C10C11)\rho=\begin{pmatrix}C_{00}&C_{01}\\C_{10}&C_{11}\end{pmatrix}

Then TrA\mathrm{Tr}_{\mathsf{A}} adds the two diagonal blocks, while TrB\mathrm{Tr}_{\mathsf{B}} takes the ordinary trace of each block:

TrA(ρ)=C00+C11\mathrm{Tr}_{\mathsf{A}}(\rho)=C_{00}+C_{11}
TrB(ρ)=(Tr(C00)Tr(C01)Tr(C10)Tr(C11))\mathrm{Tr}_{\mathsf{B}}(\rho)=\begin{pmatrix}\mathrm{Tr}(C_{00})&\mathrm{Tr}(C_{01})\\\mathrm{Tr}(C_{10})&\mathrm{Tr}(C_{11})\end{pmatrix}

For example, suppose the pair is prepared as ∣0⟩⊗∣0⟩\lvert0\rangle\otimes\lvert0\rangle or ∣1⟩⊗∣+⟩\lvert1\rangle\otimes\lvert+\rangle with equal probability:

ρ=12 ∣0⟩⟨0∣⊗∣0⟩⟨0∣+12 ∣1⟩⟨1∣⊗∣+⟩⟨+∣\rho=\tfrac12\,\lvert0\rangle\langle0\rvert\otimes\lvert0\rangle\langle0\rvert+\tfrac12\,\lvert1\rangle\langle1\rvert\otimes\lvert+\rangle\langle+\rvert

Applying the partial trace to each term gives

ρA=12∣0⟩⟨0∣+12∣1⟩⟨1∣=12I\rho_{\mathsf{A}}=\tfrac12\lvert0\rangle\langle0\rvert+\tfrac12\lvert1\rangle\langle1\rvert=\tfrac12 I
ρB=12∣0⟩⟨0∣+12∣+⟩⟨+∣\rho_{\mathsf{B}}=\tfrac12\lvert0\rangle\langle0\rvert+\tfrac12\lvert+\rangle\langle+\rvert

Tracing out multiple systems

The same idea works for any number of systems. We can divide a compound system into the part we keep and the part we discard, and then trace out whichever systems we do not need.

For a triple (A,B,C)(\mathsf{A},\mathsf{B},\mathsf{C}) in the state ρ\rho, tracing out B\mathsf{B} leaves the state of (A,C)(\mathsf{A},\mathsf{C}):

ρAC=∑b∈Γ(IA⊗⟨b∣⊗IC)ρ(IA⊗∣b⟩⊗IC)\rho_{\mathsf{AC}}=\sum_{b\in\Gamma}\bigl(I_{\mathsf{A}}\otimes\langle b\rvert\otimes I_{\mathsf{C}}\bigr)\rho\bigl(I_{\mathsf{A}}\otimes\lvert b\rangle\otimes I_{\mathsf{C}}\bigr)

Tracing out both A\mathsf{A} and B\mathsf{B} leaves the state of C\mathsf{C}:

ρC=∑a∈Σ∑b∈Γ(⟨a∣⊗⟨b∣⊗IC)ρ(∣a⟩⊗∣b⟩⊗IC)\rho_{\mathsf{C}}=\sum_{a\in\Sigma}\sum_{b\in\Gamma}\bigl(\langle a\rvert\otimes\langle b\rvert\otimes I_{\mathsf{C}}\bigr)\rho\bigl(\lvert a\rangle\otimes\lvert b\rangle\otimes I_{\mathsf{C}}\bigr)

Systems can also be traced out one at a time. Tracing out B\mathsf{B} and then A\mathsf{A} gives the same ρC\rho_{\mathsf{C}} as tracing out both together.

What information do reduced states lose?

The reduced states of a pair do not contain enough information to reconstruct the joint state. Two different joint states can give exactly the same state for A\mathsf{A} and exactly the same state for B\mathsf{B}. The difference can live entirely in the correlations between them.

Joint state
∣ϕ+⟩=∣00⟩+∣11⟩2\lvert\phi^+\rangle=\frac{\lvert00\rangle+\lvert11\rangle}{\sqrt2}

Joint state of A and B

Basis: ∣00⟩,∣01⟩,∣10⟩,∣11⟩\lvert00\rangle,\lvert01\rangle,\lvert10\rangle,\lvert11\rangle

ρAB=(12001200000000120012)\rho_{\mathsf{AB}}=\begin{pmatrix}\textcolor{#7c3aed}{\tfrac12}&0&0&\textcolor{#7c3aed}{\tfrac12}\\0&0&0&0\\0&0&0&0\\\textcolor{#7c3aed}{\tfrac12}&0&0&\textcolor{#7c3aed}{\tfrac12}\end{pmatrix}
Standard-basis measurement

Reduced state of A

ρA=(120012)\rho_{\mathsf{A}}=\begin{pmatrix}\tfrac12&0\\0&\tfrac12\end{pmatrix}

Trace out B.

Reduced state of B

ρB=(120012)\rho_{\mathsf{B}}=\begin{pmatrix}\tfrac12&0\\0&\tfrac12\end{pmatrix}

Trace out A.

Quantum channels

So far, density matrices have been used to describe quantum states, while matrices such as UU have been used to describe unitary transformations of pure states. But a quantum system does not undergo only ideal unitary transformations. It can be measured, reset, sent through a noisy device, or interact with another system that is then discarded. A general framework is therefore needed to describe these processes.

That framework is provided by quantum channels. A quantum channel is a physical transformation that takes a quantum state as input and produces another quantum state as output. The name channel comes from viewing a quantum system as passing through a process: the input is the state before the process, and the output is the state afterwards.

Channels are usually denoted by capital Greek letters such as Φ\Phi, Ψ\Psi, and Ξ\Xi. If a channel Φ\Phi is applied to a system in the state ρ\rho, the resulting state is written Φ(ρ)\Phi(\rho).

What makes a mapping a channel

Not every mapping from matrices to matrices represents a physically possible transformation. Two requirements distinguish valid channels.

  • Channels are linear mappings. If a state is prepared as a mixture of two states, the channel must produce the same mixture of the two corresponding outputs:

    Φ(pρ+(1−p)σ)=p Φ(ρ)+(1−p) Φ(σ)\Phi\bigl(p\rho+(1-p)\sigma\bigr)=p\,\Phi(\rho)+(1-p)\,\Phi(\sigma)
  • Channels preserve density matrices, even as part of a larger system. Applying a channel to a density matrix must produce another density matrix. This must remain true when the input system is one part of a larger system, including when the two systems are entangled.

The second requirement is stronger than simply checking that Φ(ρ)\Phi(\rho) is a density matrix for every density matrix ρ\rho. It is what allows the channel to act on a subsystem of a larger quantum system.

Input and output systems

Every channel Φ\Phi has an input system X\mathsf{X} and an output system Y\mathsf{Y}. Conceptually, Φ\Phi transforms X\mathsf{X} into Y\mathsf{Y}. The input is the system before the transformation, and the output is the system afterwards.

The input and output systems can also be the same. This is the case we encounter most often. Then Φ\Phi simply changes the state of a system, as a gate changes the state of the qubit it acts on.

Acting on part of a compound system

To see why the second requirement matters, suppose Z\mathsf{Z} is an additional system with classical state set Γ\Gamma, and the pair (Z,X)(\mathsf{Z},\mathsf{X}) is in the state ρ\rho. Choose the basis {∣a⟩:a∈Γ}\{\lvert a\rangle:a\in\Gamma\} for Z\mathsf{Z}. Every density matrix of the pair can then be written by grouping its entries according to the basis states of Z\mathsf{Z}:

ρ=∑a,b∈Γ∣a⟩⟨b∣⊗ρa,b\rho=\sum_{a,b\in\Gamma}\lvert a\rangle\langle b\rvert\otimes\rho_{a,b}

The matrices ρa,b\rho_{a,b} are not generally density matrices themselves. They are simply the blocks of ρ\rho corresponding to the pair of basis states aa and bb of Z\mathsf{Z}.

Now apply Φ\Phi to X\mathsf{X} alone. The system Z\mathsf{Z} is left untouched, while X\mathsf{X} is transformed into Y\mathsf{Y}. The resulting state of (Z,Y)(\mathsf{Z},\mathsf{Y}) is

∑a,b∈Γ∣a⟩⟨b∣⊗Φ(ρa,b)\sum_{a,b\in\Gamma}\lvert a\rangle\langle b\rvert\otimes\Phi(\rho_{a,b})

Nothing was done to Z\mathsf{Z}, so the factors ∣a⟩⟨b∣\lvert a\rangle\langle b\rvert remain unchanged. By linearity, Φ\Phi acts separately on every block ρa,b\rho_{a,b}.

If Γ={0,…,m−1}\Gamma=\{0,\ldots,m-1\}, we can see the same transformation as a statement about block matrices. Write ρ\rho as an m×mm\times m grid of blocks, one for each pair of basis states of Z\mathsf{Z}. The channel is then applied to every block:

ρ=(ρ0,0⋯ρ0,m−1⋮⋱⋮ρm−1,0⋯ρm−1,m−1)  ⟼  (Φ(ρ0,0)⋯Φ(ρ0,m−1)⋮⋱⋮Φ(ρm−1,0)⋯Φ(ρm−1,m−1))\rho=\begin{pmatrix}\rho_{0,0}&\cdots&\rho_{0,m-1}\\\vdots&\ddots&\vdots\\\rho_{m-1,0}&\cdots&\rho_{m-1,m-1}\end{pmatrix}\;\longmapsto\;\begin{pmatrix}\Phi(\rho_{0,0})&\cdots&\Phi(\rho_{0,m-1})\\\vdots&\ddots&\vdots\\\Phi(\rho_{m-1,0})&\cdots&\Phi(\rho_{m-1,m-1})\end{pmatrix}

For Φ\Phi to be a valid channel, the matrix on the right must be a density matrix for every choice of the system Z\mathsf{Z} and every density matrix ρ\rho on the joint system Z⊗X\mathsf{Z}\otimes\mathsf{X}.

Testing whether a mapping is a quantum channel

Valid channels
Invalid mappings
Φ(A)=AT\Phi(A)=A^{\mathsf T}
One qubit on its own∣+i⟩=∣0⟩+i∣1⟩2|{+i}\rangle=\frac{|0\rangle+i|1\rangle}{\sqrt2}
12(1−ii1)  ⟼  12(1i−i1)\frac12\begin{pmatrix}1&-i\\i&1\end{pmatrix}\;\longmapsto\;\frac12\begin{pmatrix}1&i\\-i&1\end{pmatrix}
−½01210
Pass: All ≥ 0Pass: Trace = 1
Half of an entangled pair∣ϕ+⟩=∣00⟩+∣11⟩2|\phi^+\rangle=\frac{|00\rangle+|11\rangle}{\sqrt2}
12(1001000000001001)  ⟼  12(1000001001000001)\frac12\begin{pmatrix}1&0&0&1\\0&0&0&0\\0&0&0&0\\1&0&0&1\end{pmatrix}\;\longmapsto\;\frac12\begin{pmatrix}1&0&0&0\\0&0&1&0\\0&1&0&0\\0&0&0&1\end{pmatrix}
−½012½½½−½
Fail: Negative eigenvaluePass: Trace = 1

Unitary channels

The familiar has a direct description in terms of density matrices. If UU is a unitary matrix acting on a system X\mathsf{X}, then ∣ψ⟩→U∣ψ⟩\lvert\psi\rangle\to U\lvert\psi\rangle becomes ρ→UρU†\rho\to U\rho U^{\dagger}, so the corresponding channel is Φ(ρ)=UρU†\Phi(\rho)=U\rho U^{\dagger}.

Channels of this form are called unitary channels. They transform X\mathsf{X} into itself, so the input and output systems are the same. For a pure state ρ=∣ψ⟩⟨ψ∣\rho=\lvert\psi\rangle\langle\psi\rvert, this gives Φ(ρ)=U∣ψ⟩⟨ψ∣U†=(U∣ψ⟩)(U∣ψ⟩)†\Phi(\rho)=U\lvert\psi\rangle\langle\psi\rvert U^{\dagger}=\bigl(U\lvert\psi\rangle\bigr)\bigl(U\lvert\psi\rangle\bigr)^{\dagger}, exactly recovering the familiar transformation.

A unitary channel satisfies both requirements for a physical channel:

It is linear: if the input state is a mixture of two states, the output is the same mixture of the two outputs. In other words, applying the channel does not change the probabilities in the mixture. For a mixture pρ+(1−p)σp\rho+(1-p)\sigma, the channel must act as Φ(pρ+(1−p)σ)=p Φ(ρ)+(1−p) Φ(σ)\Phi\bigl(p\rho+(1-p)\sigma\bigr)=p\,\Phi(\rho)+(1-p)\,\Phi(\sigma).

For the unitary channel Φ(ρ)=UρU†\Phi(\rho)=U\rho U^{\dagger}, this follows directly:

Φ(pρ+(1−p)σ)=U(pρ+(1−p)σ)U†=p UρU†+(1−p) UσU†=p Φ(ρ)+(1−p) Φ(σ)\Phi\bigl(p\rho+(1-p)\sigma\bigr)=U\bigl(p\rho+(1-p)\sigma\bigr)U^{\dagger}=p\,U\rho U^{\dagger}+(1-p)\,U\sigma U^{\dagger}=p\,\Phi(\rho)+(1-p)\,\Phi(\sigma)

It is also completely positive. This means that the channel must remain valid when X\mathsf{X} is part of a larger system. Suppose Z\mathsf{Z} is another system and the joint state of (Z,X)(\mathsf{Z},\mathsf{X}) is ρ\rho. Applying Φ\Phi to X\mathsf{X} alone gives (IZ⊗U)ρ(IZ⊗U)†\bigl(I_{\mathsf{Z}}\otimes U\bigr)\rho\bigl(I_{\mathsf{Z}}\otimes U\bigr)^{\dagger}.

But IZ⊗UI_{\mathsf{Z}}\otimes U is itself unitary. So this is simply a unitary transformation of the joint state, which preserves the properties required of a density matrix: its trace remains one, and positive semidefiniteness is preserved. Therefore the result is a valid density matrix for every larger system Z\mathsf{Z} and every joint state ρ\rho.

The simplest choice for UU is the identity matrix II. We call the resulting channel the identity channel, Id(ρ)=ρ\mathrm{Id}(\rho)=\rho. It leaves the state unchanged.

Convex combinations of channels

Density matrices let us average states, and we can do the same with channels. Let Φ0\Phi_{0} and Φ1\Phi_{1} be channels from X\mathsf{X} to Y\mathsf{Y}, and let p∈[0,1]p\in[0,1]. If we choose Φ0\Phi_{0} with probability pp and Φ1\Phi_{1} with probability 1−p1-p, the resulting channel is

Ψ=p Φ0+(1−p) Φ1\Psi=p\,\Phi_{0}+(1-p)\,\Phi_{1}

Applied to a state ρ\rho, this channel produces the corresponding mixture of the two outputs:

Ψ(ρ)=(p Φ0+(1−p) Φ1)(ρ)=p Φ0(ρ)+(1−p) Φ1(ρ)\Psi(\rho)=\bigl(p\,\Phi_{0}+(1-p)\,\Phi_{1}\bigr)(\rho)=p\,\Phi_{0}(\rho)+(1-p)\,\Phi_{1}(\rho)

More generally, if Φ0,…,Φm−1\Phi_{0},\ldots,\Phi_{m-1} are channels and (p0,…,pm−1)(p_{0},\ldots,p_{m-1}) is a probability vector, we can choose channel Φk\Phi_{k} with probability pkp_{k}. Their convex combination is again a channel, and its action on a state ρ\rho is described by

Ψ=∑k=0m−1pkΦk\Psi=\sum_{k=0}^{m-1}p_{k}\Phi_{k}
Ψ(ρ)=∑k=0m−1pkΦk(ρ)\Psi(\rho)=\sum_{k=0}^{m-1}p_{k}\Phi_{k}(\rho)

It is linear because each Φk\Phi_{k} is linear. Its output is a probabilistic mixture of density matrices, so it is again a density matrix. The same reasoning applies when the channel acts on part of a larger system: each Φk\Phi_{k} produces a valid joint state, and their mixture is valid as well.

A random-operation channel in action

Input state
Choose between

One run · one pure output

∣+⟩\lvert+\rangle
50%ILeave unchanged
∣+⟩\lvert+\rangle
50%ZApply Z
∣−⟩\lvert-\rangle
50%

Choice unknown · average output

|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
Distance from the center0.00

Maximally mixed

Ψ(ρ)=12 ρ+12 ZρZ†=(120012)\Psi(\rho)=\textcolor{#0284c7}{\tfrac{1}{2}}\,\rho+\textcolor{#0f766e}{\tfrac{1}{2}}\,Z\rho Z^{\dagger}=\textcolor{#7c3aed}{\begin{pmatrix}\tfrac{1}{2}&0\\0&\tfrac{1}{2}\end{pmatrix}}

Common non-unitary quantum channels

Three channels appear often enough to have names of their own. They describe physical processes that cannot be represented by a unitary gate alone, including resetting a qubit and different forms of noise.

The qubit reset channel

The qubit reset channel discards the state a qubit is in and prepares ∣0⟩|0\rangle: Λ(ρ)=Tr(ρ) ∣0⟩⟨0∣\Lambda(\rho)=\mathrm{Tr}(\rho)\,|0\rangle\langle0|.

Input state
|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
Input|0⟩0%|1⟩100%Output|0⟩100%|1⟩0%

The output is always |0⟩, whatever the input.

Resetting a qubit destroys entanglement

Every input ends at the same place, so the output contains no information about the input state. The fixed output is ∣0⟩⟨0∣|0\rangle\langle0|. To extend this idea from normalized states to arbitrary matrices, the channel is defined as Λ(X)=Tr(X) ∣0⟩⟨0∣\Lambda(X)=\mathrm{Tr}(X)\,|0\rangle\langle0|.

For a density matrix, Tr(X)=1\mathrm{Tr}(X)=1, so this simply gives ∣0⟩⟨0∣|0\rangle\langle0|. The trace factor is what makes the map linear on arbitrary matrices as well.

To see how the trace factor determines the action of the reset channel, look at the four basis operators. Since the trace is the sum of the diagonal entries, each diagonal projector has trace 1. The reset channel therefore maps both of them to the fixed state ∣0⟩⟨0∣|0\rangle\langle0|:

∣0⟩⟨0∣=(1000),∣1⟩⟨1∣=(0001)Λ(∣0⟩⟨0∣)=Λ(∣1⟩⟨1∣)=∣0⟩⟨0∣\begin{aligned}&|0\rangle\langle0|=\begin{pmatrix}1&0\\0&0\end{pmatrix},\qquad|1\rangle\langle1|=\begin{pmatrix}0&0\\0&1\end{pmatrix}\\[4pt]&\Lambda(|0\rangle\langle0|)=\Lambda(|1\rangle\langle1|)=|0\rangle\langle0|\end{aligned}

The off-diagonal operators have no diagonal entries, so their trace is 0. The reset channel therefore maps both of them to zero:

∣0⟩⟨1∣=(0100),∣1⟩⟨0∣=(0010)Λ(∣0⟩⟨1∣)=Λ(∣1⟩⟨0∣)=0\begin{aligned}&|0\rangle\langle1|=\begin{pmatrix}0&1\\0&0\end{pmatrix},\qquad|1\rangle\langle0|=\begin{pmatrix}0&0\\1&0\end{pmatrix}\\[4pt]&\Lambda(|0\rangle\langle1|)=\Lambda(|1\rangle\langle0|)=0\end{aligned}

Now consider two qubits, A and B, with A first, in the Bell state:

∣ϕ+⟩=(∣00⟩+∣11⟩)/2|\phi^+\rangle=(|00\rangle+|11\rangle)/\sqrt2

Writing the Bell state as a density matrix gives four terms:

∣ϕ+⟩⟨ϕ+∣=12∣0⟩⟨0∣⊗∣0⟩⟨0∣+12∣0⟩⟨1∣⊗∣0⟩⟨1∣+12∣1⟩⟨0∣⊗∣1⟩⟨0∣+12∣1⟩⟨1∣⊗∣1⟩⟨1∣|\phi^+\rangle\langle\phi^+|=\tfrac12|0\rangle\langle0|\otimes|0\rangle\langle0|+\tfrac12|0\rangle\langle1|\otimes|0\rangle\langle1|+\tfrac12|1\rangle\langle0|\otimes|1\rangle\langle0|+\tfrac12|1\rangle\langle1|\otimes|1\rangle\langle1|

Now apply the reset channel to qubit A, leaving B unchanged. In each term, the channel therefore acts only on the first factor. The two off-diagonal terms disappear, while both diagonal terms acquire the same first factor, ∣0⟩⟨0∣|0\rangle\langle0|:

(Λ⊗Id)(∣ϕ+⟩⟨ϕ+∣)=12Λ(∣0⟩⟨0∣)⊗∣0⟩⟨0∣+12Λ(∣0⟩⟨1∣)⊗∣0⟩⟨1∣+12Λ(∣1⟩⟨0∣)⊗∣1⟩⟨0∣+12Λ(∣1⟩⟨1∣)⊗∣1⟩⟨1∣=12∣0⟩⟨0∣⊗∣0⟩⟨0∣+0+0+12∣0⟩⟨0∣⊗∣1⟩⟨1∣=∣0⟩⟨0∣⊗12(∣0⟩⟨0∣+∣1⟩⟨1∣)=∣0⟩⟨0∣⊗I2\begin{aligned}(\Lambda\otimes\mathrm{Id})(|\phi^+\rangle\langle\phi^+|)&=\tfrac12\Lambda(|0\rangle\langle0|)\otimes|0\rangle\langle0|+\tfrac12\Lambda(|0\rangle\langle1|)\otimes|0\rangle\langle1|+\tfrac12\Lambda(|1\rangle\langle0|)\otimes|1\rangle\langle0|+\tfrac12\Lambda(|1\rangle\langle1|)\otimes|1\rangle\langle1|\\&=\tfrac12|0\rangle\langle0|\otimes|0\rangle\langle0|+0+0+\tfrac12|0\rangle\langle0|\otimes|1\rangle\langle1|\\&=|0\rangle\langle0|\otimes\tfrac12\bigl(|0\rangle\langle0|+|1\rangle\langle1|\bigr)\\&=|0\rangle\langle0|\otimes\frac{I}{2}\end{aligned}

The reset channel breaks the entanglement between A and B: A is replaced by the fixed state ∣0⟩|0\rangle, while B is left in its original reduced state I/2I/2. The pair is no longer correlated—the output is simply the product state ∣0⟩⟨0∣⊗I/2|0\rangle\langle0|\otimes I/2.

The completely dephasing channel

The completely dephasing channel zeros out the off-diagonal matrix entries, keeping the chances of measuring 0 or 1 while erasing the phase coherence that lets those possibilities interfere:

Δ ⁣(α00α01α10α11)=(α0000α11)\Delta\!\begin{pmatrix}\alpha_{00}&\alpha_{01}\\\alpha_{10}&\alpha_{11}\end{pmatrix}=\begin{pmatrix}\alpha_{00}&0\\0&\alpha_{11}\end{pmatrix}
Input state
|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
Input|0⟩80%|1⟩20%Output|0⟩80%|1⟩20%
Phase coherence0.80 → 0.00
100%
No noiseComplete dephasing

The 0/1 probabilities stay fixed while phase coherence fades.

Dephasing destroys coherence but keeps correlation

Averaging the complete channel with the identity channel gives a partial one, applying Δ\Delta with probability ε\varepsilon and leaving the state alone otherwise. It shrinks the off-diagonal entries by 1−ε1-\varepsilon rather than removing them, so ε=0\varepsilon=0 leaves the state unchanged and ε=1\varepsilon=1 is complete dephasing:

Δε=(1−ε) Id+ε ΔΔε ⁣(α00α01α10α11)=(α00(1−ε)α01(1−ε)α10α11)\begin{aligned}&\Delta_{\varepsilon}=(1-\varepsilon)\,\mathrm{Id}+\varepsilon\,\Delta\\[4pt]&\Delta_{\varepsilon}\!\begin{pmatrix}\alpha_{00}&\alpha_{01}\\\alpha_{10}&\alpha_{11}\end{pmatrix}=\begin{pmatrix}\alpha_{00}&(1-\varepsilon)\alpha_{01}\\(1-\varepsilon)\alpha_{10}&\alpha_{11}\end{pmatrix}\end{aligned}

The diagonal entries store the 0/1 probabilities, so they stay unchanged. The off-diagonal entries carry the coherence, measured by 2∣α01∣2|\alpha_{01}|. The channel can also be written by what it does to the four basis operators: the diagonal projectors are left alone, and the off-diagonal operators are sent to zero:

Δ(∣0⟩⟨0∣)=∣0⟩⟨0∣,Δ(∣1⟩⟨1∣)=∣1⟩⟨1∣Δ(∣0⟩⟨1∣)=Δ(∣1⟩⟨0∣)=0\begin{aligned}&\Delta(|0\rangle\langle0|)=|0\rangle\langle0|,\qquad\Delta(|1\rangle\langle1|)=|1\rangle\langle1|\\[4pt]&\Delta(|0\rangle\langle1|)=\Delta(|1\rangle\langle0|)=0\end{aligned}

Now consider two qubits, A and B, with A first, in the Bell state ∣ϕ+⟩=(∣00⟩+∣11⟩)/2|\phi^+\rangle=(|00\rangle+|11\rangle)/\sqrt2, and apply the dephasing channel to A while leaving B unchanged. Expanding the pair into its four terms, the channel acts only on the first factor of each:

(Δ⊗Id)(∣ϕ+⟩⟨ϕ+∣)=12Δ(∣0⟩⟨0∣)⊗∣0⟩⟨0∣+12Δ(∣0⟩⟨1∣)⊗∣0⟩⟨1∣+12Δ(∣1⟩⟨0∣)⊗∣1⟩⟨0∣+12Δ(∣1⟩⟨1∣)⊗∣1⟩⟨1∣=12∣0⟩⟨0∣⊗∣0⟩⟨0∣+0+0+12∣1⟩⟨1∣⊗∣1⟩⟨1∣=12∣00⟩⟨00∣+12∣11⟩⟨11∣\begin{aligned}(\Delta\otimes\mathrm{Id})(|\phi^+\rangle\langle\phi^+|)&=\tfrac12\Delta(|0\rangle\langle0|)\otimes|0\rangle\langle0|+\tfrac12\Delta(|0\rangle\langle1|)\otimes|0\rangle\langle1|+\tfrac12\Delta(|1\rangle\langle0|)\otimes|1\rangle\langle0|+\tfrac12\Delta(|1\rangle\langle1|)\otimes|1\rangle\langle1|\\&=\tfrac12|0\rangle\langle0|\otimes|0\rangle\langle0|+0+0+\tfrac12|1\rangle\langle1|\otimes|1\rangle\langle1|\\&=\tfrac12|00\rangle\langle00|+\tfrac12|11\rangle\langle11|\end{aligned}

Unlike reset, dephasing keeps both diagonal projectors in place, and only the terms linking 00 with 11 disappear. The pair’s quantum coherence is gone, but its classical correlation survives: measurements of A and B in the 0/1 basis still always agree, so the output is an equal mixture of 00 and 11.

The completely depolarizing channel

The completely depolarizing channel erases all information about the input and always outputs the completely mixed state:

Ω(ρ)=Tr(ρ) I2\Omega(\rho)=\mathrm{Tr}(\rho)\,\frac{I}{2}
Input state
|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
Input|0⟩80%|1⟩20%Output|0⟩50%|1⟩50%
100%
No noiseComplete depolarizing

Every input reaches the center: 0 and 1 each have probability 50%.

Depolarizing is the extreme end of a family of noise channels

As with reset, the trace factor lets the channel act on arbitrary matrices, not just normalized states. The completely depolarizing channel is defined by Ω(X)=Tr(X) I/2\Omega(X)=\mathrm{Tr}(X)\,I/2.

For a density matrix, Tr(X)=1\mathrm{Tr}(X)=1, so every input is mapped to the same output, Ω(ρ)=I/2\Omega(\rho)=I/2.

This is the completely mixed state: the center of the Bloch sphere, with no preferred direction. Every measurement basis therefore gives equal probabilities. This contrasts with reset, which always produces the pure state ∣0⟩⟨0∣|0\rangle\langle0|.

Complete depolarization is the extreme case. A weaker form of depolarization leaves the state unchanged with probability 1−ε1-\varepsilon and completely depolarizes it with probability ε\varepsilon:

Ωε=(1−ε) Id+ε ΩΩε(ρ)=(1−ε)ρ+εI2\begin{aligned}&\Omega_{\varepsilon}=(1-\varepsilon)\,\mathrm{Id}+\varepsilon\,\Omega\\[4pt]&\Omega_{\varepsilon}(\rho)=(1-\varepsilon)\rho+\varepsilon\frac{I}{2}\end{aligned}

At ε=0\varepsilon=0, nothing changes. At ε=1\varepsilon=1, every state is mapped to I/2I/2. For values in between, the Bloch vector keeps its direction but shrinks by the factor 1−ε1-\varepsilon.

Channel representations

How do we write a channel down?

A linear mapping from vectors to vectors is represented by a matrix, in the familiar way: the matrix multiplies a column vector and returns another one. Channels are linear as well, but they map matrices to matrices, so a matrix acting on a vector is the wrong shape to describe one.

Sometimes a simple formula expresses the action of a channel, such as Λ(ρ)=Tr(ρ) ∣0⟩⟨0∣\Lambda(\rho)=\mathrm{Tr}(\rho)\,|0\rangle\langle0| for the qubit reset channel. That is not practical in general, so the question is how to express an arbitrary channel in mathematical terms.

Stinespring representations

Every channel can be implemented in the same three steps:

  1. Form a compound system from the input system and an initialized workspace system.
  2. Perform a unitary operation on the compound system.
  3. Discard everything except the output system.

For a channel Φ\Phi from a system X\mathsf{X} to a system Y\mathsf{Y}, this is a circuit on two wires, for a suitable choice of the workspace W\mathsf{W} and the discarded system G\mathsf{G}:

ρ\rho
∣0⟩|0\rangle
UXYWG
Φ(ρ)\Phi(\rho)

Such a description, consisting of the unitary operation together with a specification of the input and output systems, is a Stinespring representation of the channel. Nothing about the channel is left outside it: the randomness and the information loss both come from discarding G\mathsf{G} at the end.

For a channel from a system to itself, the picture is the same with Y=X\mathsf{Y}=\mathsf{X} and G=W\mathsf{G}=\mathsf{W}: the workspace goes in initialized and comes back out to be thrown away.

Kraus representations

A Kraus representation describes a quantum channel using matrix multiplication and addition, making it especially convenient for calculations. In general, a channel can be written as

Φ(ρ)=∑k=0N−1Ak ρ Ak†\Phi(\rho)=\sum_{k=0}^{N-1}A_{k}\,\rho\,A_{k}^{\dagger}

The matrices A0,…,AN−1A_{0},\ldots,A_{N-1} are called Kraus operators. They all have the same dimensions. If the channel maps an input system to an output system, each column of a Kraus operator corresponds to an input basis state, and each row corresponds to an output basis state. The operators therefore need not be square when the input and output systems have different dimensions.

The Kraus operators are not arbitrary. To ensure that the channel maps density matrices to density matrices, they must satisfy the completeness condition

∑k=0N−1Ak†Ak=I\sum_{k=0}^{N-1}A_{k}^{\dagger}A_{k}=I

This condition guarantees that the trace of the density matrix is preserved.

Choi representations

The Choi representation packages a channel into a single matrix, called Choi matrix and denoted J(Φ)J(\Phi). If the input system has nn basis states and the output system has mm, then J(Φ)J(\Phi) is an (nm)×(nm)(nm)\times(nm) matrix.

Let Φ\Phi be a channel from a system X\mathsf{X} to a system Y\mathsf{Y}, and let Σ\Sigma be the set of basis states of X\mathsf{X}. The Choi matrix is defined by

J(Φ)=∑a,b∈Σ∣a⟩⟨b∣⊗Φ(∣a⟩⟨b∣)J(\Phi)=\sum_{a,b\in\Sigma}\textcolor{#0284c7}{|a\rangle\langle b|}\otimes\textcolor{#7c3aed}{\Phi(|a\rangle\langle b|)}

This definition has a simple interpretation. The operators ∣a⟩⟨b∣|a\rangle\langle b| form a basis for the space of all n×nn\times n matrices, so every matrix is a combination of them. For a qubit there are four of them: ∣0⟩⟨0∣|0\rangle\langle0|, ∣0⟩⟨1∣|0\rangle\langle1|, ∣1⟩⟨0∣|1\rangle\langle0| and ∣1⟩⟨1∣|1\rangle\langle1|.

Notice that only the second factor of each term passes through the channel. The first is the original basis operator, left exactly as it was, and it is there to say which input produced the output beside it. Each term is therefore a record of one input and what Φ\Phi did to it, and J(Φ)J(\Phi) packs all n2n^{2} of those records into a single matrix.

Taking Σ={0,…,n−1}\Sigma=\{0,\ldots,n-1\}, the Choi matrix can therefore be viewed as an n×nn\times n block matrix:

J(Φ)=(Φ(∣0⟩⟨0∣)Φ(∣0⟩⟨1∣)⋯Φ(∣0⟩⟨n−1∣)Φ(∣1⟩⟨0∣)Φ(∣1⟩⟨1∣)⋯Φ(∣1⟩⟨n−1∣)⋮⋮⋱⋮Φ(∣n−1⟩⟨0∣)Φ(∣n−1⟩⟨1∣)⋯Φ(∣n−1⟩⟨n−1∣))J(\Phi)=\begin{pmatrix}\Phi(|0\rangle\langle0|)&\Phi(|0\rangle\langle1|)&\cdots&\Phi(|0\rangle\langle n-1|)\\\Phi(|1\rangle\langle0|)&\Phi(|1\rangle\langle1|)&\cdots&\Phi(|1\rangle\langle n-1|)\\\vdots&\vdots&\ddots&\vdots\\\Phi(|n-1\rangle\langle0|)&\Phi(|n-1\rangle\langle1|)&\cdots&\Phi(|n-1\rangle\langle n-1|)\end{pmatrix}

Because these basis operators span all matrices, the blocks of J(Φ)J(\Phi) completely determine the channel. In other words, the representation is faithful:

J(Φ)=J(Ψ)  ⟺  Φ=ΨJ(\Phi)=J(\Psi)\iff\Phi=\Psi

The Choi matrix also turns the conditions for being a valid quantum channel into simple matrix conditions. A map is a quantum channel exactly when its Choi matrix is positive semidefinite, which expresses complete positivity, and has the correct partial trace, which is what preserves the trace of the state:

J(Φ)≥0complete positivityTrY(J(Φ))=IXtrace preservation\begin{aligned}J(\Phi)&\ge0&&\quad\textcolor{#94a3b8}{\text{complete positivity}}\\[6pt]\mathrm{Tr}_{\mathsf{Y}}\bigl(J(\Phi)\bigr)&=I_{\mathsf{X}}&&\quad\textcolor{#94a3b8}{\text{trace preservation}}\end{aligned}

One important distinction is that J(Φ)J(\Phi) is a representation of the channel, not the channel itself. An ordinary matrix such as UU acts directly on a state by matrix multiplication, for example UρU†U\rho U^{\dagger}. The Choi matrix works differently: its blocks contain the outputs Φ(∣a⟩⟨b∣)\Phi(|a\rangle\langle b|) for all basis operators ∣a⟩⟨b∣|a\rangle\langle b|. In this sense, J(Φ)J(\Phi) records how the channel acts rather than applying the channel to a state. Since those basis operators span all matrices, the complete action of Φ\Phi can be reconstructed from J(Φ)J(\Phi).

There is also a useful way to turn the Choi matrix into a quantum state. Divide it by the dimension of the input system, n=∣Σ∣n=|\Sigma|, to get J(Φ)/nJ(\Phi)/n. This is a density matrix, called the Choi state of Φ\Phi, and it has a direct physical interpretation. Take two copies of the input system, prepare them in the maximally entangled state, and write out its density matrix:

∣ψ⟩=1n∑a∈Σ∣a⟩⊗∣a⟩∣ψ⟩⟨ψ∣=1n∑a,b∈Σ∣a⟩⟨b∣⊗∣a⟩⟨b∣\begin{aligned}|\psi\rangle&=\frac{1}{\sqrt n}\sum_{a\in\Sigma}|a\rangle\otimes|a\rangle\\[6pt]|\psi\rangle\langle\psi|&=\frac1n\sum_{a,b\in\Sigma}|a\rangle\langle b|\otimes|a\rangle\langle b|\end{aligned}

Now apply the channel Φ\Phi to the second system, while leaving the first system unchanged. The identity channel Id\mathrm{Id} represents doing nothing to the first system:

(Id⊗Φ)(∣ψ⟩⟨ψ∣)=1n∑a,b∈Σ∣a⟩⟨b∣⊗Φ(∣a⟩⟨b∣)=J(Φ)n(\mathrm{Id}\otimes\Phi)(|\psi\rangle\langle\psi|)=\frac1n\sum_{a,b\in\Sigma}|a\rangle\langle b|\otimes\Phi(|a\rangle\langle b|)=\frac{J(\Phi)}{n}

Thus, the Choi state is exactly the state produced by applying Φ\Phi to one half of a maximally entangled pair. The Choi matrix is therefore not just an abstract way to represent the channel: after normalization, it describes the physical state that results from this experiment.

∣ψ⟩⟨ψ∣|\psi\rangle\langle\psi|
Φ\Phi
XXY
J(Φ)n\frac{J(\Phi)}{n}

Channel examples in three representations

Channel

Action

Δ(ρ)=∣0⟩⟨0∣ρ∣0⟩⟨0∣+∣1⟩⟨1∣ρ∣1⟩⟨1∣\Delta(\rho)=|0\rangle\langle0|\rho|0\rangle\langle0|+|1\rangle\langle1|\rho|1\rangle\langle1|

Stinespring

Two circuits for it:
ρ\rho
∣0⟩|0\rangle
Δ(ρ)\Delta(\rho)

Copy the qubit into a fresh |0⟩ with a CNOT, then discard the copy.

Follow the matrices through

Write ρab=⟨a∣ρ∣b⟩\rho_{ab}=\langle a|\rho|b\rangle for the entries of the input state. The workspace starts in ∣0⟩|0\rangle, so the joint state entering the unitary is

ρ⊗∣0⟩⟨0∣=(ρ00ρ01ρ10ρ11)⊗(1000)=(ρ00(1000)ρ01(1000)ρ10(1000)ρ11(1000))=(ρ000ρ0100000ρ100ρ1100000)\begin{aligned}\rho\otimes|0\rangle\langle0|&=\begin{pmatrix}\rho_{00}&\rho_{01}\\\rho_{10}&\rho_{11}\end{pmatrix}\otimes\begin{pmatrix}1&0\\0&0\end{pmatrix}\\[8pt]&=\begin{pmatrix}\rho_{00}\begin{pmatrix}1&0\\0&0\end{pmatrix}&\rho_{01}\begin{pmatrix}1&0\\0&0\end{pmatrix}\\[8pt]\rho_{10}\begin{pmatrix}1&0\\0&0\end{pmatrix}&\rho_{11}\begin{pmatrix}1&0\\0&0\end{pmatrix}\end{pmatrix}\\[8pt]&=\begin{pmatrix}\rho_{00}&0&\rho_{01}&0\\0&0&0&0\\\rho_{10}&0&\rho_{11}&0\\0&0&0&0\end{pmatrix}\end{aligned}

The controlled-NOT uses the input as its control and the workspace as its target. It leaves ∣00⟩|00\rangle and ∣01⟩|01\rangle unchanged, while swapping ∣10⟩|10\rangle with ∣11⟩|11\rangle. Since it only permutes basis states, multiplying on the left exchanges the last two rows of the joint state and multiplying on the right exchanges the last two columns. Between them they move the off-diagonal entries into different workspace sectors:

U(ρ⊗∣0⟩⟨0∣)U†=(1000010000010010)(ρ000ρ0100000ρ100ρ1100000)(1000010000010010)=(ρ000ρ01000000000ρ100ρ110)(1000010000010010)=(ρ0000ρ0100000000ρ1000ρ11)\begin{aligned}U(\rho\otimes|0\rangle\langle0|)U^{\dagger}&=\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}\begin{pmatrix}\rho_{00}&0&\rho_{01}&0\\0&0&0&0\\\rho_{10}&0&\rho_{11}&0\\0&0&0&0\end{pmatrix}\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}\\[8pt]&=\begin{pmatrix}\rho_{00}&0&\rho_{01}&0\\0&0&0&0\\0&0&0&0\\\rho_{10}&0&\rho_{11}&0\end{pmatrix}\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}\\[8pt]&=\begin{pmatrix}\rho_{00}&0&0&\rho_{01}\\0&0&0&0\\0&0&0&0\\\rho_{10}&0&0&\rho_{11}\end{pmatrix}\end{aligned}

To read that back in terms of the two systems, note that the rows and columns are indexed by ∣00⟩|00\rangle, ∣01⟩|01\rangle, ∣10⟩|10\rangle and ∣11⟩|11\rangle, with the input first and the workspace second, so the entry in row ∣ij⟩|ij\rangle and column ∣kl⟩|kl\rangle multiplies ∣ij⟩⟨kl∣=∣i⟩⟨k∣⊗∣j⟩⟨l∣|ij\rangle\langle kl|=|i\rangle\langle k|\otimes|j\rangle\langle l|. The four surviving entries sit where both indices are 0000 or 1111, so the state is

+ρ00 ∣0⟩⟨0∣⊗∣0⟩⟨0∣+ρ01 ∣0⟩⟨1∣⊗∣0⟩⟨1∣+ρ10 ∣1⟩⟨0∣⊗∣1⟩⟨0∣+ρ11 ∣1⟩⟨1∣⊗∣1⟩⟨1∣\begin{aligned}&\phantom{{}+{}}\rho_{00}\,|0\rangle\langle0|\otimes|0\rangle\langle0|\\&+\rho_{01}\,|0\rangle\langle1|\otimes|0\rangle\langle1|\\&+\rho_{10}\,|1\rangle\langle0|\otimes|1\rangle\langle0|\\&+\rho_{11}\,|1\rangle\langle1|\otimes|1\rangle\langle1|\end{aligned}

Now discard the workspace by taking the partial trace over the second system. For each term, the workspace factor contributes its trace. The diagonal projectors ∣0⟩⟨0∣|0\rangle\langle0| and ∣1⟩⟨1∣|1\rangle\langle1| have trace 1, while the off-diagonal operators ∣0⟩⟨1∣|0\rangle\langle1| and ∣1⟩⟨0∣|1\rangle\langle0| have trace 0. Thus the two coherence terms vanish:

TrG(U(ρ⊗∣0⟩⟨0∣)U†)=ρ00 ∣0⟩⟨0∣+ρ11 ∣1⟩⟨1∣=Δ(ρ)\mathrm{Tr}_{\mathsf{G}}\bigl(U(\rho\otimes|0\rangle\langle0|)U^{\dagger}\bigr)=\rho_{00}\,|0\rangle\langle0|+\rho_{11}\,|1\rangle\langle1|=\Delta(\rho)

The workspace has therefore recorded which computational-basis state the input occupied. Discarding that workspace removes the corresponding off-diagonal terms while leaving the diagonal probabilities unchanged — exactly the action of the dephasing channel.

Kraus

Two sets of operators:
A0=∣0⟩⟨0∣=(1000)A_{0}=|0\rangle\langle0|=\begin{pmatrix}1&0\\0&0\end{pmatrix}
A1=∣1⟩⟨1∣=(0001)A_{1}=|1\rangle\langle1|=\begin{pmatrix}0&0\\0&1\end{pmatrix}

The two projectors onto the classical states.

Expand the Kraus sum

Take A0=∣0⟩⟨0∣A_{0}=|0\rangle\langle0| and A1=∣1⟩⟨1∣A_{1}=|1\rangle\langle1|, the two projectors onto the classical states. Each keeps one component of the input and returns it in the same direction, so the two terms in the Kraus sum are

∑k=01Ak ρ Ak†=∣0⟩⟨0∣ρ∣0⟩⟨0∣+∣1⟩⟨1∣ρ∣1⟩⟨1∣\sum_{k=0}^{1}A_{k}\,\rho\,A_{k}^{\dagger}=|0\rangle\langle0|\rho|0\rangle\langle0|+|1\rangle\langle1|\rho|1\rangle\langle1|

Now use the inner products in the middle:

=⟨0∣ρ∣0⟩ ∣0⟩⟨0∣+⟨1∣ρ∣1⟩ ∣1⟩⟨1∣=Δ(ρ)\begin{aligned}&=\langle0|\rho|0\rangle\,|0\rangle\langle0|+\langle1|\rho|1\rangle\,|1\rangle\langle1|\\[6pt]&=\Delta(\rho)\end{aligned}

The two coefficients are the diagonal entries of ρ\rho, so the probabilities of measuring 0 and 1 survive unchanged. Nothing in the sum carries the off-diagonal entries across, which is the coherence being lost. The completeness condition holds as well:

∑k=01Ak†Ak=∣0⟩⟨0∣0⟩⟨0∣+∣1⟩⟨1∣1⟩⟨1∣=∣0⟩⟨0∣+∣1⟩⟨1∣=I\begin{aligned}\sum_{k=0}^{1}A_{k}^{\dagger}A_{k}&=|0\rangle\langle0|0\rangle\langle0|+|1\rangle\langle1|1\rangle\langle1|\\[6pt]&=|0\rangle\langle0|+|1\rangle\langle1|\\[6pt]&=I\end{aligned}

Choi

J(Φ)=(1000000000000001)J(\Phi)=\begin{pmatrix}1&0&0&0\\0&0&0&0\\0&0&0&0\\0&0&0&1\end{pmatrix}

Only the diagonal blocks survive, which is the coherence going away.

Build the Choi matrix

The sum runs over the four basis operators of a qubit. In each term, the first factor records the input operator, while the second is that operator after passing through the channel. Each term therefore pairs an input with its output:

J(Δ)=∑a,b=01∣a⟩⟨b∣⊗Δ(∣a⟩⟨b∣)=∣0⟩⟨0∣⊗Δ(∣0⟩⟨0∣)+∣0⟩⟨1∣⊗Δ(∣0⟩⟨1∣)+∣1⟩⟨0∣⊗Δ(∣1⟩⟨0∣)+∣1⟩⟨1∣⊗Δ(∣1⟩⟨1∣)\begin{aligned}J(\Delta)&=\sum_{a,b=0}^{1}\textcolor{#0284c7}{|a\rangle\langle b|}\otimes\textcolor{#7c3aed}{\Delta(|a\rangle\langle b|)}\\[6pt]&=\textcolor{#0284c7}{|0\rangle\langle0|}\otimes\textcolor{#7c3aed}{\Delta(|0\rangle\langle0|)}+\textcolor{#0284c7}{|0\rangle\langle1|}\otimes\textcolor{#7c3aed}{\Delta(|0\rangle\langle1|)}+\textcolor{#0284c7}{|1\rangle\langle0|}\otimes\textcolor{#7c3aed}{\Delta(|1\rangle\langle0|)}+\textcolor{#0284c7}{|1\rangle\langle1|}\otimes\textcolor{#7c3aed}{\Delta(|1\rangle\langle1|)}\end{aligned}

Dephasing leaves the two diagonal operators unchanged and sends the two off-diagonal operators to zero, so the middle two terms vanish and only two survive:

=∣0⟩⟨0∣⊗∣0⟩⟨0∣+∣0⟩⟨1∣⊗0+∣1⟩⟨0∣⊗0+∣1⟩⟨1∣⊗∣1⟩⟨1∣=∣0⟩⟨0∣⊗∣0⟩⟨0∣+∣1⟩⟨1∣⊗∣1⟩⟨1∣\begin{aligned}&=\textcolor{#0284c7}{|0\rangle\langle0|}\otimes\textcolor{#7c3aed}{|0\rangle\langle0|}+\textcolor{#0284c7}{|0\rangle\langle1|}\otimes\textcolor{#7c3aed}{0}+\textcolor{#0284c7}{|1\rangle\langle0|}\otimes\textcolor{#7c3aed}{0}+\textcolor{#0284c7}{|1\rangle\langle1|}\otimes\textcolor{#7c3aed}{|1\rangle\langle1|}\\[6pt]&=\textcolor{#0284c7}{|0\rangle\langle0|}\otimes\textcolor{#7c3aed}{|0\rangle\langle0|}+\textcolor{#0284c7}{|1\rangle\langle1|}\otimes\textcolor{#7c3aed}{|1\rangle\langle1|}\end{aligned}

As a block matrix, each block is the output of the channel on one basis operator:

J(Δ)=(Δ(1000)Δ(0100)Δ(0010)Δ(0001))=(1000000000000001)\begin{aligned}J(\Delta)&=\begin{pmatrix}\Delta\begin{pmatrix}1&0\\0&0\end{pmatrix}&\Delta\begin{pmatrix}0&1\\0&0\end{pmatrix}\\[10pt]\Delta\begin{pmatrix}0&0\\1&0\end{pmatrix}&\Delta\begin{pmatrix}0&0\\0&1\end{pmatrix}\end{pmatrix}\\[10pt]&=\begin{pmatrix}1&0&0&0\\0&0&0&0\\0&0&0&0\\0&0&0&1\end{pmatrix}\end{aligned}

Both conditions from the Choi representation can now be read directly from this matrix. It is diagonal with non-negative entries, so it is positive semidefinite, and tracing out the output system returns the identity, so the channel is trace-preserving:

TrY(J(Δ))=∣0⟩⟨0∣+∣1⟩⟨1∣=I\mathrm{Tr}_{\mathsf{Y}}\bigl(J(\Delta)\bigr)=|0\rangle\langle0|+|1\rangle\langle1|=I

The three representations are equivalent

The three forms are different ways of describing the same quantum channel. This equivalence holds for any quantum channel: a channel written in one representation can always be transformed into either of the other two, and the transformations can be reversed. Thus, the choice of representation changes how the channel is expressed, but not the information it contains.

Definition

Φ is a channel from X to Y: a linear map transforming density matrices to density matrices.

Choi representation
J(Φ)≥0J(\Phi)\ge0TrY(J(Φ))=IX\mathrm{Tr}_{\mathsf{Y}}(J(\Phi))=I_{\mathsf{X}}
Kraus representation
Φ(ρ)=∑kAkρAk†\Phi(\rho)=\sum_{k}A_{k}\rho A_{k}^{\dagger}∑kAk†Ak=I\sum_{k}A_{k}^{\dagger}A_{k}=I
Stinespring representationρ|0⟩UΦ(ρ)
Definition

Φ is a channel from X to Y: a linear map transforming density matrices to density matrices.

Choi representation
J(Φ)≥0J(\Phi)\ge0TrY(J(Φ))=IX\mathrm{Tr}_{\mathsf{Y}}(J(\Phi))=I_{\mathsf{X}}
Kraus representation
Φ(ρ)=∑kAkρAk†\Phi(\rho)=\sum_{k}A_{k}\rho A_{k}^{\dagger}∑kAk†Ak=I\sum_{k}A_{k}^{\dagger}A_{k}=I
Stinespring representationρ|0⟩UΦ(ρ)

Definition → Choi representation

A channel satisfies the Choi conditions

Let Φ\Phi be a quantum channel acting on an nn-dimensional system. Its Choi matrix pairs each input basis operator with its output:

J(Φ)=∑a,b=0n−1∣a⟩⟨b∣⊗Φ(∣a⟩⟨b∣)J(\Phi)=\sum_{a,b=0}^{n-1}\textcolor{#0284c7}{|a\rangle\langle b|}\otimes\textcolor{#7c3aed}{\Phi(|a\rangle\langle b|)}

To see what properties this matrix must have, take two copies of the system, X\mathsf{X} and Y\mathsf{Y}, prepare them in the , and write out its density matrix:

∣Ω⟩=1n∑a=0n−1∣a⟩X∣a⟩Y∣Ω⟩⟨Ω∣=1n∑a,b=0n−1∣a⟩⟨b∣X⊗∣a⟩⟨b∣Y\begin{aligned}|\Omega\rangle&=\frac{1}{\sqrt n}\sum_{a=0}^{n-1}|a\rangle_{\mathsf{X}}|a\rangle_{\mathsf{Y}}\\[6pt]|\Omega\rangle\langle\Omega|&=\frac1n\sum_{a,b=0}^{n-1}|a\rangle\langle b|_{\mathsf{X}}\otimes|a\rangle\langle b|_{\mathsf{Y}}\end{aligned}

Now apply Φ\Phi to Y\mathsf{Y} while leaving X\mathsf{X} unchanged. The map Id⊗Φ\mathrm{Id}\otimes\Phi is linear, so it acts on the terms of the sum one at a time, leaving every first factor alone and sending every second factor through the channel. What comes back is the normalized Choi state:

(Id⊗Φ)(∣Ω⟩⟨Ω∣)=(Id⊗Φ) ⁣(1n∑a,b=0n−1∣a⟩⟨b∣X⊗∣a⟩⟨b∣Y)=1n∑a,b=0n−1Id(∣a⟩⟨b∣X)⊗Φ(∣a⟩⟨b∣Y)=1n∑a,b=0n−1∣a⟩⟨b∣X⊗Φ(∣a⟩⟨b∣Y)=J(Φ)n\begin{aligned}(\mathrm{Id}\otimes\Phi)(|\Omega\rangle\langle\Omega|)&=(\mathrm{Id}\otimes\Phi)\!\left(\frac1n\sum_{a,b=0}^{n-1}|a\rangle\langle b|_{\mathsf{X}}\otimes|a\rangle\langle b|_{\mathsf{Y}}\right)\\[6pt]&=\frac1n\sum_{a,b=0}^{n-1}\textcolor{#0284c7}{\mathrm{Id}(|a\rangle\langle b|_{\mathsf{X}})}\otimes\textcolor{#7c3aed}{\Phi(|a\rangle\langle b|_{\mathsf{Y}})}\\[6pt]&=\frac1n\sum_{a,b=0}^{n-1}\textcolor{#0284c7}{|a\rangle\langle b|_{\mathsf{X}}}\otimes\textcolor{#7c3aed}{\Phi(|a\rangle\langle b|_{\mathsf{Y}})}\\[6pt]&=\frac{J(\Phi)}{n}\end{aligned}

Because Φ\Phi is a quantum channel, this output must be a valid density matrix. In particular it must be positive semidefinite:

J(Φ)≥0J(\Phi)\ge0

There is also a condition on the system X\mathsf{X}. The channel acts only on Y\mathsf{Y}, so it cannot change the state of X\mathsf{X}, which was maximally mixed before the channel and stays that way after it. Multiplying by nn then clears the factor:

TrY(∣Ω⟩⟨Ω∣)=IXnTrY ⁣(J(Φ)n)=IXnTrYJ(Φ)=IX\begin{aligned}\mathrm{Tr}_{\mathsf{Y}}\bigl(|\Omega\rangle\langle\Omega|\bigr)&=\frac{I_{\mathsf{X}}}{n}\\[6pt]\mathrm{Tr}_{\mathsf{Y}}\!\left(\frac{J(\Phi)}{n}\right)&=\frac{I_{\mathsf{X}}}{n}\\[6pt]\mathrm{Tr}_{\mathsf{Y}}J(\Phi)&=I_{\mathsf{X}}\end{aligned}

Thus every quantum channel produces a Choi matrix satisfying the two conditions

J(Φ)≥0,TrYJ(Φ)=IXJ(\Phi)\ge0,\qquad\mathrm{Tr}_{\mathsf{Y}}J(\Phi)=I_{\mathsf{X}}

These conditions are also sufficient: any matrix satisfying them is the Choi matrix of a valid quantum channel.

Three representations are a lot of machinery for one channel, and at this point I’m not even sure the payoff is worth all the setup. For simple channels like reset, the original definition is often clearer.

The point is that each view is useful for a different purpose: Stinespring gives the physical picture, Kraus is convenient for calculations, and Choi turns the channel into a matrix. The useful part is knowing that the same channel can be viewed in whichever form makes a particular problem easier.

General measurements

Measurements are the interface between quantum and classical information. Performing a measurement extracts classical information from a quantum state and, in general, changes or destroys the system in the process.

Destructive measurements produce only a classical outcome. What happens to the system afterwards is not part of the description. This means that only the outcome probabilities need to be described, without specifying the state of the system after the measurement.

There are two equivalent ways to describe a destructive measurement, both of which will be useful below. The first is a collection of matrices, one for each measurement outcome. The second is a channel whose outputs are always classical states, represented by diagonal density matrices. The two descriptions contain exactly the same information, but each is convenient in a different setting: the matrices for calculating probabilities, and the channel when the measurement appears in a circuit alongside other operations.

Non-destructive measurements, where the system survives and its post-measurement state matters, do not require a separate theory: any such measurement can be described as a destructive measurement followed by a channel that prepares the resulting post-measurement state. It is therefore useful to get the destructive case right first, and to once it is in place.

Measurements as matrices

Start with the physical question: what information do we need to describe a destructive measurement? For each possible outcome aa, we need something that tells us how likely that outcome is for a given state ρ\rho. Let that something be a matrix PaP_{a}, so that Pr(outcome=a)=Tr(Paρ)\mathrm{Pr}(\text{outcome}=a)=\mathrm{Tr}(P_{a}\rho).

The matrices must satisfy a few conditions for these numbers to be valid probabilities. They must give nonnegative values for every quantum state, and all the probabilities must add up to one.

Projective measurements

In a , each outcome aa is associated with a projection matrix Πa\Pi_{a}. The projector Πa\Pi_{a} selects the subspace corresponding to that outcome. Since the measurement must account for every possible outcome, the projectors form a complete decomposition of the identity:

Π0+⋯+Πm−1=IX\Pi_{0}+\cdots+\Pi_{m-1}=I_{\mathsf{X}}

For a pure state ∣ψ⟩\lvert\psi\rangle, the probability of obtaining outcome aa is the squared size of the component of ∣ψ⟩\lvert\psi\rangle in the corresponding subspace:

Pr(outcome=a)=∥Πa∣ψ⟩∥2=⟨ψ∣Πa∣ψ⟩\mathrm{Pr}(\text{outcome}=a)=\bigl\lVert\Pi_{a}\lvert\psi\rangle\bigr\rVert^{2}=\langle\psi\rvert\Pi_{a}\lvert\psi\rangle

How does this number become a trace? Write u=Πa∣ψ⟩u=\Pi_a\lvert\psi\rangle, so the probability is ⟨ψ∣u\langle\psi\rvert u. There are two ways to multiply this column and the row ⟨ψ∣\langle\psi\rvert. Row times column gives a single number, and column times row gives a matrix. Adding that matrix's diagonal gives the same number.

Row times column

⟨ψ∣u=\langle\psi\rvert u=[ψ1‾ψ2‾]⏟1×2\underbrace{\begin{bmatrix}\overline{\psi_1}&\overline{\psi_2}\end{bmatrix}}_{1\times2}[u1u2]⏟2×1\underbrace{\begin{bmatrix}u_1\\u_2\end{bmatrix}}_{2\times1}
⟨ψ∣u=ψ1‾u1+ψ2‾u2=Pr(a)\begin{aligned} \phantom{\langle\psi\rvert u}&=\textcolor{#0284c7}{\overline{\psi_1}u_1}+\textcolor{#4f46e5}{\overline{\psi_2}u_2}\\[2pt] &=\mathrm{Pr}(a) \end{aligned}

Column times row

u⟨ψ∣=[u1u2]⏟2×1[ψ1‾ψ2‾]⏟1×2u\langle\psi\rvert= \underbrace{\begin{bmatrix}u_1\\u_2\end{bmatrix}}_{2\times1} \underbrace{\begin{bmatrix}\overline{\psi_1}&\overline{\psi_2}\end{bmatrix}}_{1\times2}
u⟨ψ∣=\phantom{u\langle\psi\rvert}=
u1ψ1‾\textcolor{#0284c7}{u_1\overline{\psi_1}}
u1ψ2‾\textcolor{#94a3b8}{u_1\overline{\psi_2}}
u2ψ1‾\textcolor{#94a3b8}{u_2\overline{\psi_1}}
u2ψ2‾\textcolor{#4f46e5}{u_2\overline{\psi_2}}
Tr(u⟨ψ∣)=u1ψ1‾+u2ψ2‾=Pr(a)\begin{aligned} \mathrm{Tr}(u\langle\psi\rvert)&=\textcolor{#0284c7}{u_1\overline{\psi_1}}+\textcolor{#4f46e5}{u_2\overline{\psi_2}}\\[2pt] &=\mathrm{Pr}(a) \end{aligned}

The two sums agree because each entry is a scalar: ψi‾ui=uiψi‾\overline{\psi_i}u_i=u_i\overline{\psi_i}. In any dimension, the same calculation reads:

⟨ψ∣u=∑iψi‾uirow times column=∑iuiψi‾scalar factors commute=∑i(u⟨ψ∣)iithese are the diagonal entries=Tr(u⟨ψ∣)trace means sum the diagonal\begin{aligned} \langle\psi\rvert u &=\sum_i\overline{\psi_i}u_i &&\textcolor{#64748b}{\text{row times column}}\\[4pt] &=\sum_i u_i\overline{\psi_i} &&\textcolor{#64748b}{\text{scalar factors commute}}\\[4pt] &=\sum_i\bigl(u\langle\psi\rvert\bigr)_{ii} &&\textcolor{#64748b}{\text{these are the diagonal entries}}\\[4pt] &=\mathrm{Tr}\bigl(u\langle\psi\rvert\bigr) &&\textcolor{#64748b}{\text{trace means sum the diagonal}} \end{aligned}

Substitute u=Πa∣ψ⟩u=\Pi_a\lvert\psi\rangle back in. The outer product ∣ψ⟩⟨ψ∣\lvert\psi\rangle\langle\psi\rvert is exactly the pure state's density matrix:

Pr(outcome=a)=⟨ψ∣Πa∣ψ⟩=Tr ⁣(Πa∣ψ⟩⟨ψ∣⏟ρ)=Tr(Πaρ).\begin{aligned} \mathrm{Pr}(\text{outcome}=a) &=\langle\psi\rvert\Pi_a\lvert\psi\rangle\\[4pt] &=\mathrm{Tr}\!\left(\Pi_a\underbrace{\lvert\psi\rangle\langle\psi\rvert}_{\rho}\right)\\[4pt] &=\mathrm{Tr}(\Pi_a\rho). \end{aligned}

This identity works for any matrix AA in place of Πa\Pi_a: the trace calculation did not use any special property of a projector.

This form immediately extends the rule from pure states to arbitrary density matrices. For a general state ρ\rho, a projective measurement therefore gives

Pr(outcome=a)=Tr(Πaρ)\mathrm{Pr}(\text{outcome}=a)=\mathrm{Tr}(\Pi_{a}\rho)

At this point, the role of the projectors is clear: they specify the possible outcomes, while the trace expression gives their probabilities. The next step is to ask what happens if we keep this probability rule but no longer require the matrices associated with outcomes to be projections.

General measurements

We now keep the rule Pr(outcome=a)=Tr(Paρ)\mathrm{Pr}(\text{outcome}=a)=\mathrm{Tr}(P_a\rho), but allow PaP_a to be more general than a projector. What conditions must these matrices satisfy to give valid probabilities for every state?

  • Each probability must be nonnegative. For a pure state, the identity we just derived gives

    Pr(outcome=a)=Tr(Pa∣ψ⟩⟨ψ∣)=⟨ψ∣Pa∣ψ⟩≥0\mathrm{Pr}(\text{outcome}=a)=\mathrm{Tr}\bigl(P_a\lvert\psi\rangle\langle\psi\rvert\bigr)=\langle\psi\rvert P_a\lvert\psi\rangle\geq0

    Requiring this value to be real and nonnegative for every normalized ∣ψ⟩\lvert\psi\rangle is exactly the condition that PaP_a is positive semidefinite, written Pa≥0P_a\geq0. Mixed states are weighted mixtures of pure states. By linearity of the trace, their probabilities are weighted averages of these nonnegative probabilities, so they are nonnegative too.

  • The probabilities must add up to one. Adding over all outcomes and using linearity again gives

    ∑aPr(outcome=a)=∑aTr(Paρ)=Tr ⁣([∑aPa]ρ)\sum_a\mathrm{Pr}(\text{outcome}=a)=\sum_a\mathrm{Tr}(P_a\rho)=\mathrm{Tr}\!\left(\left[\sum_a P_a\right]\rho\right)

    If the matrices sum to the identity, this becomes Tr(IXρ)=Tr(ρ)=1\mathrm{Tr}(I_{\mathsf X}\rho)=\mathrm{Tr}(\rho)=1. Requiring the total to be one for every state also forces this condition: choosing any normalized eigenvector of ∑aPa\sum_a P_a as the state makes the total equal to its eigenvalue. Every eigenvalue must therefore be one, so

    P0+⋯+Pm−1=IXP_0+\cdots+P_{m-1}=I_{\mathsf X}

A general measurement is therefore a collection of matrices {P0,…,Pm−1}\{P_0,\ldots,P_{m-1}\} satisfying these two conditions, with outcome probabilities given by the same trace rule:

Pa≥0 for every a,∑aPa=IXPr(outcome=a)=Tr(Paρ).\begin{gathered} P_a\geq0\ \text{for every }a,\qquad \sum_a P_a=I_{\mathsf X}\\[6pt] \mathrm{Pr}(\text{outcome}=a)=\mathrm{Tr}(P_a\rho). \end{gathered}

Every projective measurement satisfies these conditions. The new possibilities come from allowing eigenvalues between zero and one, whereas a projector has only zero or one as eigenvalues. The upper bound follows because IX−Pa=∑b≠aPb≥0I_{\mathsf X}-P_a=\sum_{b\ne a}P_b\geq0. For a state that is an eigenvector of PaP_a, the probability of outcome aa is its eigenvalue, which can now lie strictly between zero and one.

A standard basis measurement

A standard basis measurement of a qubit is the collection {P0,P1}\{P_{0},P_{1}\} where

P0=∣0⟩⟨0∣=(1000)P1=∣1⟩⟨1∣=(0001)P_{0}=\lvert0\rangle\langle0\rvert=\begin{pmatrix}1&0\\0&0\end{pmatrix}\qquad P_{1}=\lvert1\rangle\langle1\rvert=\begin{pmatrix}0&0\\0&1\end{pmatrix}

Measuring a qubit in the state ρ\rho gives outcome probabilities that are just the diagonal entries of ρ\rho:

Pr(outcome=0)=Tr(P0ρ)=Tr(∣0⟩⟨0∣ρ)=⟨0∣ρ∣0⟩Pr(outcome=1)=Tr(P1ρ)=Tr(∣1⟩⟨1∣ρ)=⟨1∣ρ∣1⟩\begin{aligned}\mathrm{Pr}(\text{outcome}=0)&=\mathrm{Tr}(P_{0}\rho)=\mathrm{Tr}\bigl(\lvert0\rangle\langle0\rvert\rho\bigr)=\langle0\rvert\rho\lvert0\rangle\\[4pt]\mathrm{Pr}(\text{outcome}=1)&=\mathrm{Tr}(P_{1}\rho)=\mathrm{Tr}\bigl(\lvert1\rangle\langle1\rvert\rho\bigr)=\langle1\rvert\rho\lvert1\rangle\end{aligned}
State
01|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
1.00
Pr(outcome=0)=Tr((1000)⏟P0(1000)⏟ρ)=Tr(1000)=1+0≈1\begin{aligned} \mathrm{Pr}(\text{outcome}=0) &=\mathrm{Tr}\left(\underbrace{\begin{pmatrix}1&0\\0&0\end{pmatrix}}_{P_0}\underbrace{\begin{pmatrix}1&0\\0&0\end{pmatrix}}_{\rho}\right)\\[6pt] &=\mathrm{Tr}\begin{pmatrix}\textcolor{#7c3aed}{1}&0\\0&\textcolor{#7c3aed}{0}\end{pmatrix}\\[6pt] &=\textcolor{#7c3aed}{1}+\textcolor{#7c3aed}{0}\approx\textcolor{#7c3aed}{1} \end{aligned}

Two outcomes without projections

The definition requires only positive semidefinite matrices that sum to the identity. It does not require the measurement operators to be projections. For example, consider

P0=(23131313)P1=(13−13−1323)P_{0}=\begin{pmatrix}\tfrac23&\tfrac13\\[2pt]\tfrac13&\tfrac13\end{pmatrix}\qquad P_{1}=\begin{pmatrix}\tfrac13&-\tfrac13\\[2pt]-\tfrac13&\tfrac23\end{pmatrix}

They sum to the identity, and each has trace 11 and determinant 19\tfrac19, so both eigenvalues are positive and both matrices are positive semidefinite. Neither is a projection, since P02≠P0P_{0}^{2}\neq P_{0} and P12≠P1P_{1}^{2}\neq P_{1}.

For a qubit in the ∣+⟩\lvert+\rangle state, with ∣+⟩⟨+∣=(12121212)\lvert+\rangle\langle+\rvert=\begin{pmatrix}\tfrac12&\tfrac12\\[2pt]\tfrac12&\tfrac12\end{pmatrix}, the outcome probabilities add to 11 even though neither measurement operator is a projection:

Pr(outcome=0)=Tr(P0∣+⟩⟨+∣)=Tr((23131313)(12121212))=56Pr(outcome=1)=Tr(P1∣+⟩⟨+∣)=Tr((13−13−1323)(12121212))=16\begin{aligned}\mathrm{Pr}(\text{outcome}=0)&=\mathrm{Tr}\bigl(P_{0}\lvert+\rangle\langle+\rvert\bigr)=\mathrm{Tr}\left(\begin{pmatrix}\tfrac23&\tfrac13\\[2pt]\tfrac13&\tfrac13\end{pmatrix}\begin{pmatrix}\tfrac12&\tfrac12\\[2pt]\tfrac12&\tfrac12\end{pmatrix}\right)=\tfrac56\\[10pt]\mathrm{Pr}(\text{outcome}=1)&=\mathrm{Tr}\bigl(P_{1}\lvert+\rangle\langle+\rvert\bigr)=\mathrm{Tr}\left(\begin{pmatrix}\tfrac13&-\tfrac13\\[2pt]-\tfrac13&\tfrac23\end{pmatrix}\begin{pmatrix}\tfrac12&\tfrac12\\[2pt]\tfrac12&\tfrac12\end{pmatrix}\right)=\tfrac16\end{aligned}

On the Bloch sphere, the two directions are still opposite, but even perfect alignment with one direction cannot make the corresponding outcome 100% certain.

State
01|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
1.00
Pr(outcome=0)=Tr((0.6670.3330.3330.333)⏟P0(1000)⏟ρ)=Tr(0.66700.3330)=0.667+0≈0.667\begin{aligned} \mathrm{Pr}(\text{outcome}=0) &=\mathrm{Tr}\left(\underbrace{\begin{pmatrix}0.667&0.333\\0.333&0.333\end{pmatrix}}_{P_0}\underbrace{\begin{pmatrix}1&0\\0&0\end{pmatrix}}_{\rho}\right)\\[6pt] &=\mathrm{Tr}\begin{pmatrix}\textcolor{#7c3aed}{0.667}&0\\0.333&\textcolor{#7c3aed}{0}\end{pmatrix}\\[6pt] &=\textcolor{#7c3aed}{0.667}+\textcolor{#7c3aed}{0}\approx\textcolor{#7c3aed}{0.667} \end{aligned}

Four outcomes from a single qubit

The tetrahedral states are four qubit states whose Bloch vectors are arranged symmetrically in three dimensions. On the Bloch sphere, the four vectors point toward the corners of a regular tetrahedron, giving these states their name.

∣ϕ0⟩=∣0⟩∣ϕ1⟩=13∣0⟩+23 ∣1⟩∣ϕ2⟩=13∣0⟩+23 e2πi/3∣1⟩∣ϕ3⟩=13∣0⟩+23 e−2πi/3∣1⟩\begin{aligned}\lvert\phi_{0}\rangle&=\lvert0\rangle\\[4pt]\lvert\phi_{1}\rangle&=\tfrac{1}{\sqrt3}\lvert0\rangle+\sqrt{\tfrac23}\,\lvert1\rangle\\[4pt]\lvert\phi_{2}\rangle&=\tfrac{1}{\sqrt3}\lvert0\rangle+\sqrt{\tfrac23}\,e^{2\pi i/3}\lvert1\rangle\\[4pt]\lvert\phi_{3}\rangle&=\tfrac{1}{\sqrt3}\lvert0\rangle+\sqrt{\tfrac23}\,e^{-2\pi i/3}\lvert1\rangle\end{aligned}

A measurement {P0,P1,P2,P3}\{P_{0},P_{1},P_{2},P_{3}\} can be defined from them by halving each projection:

P0=∣ϕ0⟩⟨ϕ0∣2P1=∣ϕ1⟩⟨ϕ1∣2P2=∣ϕ2⟩⟨ϕ2∣2P3=∣ϕ3⟩⟨ϕ3∣2P_{0}=\frac{\lvert\phi_{0}\rangle\langle\phi_{0}\rvert}{2}\quad P_{1}=\frac{\lvert\phi_{1}\rangle\langle\phi_{1}\rvert}{2}\quad P_{2}=\frac{\lvert\phi_{2}\rangle\langle\phi_{2}\rvert}{2}\quad P_{3}=\frac{\lvert\phi_{3}\rangle\langle\phi_{3}\rvert}{2}

Each is positive semidefinite because it is a projection scaled by a positive number, and the four operators sum to the identity because the four tetrahedral directions cancel out. A two-dimensional system can therefore have four measurement outcomes, something that a projective measurement cannot provide. The trade-off is that the outcomes are no longer perfectly distinguishable. If the four preparations are equally likely beforehand, observing outcome aa makes ∣ϕa⟩\lvert\phi_{a}\rangle the most likely preparation, but it does not prove that it was the one prepared.

State
0123|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
1.00
Pr(outcome=0)=Tr((0.5000)⏟P0(1000)⏟ρ)=Tr(0.5000)=0.5+0≈0.5\begin{aligned} \mathrm{Pr}(\text{outcome}=0) &=\mathrm{Tr}\left(\underbrace{\begin{pmatrix}0.5&0\\0&0\end{pmatrix}}_{P_0}\underbrace{\begin{pmatrix}1&0\\0&0\end{pmatrix}}_{\rho}\right)\\[6pt] &=\mathrm{Tr}\begin{pmatrix}\textcolor{#7c3aed}{0.5}&0\\0&\textcolor{#7c3aed}{0}\end{pmatrix}\\[6pt] &=\textcolor{#7c3aed}{0.5}+\textcolor{#7c3aed}{0}\approx\textcolor{#7c3aed}{0.5} \end{aligned}

Measurements as channels

Classical probabilistic states are represented by diagonal density matrices, with the probabilities along the diagonal. That is the key idea behind the second description: a measurement is a channel whose output is always a classical state.

Any general measurement can therefore be described by a channel Φ\Phi. The input system X\mathsf{X} is the system being measured, and the output system Y\mathsf{Y} is classical, with states corresponding to the possible measurement outcomes {0,…,m−1}\{0,\ldots,m-1\}. For every input state ρ\rho of X\mathsf{X}, the output Φ(ρ)\Phi(\rho) is a diagonal density matrix whose diagonal entries are the outcome probabilities.

The completely dephasing channel Δ\Delta describes a standard basis measurement of a qubit:

Δ(ρ)=⟨0∣ρ∣0⟩ ∣0⟩⟨0∣+⟨1∣ρ∣1⟩ ∣1⟩⟨1∣\Delta(\rho)=\langle0\rvert\rho\lvert0\rangle\,\lvert0\rangle\langle0\rvert+\langle1\rvert\rho\lvert1\rangle\,\lvert1\rangle\langle1\rvert

The channel keeps the diagonal entries of ρ\rho and erases everything else. The result contains the two outcome probabilities, but no information about the coherence between them. The same channel also describes dephasing noise, which is what happens when the environment effectively measures a system in the standard basis without the outcome being read.

Equivalence to the matrix description

A channel Φ\Phi from X\mathsf{X} to Y\mathsf{Y} has diagonal output for every input state if and only if there is a {P0,…,Pm−1}\{P_{0},\ldots,P_{m-1}\} for which:

Φ(ρ)=∑a=0m−1Tr(Paρ) ∣a⟩⟨a∣\Phi(\rho)=\sum_{a=0}^{m-1}\mathrm{Tr}(P_{a}\rho)\,\lvert a\rangle\langle a\rvert

One direction is immediate. Given the measurement matrices, the formula defines a channel whose output is diagonal by construction.

For the other direction, suppose Φ(ρ)\Phi(\rho) is always diagonal. Its aa-th diagonal entry, ⟨a∣Φ(ρ)∣a⟩\langle a\rvert\Phi(\rho)\lvert a\rangle, is a linear function of ρ\rho, because Φ\Phi is linear. Any linear function of a matrix can be written as a trace against a fixed matrix, so there is a matrix PaP_{a} for which:

⟨a∣Φ(ρ)∣a⟩=Tr(Paρ)\langle a\rvert\Phi(\rho)\lvert a\rangle=\mathrm{Tr}(P_{a}\rho)

The required properties of the PaP_{a} then follow from the corresponding properties of the channel. Since Φ(ρ)\Phi(\rho) is a density matrix, its diagonal entries are nonnegative for every density matrix ρ\rho, which means that each PaP_{a} is positive semidefinite. Those entries also sum to one for every ρ\rho:

∑a=0m−1Tr(Paρ)=1\sum_{a=0}^{m-1}\mathrm{Tr}(P_{a}\rho)=1

A single matrix has that trace against every density matrix only if it is the identity, so the measurement matrices must add up to it:

P0+⋯+Pm−1=IXP_{0}+\cdots+P_{m-1}=I_{\mathsf{X}}

So the two descriptions contain exactly the same information. The matrices are convenient when the goal is to calculate an outcome probability. The channel is convenient when the measurement needs to appear in a circuit alongside other operations, since it composes with them like any other channel.

Partial measurements

A measurement channel replaces a system with a classical record of its outcome. Suppose now that the measured system X\mathsf{X} is only half of a pair (X,Z)(\mathsf{X},\mathsf{Z}) held in the joint state ρ\rho, and that the measurement {P0,…,Pm−1}\{P_0,\ldots,P_{m-1}\} acts on X\mathsf{X} alone. Nothing touches Z\mathsf{Z}.

(X,Z)→ Φ⊗IdZ (Y,Z)(\mathsf{X},\mathsf{Z})\xrightarrow{\ \Phi\otimes\mathrm{Id}_{\mathsf{Z}}\ }(\mathsf{Y},\mathsf{Z})

The setup

ρ\rho
XYZZ

Only the upper half of the pair enters the meter. Its outcome leaves on the double wire as the classical register Y, and Z runs past untouched.

Two questions follow: which outcome aa appears, and once aa appears, what state is Z\mathsf{Z} in?

What is the probability of each outcome

The first question depends only on the measured system X\mathsf{X}. Its , ρX=TrZ(ρ)\rho_{\mathsf{X}}=\mathrm{Tr}_{\mathsf{Z}}(\rho), contains everything needed to determine the probabilities of the measurement outcomes:

pa=Tr(Pa ρX)p_a=\mathrm{Tr}\bigl(P_a\,\textcolor{#0284c7}{\rho_{\mathsf{X}}}\bigr)
pa=Tr((Pa⊗IZ)ρ)\phantom{p_a}=\mathrm{Tr}\bigl((P_a\otimes I_{\mathsf{Z}})\rho\bigr)

Where the probability comes from

ρ\rho

joint state of the pair

trace out Z

ρX\textcolor{#0284c7}{\rho_{\mathsf{X}}}

local state of the measured system

measure with PaP_a

pa\textcolor{#0284c7}{p_a}

probability of outcome a

The two expressions describe the same calculation in two equivalent ways. Either first reduce the joint state to X\mathsf{X} and then measure it, or evaluate the measurement directly on the pair while leaving Z\mathsf{Z} untouched. In both cases, the measurement acts only on X\mathsf{X}.

Taking X\mathsf{X} and Z\mathsf{Z} to be qubits shows what that first step does. Cut ρ\rho into four 2×22\times2 blocks, one for each pair of basis states of X\mathsf{X}, and the partial trace acts on each block separately.

ρ=(ρ00ρ01ρ02ρ03ρ10ρ11ρ12ρ13ρ20ρ21ρ22ρ23ρ30ρ31ρ32ρ33)=(C00C01C10C11)\rho=\left(\begin{array}{cc|cc} \rho_{00}&\rho_{01}&\rho_{02}&\rho_{03}\\\rho_{10}&\rho_{11}&\rho_{12}&\rho_{13}\\\hline\rho_{20}&\rho_{21}&\rho_{22}&\rho_{23}\\\rho_{30}&\rho_{31}&\rho_{32}&\rho_{33} \end{array}\right) =\begin{pmatrix}C_{00}&C_{01}\\ C_{10}&C_{11}\end{pmatrix}

Tracing out Z

XZ0001101100011011
C00C_{00}
C01C_{01}
C10C_{10}
C11C_{11}
TrZ\mathrm{Tr}_{\mathsf{Z}}
ρX=TrZ(ρ)=(Tr(C00)Tr(C01)Tr(C10)Tr(C11))\textcolor{#0284c7}{\rho_{\mathsf{X}}}=\mathrm{Tr}_{\mathsf{Z}}(\rho)=\begin{pmatrix} \mathrm{Tr}(C_{00})&\mathrm{Tr}(C_{01})\\ \mathrm{Tr}(C_{10})&\mathrm{Tr}(C_{11})\end{pmatrix}
X0101
Tr(C00)\mathrm{Tr}(C_{00})
Tr(C01)\mathrm{Tr}(C_{01})
Tr(C10)\mathrm{Tr}(C_{10})
Tr(C11)\mathrm{Tr}(C_{11})

So, the reduced state ρX\rho_{\mathsf{X}} keeps one number from each block and drops everything else inside it. That is enough for the outcome probabilities. But ρX\rho_{\mathsf{X}} cannot tell us what state Z\mathsf{Z} is left in. That depends on the correlations between X\mathsf{X} and Z\mathsf{Z}, which are part of the joint state ρ\rho and are lost when Z\mathsf{Z} is traced out. To answer the second question, we therefore need the joint state again.

The state Z, conditioned on the outcome a

The second question depends on the whole joint state ρ\rho. When X\mathsf{X} is measured and outcome aa occurs, the pair is projected onto the corresponding subspace of X\mathsf{X}, and what remains is the state that Z\mathsf{Z} is left in:

σa=TrX((Pa⊗IZ) ρ)ρZ∣a=σapa\sigma_a=\mathrm{Tr}_{\mathsf{X}}\bigl((P_a\otimes I_{\mathsf{Z}})\,\rho\bigr) \qquad \rho_{\mathsf{Z}\mid a}=\frac{\sigma_a}{p_a}

The first expression is the unnormalized conditional state: a positive operator on Z\mathsf{Z} whose trace is exactly the probability of the outcome, Tr(σa)=pa\mathrm{Tr}(\sigma_a)=p_a. Dividing by pap_a renormalizes it to a proper state. This is the same calculation as before, run in reverse — instead of tracing out Z\mathsf{Z} to ask what we will see, we trace out X\mathsf{X} to ask what is left behind.

The two trace operations do different jobs, and keeping them apart is the whole point:

Where the conditional state comes from

ρ\rho

joint state of the pair

Pa⊗IZ\textcolor{#0284c7}{P_a\otimes I_{\mathsf{Z}}}

project X onto outcome a, leave Z untouched

σa\textcolor{#059669}{\sigma_a}

unnormalized state of Z (trace = p_a)

÷ pa\div\,p_a

ρZ∣a\textcolor{#c2410c}{\rho_{\mathsf{Z}\mid a}}

the state Z is left in, given outcome a

Taking X\mathsf{X} and Z\mathsf{Z} to be qubits shows what this projection does to the block structure. The projector PaP_a selects one row and one column of the block matrix, the part of ρ\rho where X\mathsf{X} sits in outcome aa, and the partial trace over X\mathsf{X} then collapses the selected blocks onto Z\mathsf{Z}.

Writing the blocks of ρ\rho as CijC_{ij} as before, take the simplest case of a standard basis measurement on X\mathsf{X}, with P0=∣0⟩⟨0∣P_0=\lvert0\rangle\langle0\rvert and P1=∣1⟩⟨1∣P_1=\lvert1\rangle\langle1\rvert:

Outcome 0

XZ0001101100011011
C00\textcolor{#059669}{C_{00}}
C01\textcolor{#94a3b8}{C_{01}}
0\textcolor{#94a3b8}{0}
0\textcolor{#94a3b8}{0}
σ0=TrX((P0⊗IZ) ρ)=C00ρZ∣0=C00Tr(C00)\begin{aligned} \sigma_0&=\mathrm{Tr}_{\mathsf{X}}\bigl((P_0\otimes I_{\mathsf{Z}})\,\rho\bigr)=\textcolor{#059669}{C_{00}}\\ \rho_{\mathsf{Z}\mid 0}&=\frac{C_{00}}{\mathrm{Tr}(C_{00})} \end{aligned}

Outcome 1

XZ0001101100011011
0\textcolor{#94a3b8}{0}
0\textcolor{#94a3b8}{0}
C10\textcolor{#94a3b8}{C_{10}}
C11\textcolor{#059669}{C_{11}}
σ1=TrX((P1⊗IZ) ρ)=C11ρZ∣1=C11Tr(C11)\begin{aligned} \sigma_1&=\mathrm{Tr}_{\mathsf{X}}\bigl((P_1\otimes I_{\mathsf{Z}})\,\rho\bigr)=\textcolor{#059669}{C_{11}}\\ \rho_{\mathsf{Z}\mid 1}&=\frac{C_{11}}{\mathrm{Tr}(C_{11})} \end{aligned}

In the block picture, measuring X\mathsf{X} and conditioning on outcome aa keeps the diagonal block CaaC_{aa} intact, everything inside it and not just its trace, and discards the other three. That is precisely the information ρX\rho_{\mathsf{X}} threw away. The off-diagonal blocks C01C_{01} and C10C_{10} and the rival outcome’s block C11C_{11} all contained correlation between X\mathsf{X} and Z\mathsf{Z}, but only CaaC_{aa} describes Z\mathsf{Z} in the branch where the outcome actually happened.

So the reduced state ρX\rho_{\mathsf{X}} and the conditional state ρZ∣a\rho_{\mathsf{Z}\mid a} are complementary reductions of the same joint state: one averages over Z\mathsf{Z} to predict the outcome, the other conditions on the outcome to describe what remains. Neither can be obtained from the other, because each retains exactly the information the other discards. Together they answer the two questions, but only ρ\rho contains both answers at once.

Naimark’s theorem

So far, a measurement has been specified by its matrices, with the outcome probabilities obtained from a trace. That tells us what the measurement does, but not how to implement it as a circuit.

Naimark’s theorem gives such an implementation. Any {P0,…,Pm−1}\{P_0,\ldots,P_{m-1}\} on X\mathsf{X} can be realized by adding an initialized workspace, applying one unitary operation to the combined system, and then measuring the workspace in the standard basis.

XYY

This is the measurement counterpart of a . The same kind of workspace is introduced in the state ∣0⟩\lvert0\rangle, and the same single unitary acts on the combined system. The difference is in the final step: for a channel the workspace is discarded, and for a measurement it is measured.

The proof is constructive. We choose the part of UU that the initial state can actually reach, show that this part is compatible with a unitary, and then check that the resulting circuit has the required outcome probabilities.

The square root of a positive semidefinite matrix

The construction relies on one fact about positive semidefinite matrices. Every positive semidefinite matrix PP has a unique positive semidefinite square root, written P\sqrt{P} and characterized by (P)2=P\bigl(\sqrt{P}\bigr)^{2}=P.

Take a spectral decomposition of PP:

P=∑k=0n−1λk∣ψk⟩⟨ψk∣P=\sum_{k=0}^{n-1}\lambda_k\lvert\psi_k\rangle\langle\psi_k\rvert

Then P\sqrt{P} is obtained by replacing each eigenvalue with its square root:

P=∑k=0n−1λk∣ψk⟩⟨ψk∣\sqrt{P}=\sum_{k=0}^{n-1}\sqrt{\lambda_k}\lvert\psi_k\rangle\langle\psi_k\rvert

The eigenvalues λk\lambda_k are nonnegative, so their square roots are real. The eigenvectors stay the same, and only the eigenvalues change.

Because every measurement matrix PaP_a is positive semidefinite, each one has such a square root Pa\sqrt{P_a}.

Choosing the unitary

Naimark’s construction uses a unitary on the system X\mathsf{X} and an auxiliary workspace Y\mathsf{Y} to implement the measurement. Arrange the combined system in the order (Y,X)(\mathsf{Y},\mathsf{X}), so the matrix is divided into blocks according to Y\mathsf{Y}, with each block containing the action on X\mathsf{X}.

We want UU to transform the initial state ∣0⟩Y⊗ρX\lvert0\rangle_{\mathsf{Y}}\otimes\rho_{\mathsf{X}} so that the different measurement outcomes are recorded in Y\mathsf{Y}. Since the input workspace state is ∣0⟩\lvert0\rangle, only block column 0 of UU determines this transformation.

If the workspace started in a different state, other block columns could contribute. For the present input, however, the remaining columns are not reached. They only need to complete UU to a valid unitary, so their specific values are irrelevant:

V=(P0P1⋮Pm−1)V=\begin{pmatrix}\sqrt{P_0}\\\sqrt{P_1}\\\vdots\\\sqrt{P_{m-1}}\end{pmatrix}
U=U=
P0\sqrt{P_0}
P1\sqrt{P_1}
⋮\vdots
Pm−1\sqrt{P_{m-1}}
??

Completing the unitary

A matrix is unitary exactly when its columns form an . So to complete UU, the specified block column must first have orthonormal columns.

The measurement condition gives exactly this. The conjugate-transpose product of VV collapses to the identity on X\mathsf{X}, so VV is an isometry — its columns are orthonormal in the larger (Y,X)(\mathsf{Y},\mathsf{X}) space:

V†V=∑a=0m−1(Pa)†Pa=∑a=0m−1Pa=IXV^{\dagger}V=\sum_{a=0}^{m-1}\bigl(\sqrt{P_a}\bigr)^{\dagger}\sqrt{P_a} =\sum_{a=0}^{m-1}P_a=I_{\mathsf{X}}

Any can be extended to an orthonormal basis. We can therefore choose additional columns to complete the columns of VV to a basis and place them in the remaining part of UU. The resulting matrix is unitary.

The unspecified blocks therefore represent exactly this freedom: different valid choices of the remaining columns give different unitary completions, while the required action on the initial Y=0\mathsf{Y}=0 input stays the same.

Why the circuit works

The unitary was chosen to have the required action on the initial state ∣0⟩Y⊗ρX\lvert0\rangle_{\mathsf{Y}}\otimes\rho_{\mathsf{X}}. We can now run the circuit and verify that measuring Y\mathsf{Y} produces the correct probabilities.

The circuit starts with the workspace in ∣0⟩Y\lvert0\rangle_{\mathsf{Y}}, applies UU, and then measures Y\mathsf{Y} in the standard basis. Order the joint system as (Y,X)(\mathsf{Y},\mathsf{X}), so each block is indexed by a workspace state and contains an operator on X\mathsf{X}. Before applying UU, the joint state is σin=∣0⟩⟨0∣Y⊗ρX\sigma_{\mathrm{in}}=\lvert0\rangle\langle0\rvert_{\mathsf{Y}}\otimes\rho_{\mathsf{X}}.

Its block with row aa and column bb is (σin)ab=⟨a∣0⟩⟨0∣b⟩ ρ\bigl(\sigma_{\mathrm{in}}\bigr)_{ab}=\langle a\rvert0\rangle\langle0\rvert b\rangle\,\rho. Because the workspace basis states are orthonormal, ⟨a∣0⟩\langle a\rvert0\rangle is zero unless a=0a=0, and ⟨0∣b⟩\langle0\rvert b\rangle is zero unless b=0b=0. Therefore every block is zero except the (0,0)(0,0) block:

σin=(ρ00ρ01⋯ρ0,m−1ρ10ρ11⋯ρ1,m−1⋮⋮⋱⋮ρm−1,0ρm−1,1⋯ρm−1,m−1)=\sigma_{\mathrm{in}}= \begin{pmatrix} \rho_{00}&\rho_{01}&\cdots&\rho_{0,m-1}\\ \rho_{10}&\rho_{11}&\cdots&\rho_{1,m-1}\\ \vdots&\vdots&\ddots&\vdots\\ \rho_{m-1,0}&\rho_{m-1,1}&\cdots&\rho_{m-1,m-1} \end{pmatrix}=ρ\rho00⋯\cdots000000⋯\cdots00⋮\vdots⋮\vdots⋱\ddots⋮\vdots0000⋯\cdots00

What has to be shown is that, after applying UU, measuring Y\mathsf{Y} gives outcome aa with probability Tr(Paρ)\mathrm{Tr}(P_a\rho). This must hold for every input state ρ\rho and every outcome aa, so that the circuit implements exactly the original measurement {P0,…,Pm−1}\{P_0,\ldots,P_{m-1}\}.

Sandwiching the input state between UU and U†U^{\dagger} leaves only the first block column of UU and the first block row of U†U^{\dagger}. Every other block meets a zero block of the input and therefore makes no contribution. The unspecified part of UU is never reached by this particular input:

σout=UσinU†=\sigma_{\mathrm{out}}=U\sigma_{\mathrm{in}}U^{\dagger}=VVP0\sqrt{P_0}P1\sqrt{P_1}⋮\vdotsPm−1\sqrt{P_{m-1}}??ρ\rho00⋯\cdots000000⋯\cdots00⋮\vdots⋮\vdots⋱\ddots⋮\vdots0000⋯\cdots00V†V^{\dagger}P0\sqrt{P_0}P1\sqrt{P_1}⋯\cdotsPm−1\sqrt{P_{m-1}}?†?^{\dagger}
σout=VρV†  +  V 0 ?†  +  ? 0 V†  +  ? 0 ?†  +  ⋯=VρV†\phantom{\sigma_{\mathrm{out}}} =V\rho V^{\dagger}\textcolor{#94a3b8}{\;+\;V\,0\,?^{\dagger}}\textcolor{#94a3b8}{\;+\;?\,0\,V^{\dagger}}\textcolor{#94a3b8}{\;+\;?\,0\,?^{\dagger}}\textcolor{#94a3b8}{\;+\;\cdots}=V\rho V^{\dagger}
σout=(P0 ρ P0⋯P0 ρ Pm−1⋮⋱⋮Pm−1 ρ P0⋯Pm−1 ρ Pm−1)\sigma_{\mathrm{out}}=\begin{pmatrix} \sqrt{P_0}\,\rho\,\sqrt{P_0}&\cdots&\sqrt{P_0}\,\rho\,\sqrt{P_{m-1}}\\ \vdots&\ddots&\vdots\\ \sqrt{P_{m-1}}\,\rho\,\sqrt{P_0}&\cdots&\sqrt{P_{m-1}}\,\rho\,\sqrt{P_{m-1}} \end{pmatrix}

The same matrix written as a sum over blocks, with ∣a⟩⟨b∣\lvert a\rangle\langle b\rvert naming the block position in Y\mathsf{Y} and the X\mathsf{X} factor giving its contents:

σout=∑a,b=0m−1∣a⟩⟨b∣⊗Pa ρ Pb\sigma_{\mathrm{out}}=\sum_{a,b=0}^{m-1}\lvert a\rangle\langle b\rvert\otimes\sqrt{P_a}\,\rho\,\sqrt{P_b}

The final measurement acts on Y\mathsf{Y} alone, so first reduce the state to Y\mathsf{Y} by tracing out X\mathsf{X}. Taking the trace of each block gives

σY=∑a,b=0m−1Tr(Pa ρ Pb)∣a⟩⟨b∣\sigma_{\mathsf{Y}}=\sum_{a,b=0}^{m-1} \mathrm{Tr}\bigl(\sqrt{P_a}\,\rho\,\sqrt{P_b}\bigr)\lvert a\rangle\langle b\rvert
Y012012
P0 ρ P0\sqrt{P_{0}}\,\textcolor{#c2410c}{\rho}\,\sqrt{P_{0}}
P0 ρ P1\sqrt{P_{0}}\,\textcolor{#c2410c}{\rho}\,\sqrt{P_{1}}
P0 ρ P2\sqrt{P_{0}}\,\textcolor{#c2410c}{\rho}\,\sqrt{P_{2}}
P1 ρ P0\sqrt{P_{1}}\,\textcolor{#c2410c}{\rho}\,\sqrt{P_{0}}
P1 ρ P1\sqrt{P_{1}}\,\textcolor{#c2410c}{\rho}\,\sqrt{P_{1}}
P1 ρ P2\sqrt{P_{1}}\,\textcolor{#c2410c}{\rho}\,\sqrt{P_{2}}
P2 ρ P0\sqrt{P_{2}}\,\textcolor{#c2410c}{\rho}\,\sqrt{P_{0}}
P2 ρ P1\sqrt{P_{2}}\,\textcolor{#c2410c}{\rho}\,\sqrt{P_{1}}
P2 ρ P2\sqrt{P_{2}}\,\textcolor{#c2410c}{\rho}\,\sqrt{P_{2}}
TrX\mathrm{Tr}_{\mathsf{X}}
Y012012
Tr(P0ρ)\textcolor{#059669}{\mathrm{Tr}(P_{0}\rho)}
Tr(P0ρP1)\textcolor{#94a3b8}{\mathrm{Tr}\bigl(\sqrt{P_{0}}\rho\sqrt{P_{1}}\bigr)}
Tr(P0ρP2)\textcolor{#94a3b8}{\mathrm{Tr}\bigl(\sqrt{P_{0}}\rho\sqrt{P_{2}}\bigr)}
Tr(P1ρP0)\textcolor{#94a3b8}{\mathrm{Tr}\bigl(\sqrt{P_{1}}\rho\sqrt{P_{0}}\bigr)}
Tr(P1ρ)\textcolor{#059669}{\mathrm{Tr}(P_{1}\rho)}
Tr(P1ρP2)\textcolor{#94a3b8}{\mathrm{Tr}\bigl(\sqrt{P_{1}}\rho\sqrt{P_{2}}\bigr)}
Tr(P2ρP0)\textcolor{#94a3b8}{\mathrm{Tr}\bigl(\sqrt{P_{2}}\rho\sqrt{P_{0}}\bigr)}
Tr(P2ρP1)\textcolor{#94a3b8}{\mathrm{Tr}\bigl(\sqrt{P_{2}}\rho\sqrt{P_{1}}\bigr)}
Tr(P2ρ)\textcolor{#059669}{\mathrm{Tr}(P_{2}\rho)}

A standard basis measurement of Y\mathsf{Y} reads outcome aa from the aath diagonal entry. Using cyclicity of the trace,

Pr(outcome=a)=⟨a∣σY∣a⟩=Tr(Pa ρ Pa)=Tr(Paρ)\mathrm{Pr}(\text{outcome}=a)=\langle a\rvert\sigma_{\mathsf{Y}}\lvert a\rangle =\mathrm{Tr}\bigl(\sqrt{P_a}\,\rho\,\sqrt{P_a}\bigr)=\mathrm{Tr}(P_a\rho)

This is exactly the probability assigned to outcome aa by the original measurement {P0,…,Pm−1}\{P_0,\ldots,P_{m-1}\}.

So the circuit reproduces the measurement exactly. Every general measurement can therefore be implemented by adding an initialized workspace with one basis state per outcome, applying a single unitary to the workspace and the measured system, and performing a standard basis measurement on the workspace.

Non-destructive measurements

A destructive measurement describes the classical outcome probabilities, without retaining a quantum state of the measured system. A non-destructive measurement has both a classical outcome and a post-measurement quantum state of the system that was measured.

There are two useful ways to formulate a non-destructive measurement. The first reads the post-measurement state directly from the , where the measured system remains on an outgoing wire. The second gives the general mathematical form directly.

Reading the state from Naimark’s construction

Take a general (destructive) measurement {P0,…,Pm−1}\{P_0,\ldots,P_{m-1}\} of a system X\mathsf{X}, and the . The final measurement acts on the workspace Y\mathsf{Y} alone, and nothing discards X\mathsf{X}. Ignoring that outgoing wire is what made the measurement destructive. Keeping it gives a non-destructive measurement with exactly the same outcome probabilities.

XYY

The joint state just before the is:

σout=∑a,b=0m−1∣a⟩⟨b∣⊗Pa ρ Pb\sigma_{\mathrm{out}}=\sum_{a,b=0}^{m-1}\lvert a\rangle\langle b\rvert\otimes\sqrt{P_a}\,\rho\,\sqrt{P_b}

Measuring Y\mathsf{Y} in the standard basis and obtaining outcome aa leaves Y\mathsf{Y} in ∣a⟩\lvert a\rangle. This selects the (a,a)(a,a) block of σout\sigma_{\mathrm{out}}, so the state left on X\mathsf{X} is Pa ρ Pa\sqrt{P_a}\,\rho\,\sqrt{P_a}.

Dividing by the probability of outcome aa normalizes the state, so conditioned on observing aa, the state of X\mathsf{X} becomes Pa ρ PaTr(Paρ)\displaystyle\frac{\sqrt{P_a}\,\rho\,\sqrt{P_a}}{\mathrm{Tr}(P_a\rho)}.

From Kraus operators

Naimark’s construction uses Pa\sqrt{P_a} as the measurement operator for outcome aa, but this choice is not unique. The outcome probability depends on Pa\sqrt{P_a} only through (Pa)†Pa=Pa\bigl(\sqrt{P_a}\bigr)^{\dagger}\sqrt{P_a}=P_a, so different matrices can produce the same probabilities while leaving the system in different post-measurement states.

Let M0,…,Mm−1M_0,\ldots,M_{m-1} be square matrices satisfying ∑a=0m−1Ma†Ma=IX\sum_{a=0}^{m-1}M_a^{\dagger}M_a=I_{\mathsf{X}} — they define a non-destructive measurement.

For an input state ρ\rho, outcome aa occurs with probability Pr(outcome=a)=Tr(MaρMa†)=Tr(Ma†Maρ)\mathrm{Pr}(\text{outcome}=a)=\mathrm{Tr}\bigl(M_a\rho M_a^{\dagger}\bigr)=\mathrm{Tr}\bigl(M_a^{\dagger}M_a\rho\bigr).

Conditioned on that outcome, the measured system is left in MaρMa†Tr(MaρMa†)\displaystyle\frac{M_a\rho M_a^{\dagger}}{\mathrm{Tr}\bigl(M_a\rho M_a^{\dagger}\bigr)}.

If the outcome is ignored, the state changes as ρ↦∑aMa ρ Ma†\rho\mapsto\sum_a M_a\,\rho\,M_a^{\dagger}. This is a quantum channel, with the MaM_a as its . The same matrices therefore describe the non-destructive measurement when each term is associated with its corresponding outcome.

State discrimination and tomography

Quantum state discrimination

A measurement is not just something that produces an outcome. When the possible states are known in advance, the measurement can be chosen to make those outcomes as informative as possible. Quantum state discrimination asks how to choose that measurement when the goal is to identify which one of several known states was prepared.

Let ρ0,…,ρm−1\rho_0,\ldots,\rho_{m-1} be quantum states of a system X\mathsf{X}, and let (p0,…,pm−1)(p_0,\ldots,p_{m-1}) be the probabilities with which they are prepared. A label a∈{0,…,m−1}a\in\{0,\ldots,m-1\} is drawn according to (p0,…,pm−1)(p_0,\ldots,p_{m-1}), and the system is then prepared in the corresponding state ρa\rho_a. The label is hidden, while the list of possible states and their probabilities are known. The task is to measure X\mathsf{X} and guess which label was chosen.

A strategy is a {P0,…,Pm−1}\{P_0,\ldots,P_{m-1}\}, with one outcome for each possible label. If outcome aa occurs, the strategy guesses that the prepared state was ρa\rho_a. Its performance is measured by the probability of guessing correctly:

Pr(correct)=∑a=0m−1pa Tr(Paρa)\mathrm{Pr}(\text{correct})=\sum_{a=0}^{m-1}p_a\,\mathrm{Tr}(P_a\rho_a)

The measurement should therefore be chosen to make this quantity as large as possible. The probabilities pap_a matter because some states are more likely to occur than others: correctly identifying a common state contributes more to the overall success probability than correctly identifying a rare one. A measurement that maximizes the expression is called a minimum-error measurement.

This immediately gives two useful limits. If the possible states are mutually orthogonal, they can be , so the success probability can reach 11. At the other extreme, the measurement can simply be ignored: always guessing the most likely state already gives success probability max⁡apa\max_a p_a. State discrimination is interesting because the measurement can do better than this baseline by extracting information from X\mathsf{X}, while perfect discrimination is possible only when the states are sufficiently distinguishable.

The problem is therefore not to determine an unknown quantum state from scratch. The possible states are already known, and only the label identifying the prepared state is hidden. This is what distinguishes quantum state discrimination from , where the state itself is unknown and must be reconstructed from measurement data.

Discriminating pairs of states

With only two possible states, the optimal measurement for discriminating between them has a closed-form solution. A pair ρ0\rho_0 and ρ1\rho_1 is best discriminated by the Helstrom measurement—a two-outcome {Π0,Π1}\{\Pi_0,\Pi_1\} read off a single Hermitian operator built from the two states and the prior probabilities p0p_0 and p1p_1 with which they are prepared.

The Helstrom operator is the weighted difference of the two states:

Δ=p0ρ0−p1ρ1\Delta=p_0\rho_0-p_1\rho_1

Read it as a scoreboard. Every direction in the state space gets a score, and the sign of the score says which hypothesis that direction supports:

  • Δ>0\Delta>0—the weighted evidence for ρ0\rho_0 outweighs that for ρ1\rho_1. Assign this direction to outcome 00, i.e. answer “ρ0\rho_0 was prepared.”
  • Δ<0\Delta<0—ρ1\rho_1 wins. Assign it to outcome 11, answer ρ1\rho_1 instead.
  • Δ=0\Delta=0—the two weighted contributions cancel exactly. The direction carries no information, and either answer is right equally often.
θ = 72°
ρ1\rho_1
ρ0\rho_0
Δ<0→say ρ1\Delta<0\rightarrow\text{say }\rho_1λ−=−0.29\lambda_-=-0.29λ+=0.29\lambda_+=0.29Δ>0→say ρ0\Delta>0\rightarrow\text{say }\rho_0
0.50
72°
Pr(correct identification)=12(1+∥Δ∥1)=12(1+0.59)≈0.79\begin{aligned} \mathrm{Pr}(\text{correct identification}) &=\tfrac12\bigl(1+\lVert\Delta\rVert_1\bigr)\\[4pt] &=\tfrac12\bigl(1+0.59\bigr)\approx0.79 \end{aligned}

Note that Δ\Delta is not a density operator: its trace is p0−p1p_0-p_1, not 11, and its eigenvalues may be negative. Both are features—Δ\Delta encodes a comparison of two states, so its eigenvalues are scores, not probabilities.

Because Δ\Delta is Hermitian, the applies: it can be diagonalized. Thus there is an orthonormal basis {∣ψk⟩}\{\lvert\psi_k\rangle\} of eigenvectors of Δ\Delta, with real eigenvalues λk\lambda_k, such that Δ∣ψk⟩=λk∣ψk⟩\Delta\lvert\psi_k\rangle=\lambda_k\lvert\psi_k\rangle.

These directions ∣ψk⟩\lvert\psi_k\rangle are not chosen by hand—they are determined by Δ\Delta itself. They are the directions on which the action of Δ\Delta is simple: it does not change the direction and only multiplies it by the real number λk\lambda_k.

The same eigenvectors yield the spectral decomposition Δ=∑kλk∣ψk⟩⟨ψk∣\Delta=\sum_k\lambda_k\lvert\psi_k\rangle\langle\psi_k\rvert.

What matters here is the meaning of λk\lambda_k. Since Δ=p0ρ0−p1ρ1\Delta=p_0\rho_0-p_1\rho_1, along the direction ∣ψk⟩\lvert\psi_k\rangle the eigenvalue indicates which of the two weighted contributions is larger. A positive value means that p0ρ0p_0\rho_0 outweighs p1ρ1p_1\rho_1 in that direction, a negative value means the reverse, and zero means the two are exactly equal.

p0ρ0p_0\rho_0
p1ρ1p_1\rho_1
Δ\Delta
∣ψ+⟩⟨ψ+∣\lvert\psi_+\rangle\langle\psi_+\rvert
∣ψ−⟩⟨ψ−∣\lvert\psi_-\rangle\langle\psi_-\rvert
λ+\lambda_+
λ−\lambda_-
Δ\Delta
0.50
72°
Δ=λ+∣ψ+⟩⟨ψ+∣+λ−∣ψ−⟩⟨ψ−∣=0.29∣ψ+⟩⟨ψ+∣−0.29∣ψ−⟩⟨ψ−∣\begin{aligned} \Delta&=\lambda_+\lvert\psi_+\rangle\langle\psi_+\rvert +\lambda_-\lvert\psi_-\rangle\langle\psi_-\rvert\\[6pt] &=\textcolor{#0284c7}{0.29} \lvert\psi_+\rangle\langle\psi_+\rvert \textcolor{#7c3aed}{-0.29} \lvert\psi_-\rangle\langle\psi_-\rvert \end{aligned}

Grouping the indices by the sign of the corresponding eigenvalue gives two sets:

S0={k∈{0,…,n−1}:λk≥0}S1={k∈{0,…,n−1}:λk<0}\begin{aligned} S_0&=\{k\in\{0,\ldots,n-1\}:\lambda_k\geq0\}\\[4pt] S_1&=\{k\in\{0,\ldots,n-1\}:\lambda_k<0\} \end{aligned}

To turn each group into a measurement outcome, we construct a projector onto the corresponding subspace. For a single normalized direction ∣ψk⟩\lvert\psi_k\rangle, the operator ∣ψk⟩⟨ψk∣\lvert\psi_k\rangle\langle\psi_k\rvert projects onto that direction, keeping the component of a state along ∣ψk⟩\lvert\psi_k\rangle and discarding components along orthogonal directions. Since the eigenvectors within each group are mutually orthogonal, adding these individual projectors gives a projector onto the entire subspace spanned by that group. We therefore use the positive-eigenvalue subspace for outcome 00, identifying ρ0\rho_0, and the negative-eigenvalue subspace for outcome 11, identifying ρ1\rho_1:

Π0=∑k∈S0∣ψk⟩⟨ψk∣⟶ outcome 0, identify ρ0Π1=∑k∈S1∣ψk⟩⟨ψk∣⟶ outcome 1, identify ρ1\begin{aligned} \Pi_0&=\sum_{k\in S_0}\lvert\psi_k\rangle\langle\psi_k\rvert &&\quad\textcolor{#94a3b8}{\longrightarrow\ \text{outcome }0,\ \text{identify }\rho_0}\\[4pt] \Pi_1&=\sum_{k\in S_1}\lvert\psi_k\rangle\langle\psi_k\rvert &&\quad\textcolor{#94a3b8}{\longrightarrow\ \text{outcome }1,\ \text{identify }\rho_1} \end{aligned}

That {Π0,Π1}\{\Pi_0,\Pi_1\} is a valid measurement follows from the same decomposition. The eigenvectors form an orthonormal basis, so the two subspaces are orthogonal and together span the whole space, which in operator form reads

Π0Π1=0⟶ the outcomes never fire togetherΠ0+Π1=IX⟶ one of them always fires\begin{aligned} \Pi_0\Pi_1&=0 &&\quad\textcolor{#94a3b8}{\longrightarrow\ \text{the outcomes never fire together}}\\[4pt] \Pi_0+\Pi_1&=I_{\mathsf{X}} &&\quad\textcolor{#94a3b8}{\longrightarrow\ \text{one of them always fires}} \end{aligned}

The projectors determine the measurement outcomes:

Given that the system was prepared in ρ0\rho_0, the probabilities of the two possible outcomes are

Pr(0∣ρ0)=Tr(Π0ρ0),Pr(1∣ρ0)=Tr(Π1ρ0).\begin{aligned} \mathrm{Pr}(0\mid\rho_0)&=\mathrm{Tr}(\Pi_0\rho_0),\\[4pt] \mathrm{Pr}(1\mid\rho_0)&=\mathrm{Tr}(\Pi_1\rho_0). \end{aligned}

Given that the system was prepared in ρ1\rho_1, they are

Pr(0∣ρ1)=Tr(Π0ρ1),Pr(1∣ρ1)=Tr(Π1ρ1).\begin{aligned} \mathrm{Pr}(0\mid\rho_1)&=\mathrm{Tr}(\Pi_0\rho_1),\\[4pt] \mathrm{Pr}(1\mid\rho_1)&=\mathrm{Tr}(\Pi_1\rho_1). \end{aligned}

The first and last of these are the probabilities of correct identification: obtaining outcome 00 when ρ0\rho_0 was prepared, or outcome 11 when ρ1\rho_1 was prepared. Since ρ0\rho_0 is prepared with probability p0p_0 and ρ1\rho_1 with probability p1p_1, the overall success probability is the prior-weighted sum of these two correct-identification probabilities:

Pr(correct identification)=p0 Tr(Π0ρ0)+p1 Tr(Π1ρ1).\mathrm{Pr}(\text{correct identification})=p_0\,\mathrm{Tr}(\Pi_0\rho_0)+p_1\,\mathrm{Tr}(\Pi_1\rho_1).

To express this in terms of the Helstrom operator Δ\Delta, use Π1=IX−Π0\Pi_1=I_{\mathsf{X}}-\Pi_0. Because ρ1\rho_1 is a density operator, Tr(ρ1)=1\mathrm{Tr}(\rho_1)=1, so

Pr(correct identification)=p0 Tr(Π0ρ0)+p1 Tr((IX−Π0)ρ1)=p0 Tr(Π0ρ0)+p1(Tr(IXρ1)−Tr(Π0ρ1))=p0 Tr(Π0ρ0)+p1−p1 Tr(Π0ρ1)=p1+p0 Tr(Π0ρ0)−p1 Tr(Π0ρ1)=p1+Tr(Π0 p0ρ0)−Tr(Π0 p1ρ1)=p1+Tr(Π0(p0ρ0−p1ρ1)).\begin{aligned} \mathrm{Pr}(\text{correct identification}) &=p_0\,\mathrm{Tr}(\Pi_0\rho_0) +p_1\,\mathrm{Tr}\bigl((I_{\mathsf{X}}-\Pi_0)\rho_1\bigr)\\[4pt] &=p_0\,\mathrm{Tr}(\Pi_0\rho_0) +p_1\bigl(\mathrm{Tr}(I_{\mathsf{X}}\rho_1)-\mathrm{Tr}(\Pi_0\rho_1)\bigr)\\[4pt] &=p_0\,\mathrm{Tr}(\Pi_0\rho_0)+p_1-p_1\,\mathrm{Tr}(\Pi_0\rho_1)\\[4pt] &=p_1+p_0\,\mathrm{Tr}(\Pi_0\rho_0)-p_1\,\mathrm{Tr}(\Pi_0\rho_1)\\[4pt] &=p_1+\mathrm{Tr}(\Pi_0\,p_0\rho_0)-\mathrm{Tr}(\Pi_0\,p_1\rho_1)\\[4pt] &=p_1+\mathrm{Tr}\bigl(\Pi_0(p_0\rho_0-p_1\rho_1)\bigr). \end{aligned}

Since Δ=p0ρ0−p1ρ1\Delta=p_0\rho_0-p_1\rho_1, this becomes Pr(correct identification)=p1+Tr(Π0Δ)\mathrm{Pr}(\text{correct identification})=p_1+\mathrm{Tr}(\Pi_0\Delta).

Now we can see why the positive-eigenvalue subspace was chosen for Π0\Pi_0. That expression shows that, with p1p_1 fixed, maximizing the probability of correct identification means maximizing Tr(Π0Δ)\mathrm{Tr}(\Pi_0\Delta). In the eigenbasis of Δ\Delta, each direction ∣ψk⟩\lvert\psi_k\rangle contributes its eigenvalue λk\lambda_k. Therefore, including a direction with λk>0\lambda_k>0 increases the trace, while including a direction with λk<0\lambda_k<0 decreases it. The optimal projector Π0\Pi_0 must therefore include all positive-eigenvalue directions and exclude all negative-eigenvalue directions. This is exactly the projector we constructed from S0S_0:

Π0=∑k∈S0∣ψk⟩⟨ψk∣.\Pi_0=\sum_{k\in S_0}\lvert\psi_k\rangle\langle\psi_k\rvert.

We can now evaluate the value of the trace for this optimal choice. Since Δ=∑jλj∣ψj⟩⟨ψj∣\Delta=\sum_j\lambda_j\lvert\psi_j\rangle\langle\psi_j\rvert,

Π0Δ=(∑k∈S0∣ψk⟩⟨ψk∣)(∑jλj∣ψj⟩⟨ψj∣)the projector times the decomposition=∑k∈S0∑jλj∣ψk⟩⟨ψk∣ψj⟩⟨ψj∣multiplied out, every k against every j=∑k∈S0∑jλj δkj∣ψk⟩⟨ψj∣orthonormality: ⟨ψk∣ψj⟩=δkj=∑k∈S0λk∣ψk⟩⟨ψk∣δkj=0 unless j=k, so only that term survives\begin{aligned} \Pi_0\Delta &=\Bigl(\sum_{k\in S_0}\lvert\psi_k\rangle\langle\psi_k\rvert\Bigr) \Bigl(\sum_j\lambda_j\lvert\psi_j\rangle\langle\psi_j\rvert\Bigr) &&\quad\textcolor{#94a3b8}{\text{the projector times the decomposition}}\\[4pt] &=\sum_{k\in S_0}\sum_j\lambda_j \lvert\psi_k\rangle\langle\psi_k\vert\psi_j\rangle\langle\psi_j\rvert &&\quad\textcolor{#94a3b8}{\text{multiplied out, every }k\text{ against every }j}\\[4pt] &=\sum_{k\in S_0}\sum_j\lambda_j\,\delta_{kj} \lvert\psi_k\rangle\langle\psi_j\rvert &&\quad\textcolor{#94a3b8}{\text{orthonormality: }\langle\psi_k\vert\psi_j\rangle=\delta_{kj}}\\[4pt] &=\sum_{k\in S_0}\lambda_k\lvert\psi_k\rangle\langle\psi_k\rvert &&\quad\textcolor{#94a3b8}{\delta_{kj}=0\ \text{unless }j=k,\ \text{so only that term survives}} \end{aligned}

and taking the trace,

Tr(Π0Δ)=Tr(∑k∈S0λk∣ψk⟩⟨ψk∣)=∑k∈S0λk Tr(∣ψk⟩⟨ψk∣)the trace is linear=∑k∈S0λkTr(∣ψk⟩⟨ψk∣)=⟨ψk∣ψk⟩=1\begin{aligned} \mathrm{Tr}(\Pi_0\Delta) &=\mathrm{Tr}\Bigl(\sum_{k\in S_0}\lambda_k \lvert\psi_k\rangle\langle\psi_k\rvert\Bigr)\\[4pt] &=\sum_{k\in S_0}\lambda_k\, \mathrm{Tr}\bigl(\lvert\psi_k\rangle\langle\psi_k\rvert\bigr) &&\quad\textcolor{#94a3b8}{\text{the trace is linear}}\\[4pt] &=\sum_{k\in S_0}\lambda_k &&\quad\textcolor{#94a3b8}{\mathrm{Tr}\bigl(\lvert\psi_k\rangle\langle\psi_k\rvert\bigr)=\langle\psi_k\vert\psi_k\rangle=1} \end{aligned}

The success probability is therefore obtained by adding the fixed term p1p_1 to the contributions from all nonnegative eigenvalues selected by Π0\Pi_0:

Pr(correct identification)=p1+∑k∈S0λk.\mathrm{Pr}(\text{correct identification})=p_1+\sum_{k\in S_0}\lambda_k.

To put this into a symmetric form, start from the trace of Δ\Delta. Both ρ0\rho_0 and ρ1\rho_1 are density operators, so each has unit trace:

Tr(Δ)=Tr(p0ρ0−p1ρ1)the definition of Δ=p0 Tr(ρ0)−p1 Tr(ρ1)the trace is linear=p0−p1Tr(ρ0)=Tr(ρ1)=1\begin{aligned} \mathrm{Tr}(\Delta) &=\mathrm{Tr}(p_0\rho_0-p_1\rho_1) &&\quad\textcolor{#94a3b8}{\text{the definition of }\Delta}\\[4pt] &=p_0\,\mathrm{Tr}(\rho_0)-p_1\,\mathrm{Tr}(\rho_1) &&\quad\textcolor{#94a3b8}{\text{the trace is linear}}\\[4pt] &=p_0-p_1 &&\quad\textcolor{#94a3b8}{\mathrm{Tr}(\rho_0)=\mathrm{Tr}(\rho_1)=1} \end{aligned}

The trace is also the sum of the eigenvalues, and S0S_0 and S1S_1 between them account for every index, so that sum splits in two:

∑k∈S0λk+∑k∈S1λk=∑kλk=Tr(Δ)=p0−p1every index lies in one set or the other∑k∈S1λk=p0−p1−∑k∈S0λkrearranged for the negative part\begin{aligned} \sum_{k\in S_0}\lambda_k+\sum_{k\in S_1}\lambda_k &=\sum_k\lambda_k=\mathrm{Tr}(\Delta)=p_0-p_1 &&\quad\textcolor{#94a3b8}{\text{every index lies in one set or the other}}\\[4pt] \sum_{k\in S_1}\lambda_k &=p_0-p_1-\sum_{k\in S_0}\lambda_k &&\quad\textcolor{#94a3b8}{\text{rearranged for the negative part}} \end{aligned}

Every eigenvalue in S1S_1 is negative, so ∣λk∣=−λk\lvert\lambda_k\rvert=-\lambda_k there, and the total magnitude of the eigenvalues can be written using only the selected sum:

∑k∣λk∣=∑k∈S0∣λk∣+∑k∈S1∣λk∣=∑k∈S0λk−∑k∈S1λk∣λk∣=λk on S0, −λk on S1=∑k∈S0λk−(p0−p1−∑k∈S0λk)substituting the line above=2∑k∈S0λk−p0+p1collecting the two copies\begin{aligned} \sum_k\lvert\lambda_k\rvert &=\sum_{k\in S_0}\lvert\lambda_k\rvert+\sum_{k\in S_1}\lvert\lambda_k\rvert\\[4pt] &=\sum_{k\in S_0}\lambda_k-\sum_{k\in S_1}\lambda_k &&\quad\textcolor{#94a3b8}{\lvert\lambda_k\rvert=\lambda_k\text{ on }S_0,\ -\lambda_k\text{ on }S_1}\\[4pt] &=\sum_{k\in S_0}\lambda_k-\Bigl(p_0-p_1-\sum_{k\in S_0}\lambda_k\Bigr) &&\quad\textcolor{#94a3b8}{\text{substituting the line above}}\\[4pt] &=2\sum_{k\in S_0}\lambda_k-p_0+p_1 &&\quad\textcolor{#94a3b8}{\text{collecting the two copies}} \end{aligned}

Solving that for the selected sum gives ∑k∈S0λk=12(∑k∣λk∣+p0−p1)\sum_{k\in S_0}\lambda_k=\tfrac12\bigl(\sum_k\lvert\lambda_k\rvert+p_0-p_1\bigr), and substituting it into the success probability above:

Pr(correct identification)=p1+12(∑k∣λk∣+p0−p1)substituted into p1+∑k∈S0λk=p0+p12+12∑k∣λk∣collecting the priors=12+12∑k∣λk∣p0+p1=1\begin{aligned} \mathrm{Pr}(\text{correct identification}) &=p_1+\tfrac12\Bigl(\sum_k\lvert\lambda_k\rvert+p_0-p_1\Bigr) &&\quad\textcolor{#94a3b8}{\text{substituted into }p_1+\textstyle\sum_{k\in S_0}\lambda_k}\\[4pt] &=\frac{p_0+p_1}{2}+\frac12\sum_k\lvert\lambda_k\rvert &&\quad\textcolor{#94a3b8}{\text{collecting the priors}}\\[4pt] &=\frac12+\frac12\sum_k\lvert\lambda_k\rvert &&\quad\textcolor{#94a3b8}{p_0+p_1=1} \end{aligned}

Since ∑k∣λk∣\sum_k\lvert\lambda_k\rvert is the trace norm ∥Δ∥1\lVert\Delta\rVert_1, this becomes the Helstrom bound

  Pr(correct identification)=12+12∥Δ∥1.  \boxed{\;\mathrm{Pr}(\text{correct identification})=\frac12+\frac12\lVert\Delta\rVert_1.\;}

The two extremes can be read straight off this formula. If the states are identical and the priors equal, then Δ=0\Delta=0, the norm vanishes, and the measurement does no better than a coin toss at 1/21/2. If the states are orthogonal, the norm equals 11, the success probability reaches 11, and is possible. Everything else lies in between.

The Helstrom bound applies to any pair of quantum states, whether pure or mixed, and to arbitrary prior probabilities. A particularly simple special case is two pure states with equal priors, ρ0=∣ψ0⟩⟨ψ0∣\rho_0=\lvert\psi_0\rangle\langle\psi_0\rvert, ρ1=∣ψ1⟩⟨ψ1∣\rho_1=\lvert\psi_1\rangle\langle\psi_1\rvert, and p0=p1=1/2p_0=p_1=1/2. For this case, the trace norm can be expressed directly through the overlap of the two states, giving the compact formula

Pr(correct identification)=12+121−∣⟨ψ0∣ψ1⟩∣2.\mathrm{Pr}(\text{correct identification})=\frac12+\frac12\sqrt{1-\lvert\langle\psi_0\vert\psi_1\rangle\rvert^{2}}.

Here ∣⟨ψ0∣ψ1⟩∣\lvert\langle\psi_0\vert\psi_1\rangle\rvert measures how similar the two states are: the larger the overlap, the harder they are to distinguish, and the lower the achievable success probability.

States
Π0\Pi_0
Π1\Pi_1
ρ0\rho_0
ρ1\rho_1
0.50 / 0.50
Guess50%
Helstrom measurement79.4%
Partial discriminationEqual priors
∣⟨ψ0∣ψ1⟩∣\lvert\langle\psi_0\vert\psi_1\rangle\rvert0.81
λ+\lambda_+0.29
λ−\lambda_--0.29
Pr(correct identification)=1+∥p0ρ0−p1ρ1∥12=1+0.592≈0.79\begin{aligned} \mathrm{Pr}(\text{correct identification}) &=\frac{1+\lVert p_0\rho_0-p_1\rho_1\rVert_1}{2}\\[6pt] &=\frac{1+\textcolor{#0f766e}{0.59}}{2} \approx\textcolor{#0f766e}{0.79} \end{aligned}

Optimality of the Helstrom measurement

The measurement constructed from the eigenvalues of Δ\Delta achieves the Helstrom bound, but we still need to show that this bound cannot be exceeded by a more general measurement. The Helstrom–Holevo theorem establishes exactly this: the optimal success probability is unchanged even when arbitrary measurements, not only projective ones, are allowed.

To see this, consider any two-outcome {P0,P1}\{P_0,P_1\}. This is sufficient because any measurement with more outcomes can group them according to the two possible guesses, ρ0\rho_0 or ρ1\rho_1.

The derivation of the success probability did not rely on P0P_0 being a projector. It used only the completeness relation P0+P1=IXP_0+P_1=I_{\mathsf{X}} and the normalization Tr(ρ1)=1\mathrm{Tr}(\rho_1)=1. Therefore, for any two-outcome measurement {P0,P1}\{P_0,P_1\}, we may substitute P1=IX−P0P_1=I_{\mathsf{X}}-P_0 in exactly the same way to obtain Pr(correct identification)=p1+Tr(P0Δ)\mathrm{Pr}(\text{correct identification})=p_1+\mathrm{Tr}(P_0\Delta).

To compare an arbitrary P0P_0 with the projector Π0\Pi_0, look at what P0P_0 does along each eigenvector ∣ψk⟩\lvert\psi_k\rangle of Δ\Delta. Define ck=⟨ψk∣P0∣ψk⟩c_k=\langle\psi_k\rvert P_0\lvert\psi_k\rangle.

Because 0≤P0≤IX0\leq P_0\leq I_{\mathsf{X}}, each ckc_k lies between 00 and 11. It can therefore be viewed as how much of the kk-th eigendirection contributes to outcome 00. Using the eigenbasis of Δ\Delta, the trace becomes Tr(P0Δ)=∑kλkck\mathrm{Tr}(P_0\Delta)=\sum_k\lambda_k c_k.

For every nonnegative eigenvalue, the largest possible contribution is obtained by taking ck=1c_k=1, and for every negative eigenvalue, the largest contribution is obtained by taking ck=0c_k=0. Hence Tr(P0Δ)≤∑k∈S0λk\mathrm{Tr}(P_0\Delta)\leq\sum_{k\in S_0}\lambda_k.

The Helstrom projector Π0\Pi_0 has exactly this choice: ck=1c_k=1 for every k∈S0k\in S_0 and ck=0c_k=0 for every k∈S1k\in S_1. It therefore attains the maximum possible value of Tr(P0Δ)\mathrm{Tr}(P_0\Delta). Hence every two-outcome measurement satisfies, and since Π0\Pi_0 attains this bound, the Helstrom measurement is optimal:

Pr(correct identification)≤p1+∑k∈S0λk=12+12∥Δ∥1.\mathrm{Pr}(\text{correct identification})\leq p_1+\sum_{k\in S_0}\lambda_k=\frac12+\frac12\lVert\Delta\rVert_1.

Discriminating three or more states

Two-state discrimination works because there is only ever one question to answer. Every relevant direction either favors ρ0\rho_0 or favors ρ1\rho_1, and the sign of a single operator settles the choice. The is simply that sign, read off from the states.

With m≥3m\geq3 states, there is no such question. A direction that favors ρ0\rho_0 over ρ1\rho_1 may still be more useful for distinguishing ρ2\rho_2, so all outcomes compete simultaneously and no single operator has a sign by which to sort them. There is no known closed-form formula for the optimal measurement in general.

The can still be solved as an optimization. The success probability is linear in the measurement operators, while the operators themselves are constrained to be positive semidefinite and to sum to the identity. An optimization problem with exactly this structure is called a semidefinite program (SDP).

For a given set of states and priors, a numerical solver can therefore return the optimal measurement operators—usually as numerical matrices—together with the corresponding maximum success probability. What it generally cannot provide is a simple closed-form expression for those operators in terms of the states.

Verifying a candidate measurement

A measurement that was guessed, or constructed from the states using some natural recipe, can be tested directly for optimality. Given the proposed measurement {Pa}\{P_a\}, form Y=∑apaρaPaY=\sum_a p_a\rho_aP_a.

The Holevo–Yuen–Kennedy–Lax (HYKL) conditions state that this measurement is optimal exactly when both of the following hold:

  • Y=Y†Y=Y^{\dagger} (Hermiticity)
  • Y−pbρb≥0Y-p_b\rho_b\geq0 for every bb

When these conditions hold, the measurement’s success probability is Pr(correct identification)=Tr(Y)\mathrm{Pr}(\text{correct identification})=\mathrm{Tr}(Y).

For an arbitrary collection of states, verifying the conditions may still require a nontrivial calculation. Symmetry can simplify the problem: when several states are arranged symmetrically, the measurement suggested by that symmetry can sometimes be verified directly, and some such families can be solved exactly.

States evenly spaced around a circle

For example, consider mm equally likely states arranged symmetrically around a circle of the Bloch sphere. The symmetry makes it possible to construct and verify an optimal measurement explicitly. In this case, the optimal measurement succeeds with probability 2/m2/m, compared with 1/m1/m for guessing.

0
1
2
3
Guess25%
Best measurement50%
4
guess1m=14≈0.25best measurement2m=24≈0.50\begin{aligned} \text{guess}\quad&\tfrac1m=\tfrac1{4}\approx0.25\\[6pt] \text{best measurement}\quad&\tfrac2m=\tfrac2{4}\approx0.50 \end{aligned}

The tetrahedral states

The are the same idea, with the symmetry spread over the whole Bloch sphere rather than around a single circle. Four equally likely states point toward the vertices of a tetrahedron, and the corresponding symmetric four-outcome measurement uses Pa=∣ϕa⟩⟨ϕa∣/2P_a=\lvert\phi_a\rangle\langle\phi_a\rvert/2. The tetrahedral symmetry makes the HYKL conditions easy to verify, showing directly that this natural measurement is optimal.

0123|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
Guess25%
Best measurement50%

Asking which of a known list of states was prepared can look like an invented puzzle: if someone prepared the state, why not simply ask them? But theoretically, a quantum experiment can produce a state without revealing which preparation was used. A measurement then has to infer the state from the quantum system alone.

The discrimination problem sounds simple, but the mathematics turns out to be surprisingly heavy, and this treatment has not even touched the physical problem of how to construct the measurements. The payoff, however, is a fundamental limit: for non-orthogonal states, no measurement can identify the state with absolute certainty.

Quantum state tomography

Quantum state tomography is the task of reconstructing an unknown quantum state from measurement data.

Let ρ\rho be an unknown quantum state of a system. Identical systems X1,…,XN\mathsf{X}_1,\ldots,\mathsf{X}_N are each independently prepared in the state ρ\rho. The goal is to approximate ρ\rho by measuring X1,…,XN\mathsf{X}_1,\ldots,\mathsf{X}_N.

Compare this with . There, the candidates were handed over in advance and the only thing hidden was a label, so one system and one measurement sufficed to make a decision. Here, nothing is given in advance: the answer is a matrix rather than an index, and accuracy is paid for with the number of copies NN.

One copy is worth almost nothing on its own. A measurement of a single system returns a single outcome, and an outcome is merely a sample from a probability distribution rather than the distribution itself. Only by repeating a measurement across many identically prepared systems do the probabilities that describe ρ\rho begin to show through. This is why tomography is inherently a statistical procedure, while discrimination is not.

Quantum state tomography comes in several variants:

  • Local vs. global measurements. Measurements can be local, with each of X1,…,XN\mathsf{X}_1,\ldots,\mathsf{X}_N measured separately, or global, with a single joint measurement performed on all copies at once. Global measurements can extract more information from the same number of copies, but are much harder to implement.
  • Reconstruction strategies. Different strategies can be used to infer ρ\rho from the measurement data. The simplest inverts the relationship between the state and the outcome probabilities, treating the observed frequencies as exact. But this naive inversion can produce a matrix that is not a valid quantum state, motivating more sophisticated approaches.

Qubit tomography

Suppose ρ\rho is an unknown qubit state, and X1,…,XN\mathsf{X}_1,\ldots,\mathsf{X}_N are qubits independently prepared in the state ρ\rho. The goal is to determine ρ\rho by performing measurements on these qubits.

There are several possible measurement strategies. For example, one could measure the Pauli observables, or use a tetrahedral measurement, whose four outcomes correspond to the vertices of a tetrahedron on the Bloch sphere.

The Pauli observables σx\sigma_x, σy\sigma_y, and σz\sigma_z provide a natural way to reconstruct a qubit state, but they are incompatible with each other: they cannot be measured simultaneously on the same qubit, and a measurement disturbs the state. The NN qubits are therefore divided into three groups, with each group used to measure one observable, yielding estimates of the three expectation values ⟨σx⟩\langle\sigma_x\rangle, ⟨σy⟩\langle\sigma_y\rangle, and ⟨σz⟩\langle\sigma_z\rangle.

Measuring σx\sigma_x

{∣+⟩⟨+∣,∣−⟩⟨−∣}\{\lvert+\rangle\langle+\rvert,\lvert-\rangle\langle-\rvert\}

N/3N/3 qubits

The two outcomes are the eigenvectors of σx\sigma_x.

  • +1for each ∣+⟩⟨+∣\lvert+\rangle\langle+\rvert outcome
  • −1for each ∣−⟩⟨−∣\lvert-\rangle\langle-\rvert outcome

Expected value for each measurement:

Tr(σxρ)\mathrm{Tr}(\sigma_x\rho)

Measuring σy\sigma_y

{∣+i⟩⟨+i∣,∣−i⟩⟨−i∣}\{\lvert{+}i\rangle\langle{+}i\rvert,\lvert{-}i\rangle\langle{-}i\rvert\}

N/3N/3 qubits

The two outcomes are the eigenvectors of σy\sigma_y.

  • +1for each ∣+i⟩⟨+i∣\lvert{+}i\rangle\langle{+}i\rvert outcome
  • −1for each ∣−i⟩⟨−i∣\lvert{-}i\rangle\langle{-}i\rvert outcome

Expected value for each measurement:

Tr(σyρ)\mathrm{Tr}(\sigma_y\rho)

Measuring σz\sigma_z

{∣0⟩⟨0∣,∣1⟩⟨1∣}\{\lvert0\rangle\langle0\rvert,\lvert1\rangle\langle1\rvert\}

N/3N/3 qubits

The two outcomes are the eigenvectors of σz\sigma_z.

  • +1for each ∣0⟩⟨0∣\lvert0\rangle\langle0\rvert outcome
  • −1for each ∣1⟩⟨1∣\lvert1\rangle\langle1\rvert outcome

Expected value for each measurement:

Tr(σzρ)\mathrm{Tr}(\sigma_z\rho)

Reconstructing ρ\rho

ρ=I+Tr(σxρ) σx+Tr(σyρ) σy+Tr(σzρ) σz2\displaystyle\rho=\frac{I+\mathrm{Tr}(\sigma_x\rho)\,\sigma_x+\mathrm{Tr}(\sigma_y\rho)\,\sigma_y+\mathrm{Tr}(\sigma_z\rho)\,\sigma_z}{2}

Any qubit state can be written in the Pauli basis as

ρ=I+rxσx+ryσy+rzσz2,\displaystyle\rho=\frac{I+r_x\sigma_x+r_y\sigma_y+r_z\sigma_z}{2},

where rxr_x, ryr_y, and rzr_z are the three components of the state’s . For the state ρ\rho, these components are exactly the expectation values of the Pauli observables:

rx=Tr(σxρ),ry=Tr(σyρ),rz=Tr(σzρ).r_x=\mathrm{Tr}(\sigma_x\rho),\quad r_y=\mathrm{Tr}(\sigma_y\rho),\quad r_z=\mathrm{Tr}(\sigma_z\rho).

The measurement results therefore provide estimates of rxr_x, ryr_y, and rzr_z. Substituting them into the Pauli-basis expansion gives the reconstructed state above.

What finite measurements reveal

The reconstruction formula is exact only when true expectation values are available, while a real experiment yields merely finite samples. With NN copies divided into three groups, each group of N/3N/3 outcomes provides a sample average that estimates one expectation value. The typical error of such an average scales as 1/N1/\sqrt{N}, so increasing the number of copies improves the estimate but never makes it exact for any finite run.

More importantly, the three estimated coefficients need not correspond to a physical state. If rx2+ry2+rz2>1r_x^2+r_y^2+r_z^2>1, the reconstructed point lies outside the Bloch ball, and the resulting matrix has a negative eigenvalue — it is not a density matrix at all. This is why finite-data tomography is inherently approximate, and why practical reconstruction methods must do more than simply invert the sample averages.

State
|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
true stateestimate
rxr_x
true+0.60
est+0.60
ryr_y
true+0.50
est+0.72
rzr_z
true+0.62
est+0.48
-10+1
52°
40°
1.00
150
Copies per group50
∥r^∥\lVert\hat{r}\rVert1.05
ρ^=I+r^xσx+r^yσy+r^zσz2=(0.740.30−0.36i0.30+0.36i0.26)\displaystyle\hat{\rho}=\frac{I+\hat{r}_x\sigma_x+\hat{r}_y\sigma_y+\hat{r}_z\sigma_z}{2}={\small\begin{pmatrix}0.74&0.30-0.36i\\0.30+0.36i&0.26\end{pmatrix}}

Purifications

Purifications

are more difficult to work with than pure states because they represent statistical mixtures rather than a single state vector. A useful way to handle them is to represent a mixed state as part of a larger system whose overall state is pure. Purification formalizes this construction.

A purification of a density matrix ρ\rho on system X\mathsf{X} is a pure state ∣ψ⟩\lvert\psi\rangle of a larger composite system X⊗Y\mathsf{X}\otimes\mathsf{Y} such that, after ignoring the auxiliary system Y\mathsf{Y}, the state of X\mathsf{X} is exactly ρ\rho:

TrY(∣ψ⟩⟨ψ∣)=ρ.\mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\psi\rangle\langle\psi\rvert\bigr)=\rho.

Here TrY\mathrm{Tr}_{\mathsf{Y}} denotes the over Y\mathsf{Y}. The auxiliary system Y\mathsf{Y} can be thought of as containing degrees of freedom that are not accessible when only X\mathsf{X} is observed. If X\mathsf{X} and Y\mathsf{Y} are correlated, X\mathsf{X} can therefore appear mixed even though the joint state of X⊗Y\mathsf{X}\otimes\mathsf{Y} is pure.

State

mixed state of X

pure state of X and Y

ρX=0.50 ∣0⟩⟨0∣+0.50 ∣1⟩⟨1∣\rho_{\mathsf{X}}=0.50\,\lvert0\rangle\langle0\rvert+0.50\,\lvert1\rangle\langle1\rvert
∣ψ⟩=0.71 ∣00⟩+0.71 ∣11⟩\lvert\psi\rangle=0.71\,\lvert00\rangle+0.71\,\lvert11\rangle
⟨0∣\langle0\rvert⟨1∣\langle1\rvert
∣0⟩\lvert0\rangle0.500.00
∣1⟩\lvert1\rangle0.000.50
⟨00∣\langle00\rvert⟨01∣\langle01\rvert⟨10∣\langle10\rvert⟨11∣\langle11\rvert
∣00⟩\lvert00\rangle0.500.000.000.50
∣01⟩\lvert01\rangle0.000.000.000.00
∣10⟩\lvert10\rangle0.000.000.000.00
∣11⟩\lvert11\rangle0.500.000.000.50
|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
X
|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
X
|0⟩|1⟩|+⟩|−⟩|i⟩|−i⟩
Y
0.50 / 0.50
TrY(∣ψ⟩⟨ψ∣)=(0.50000.50)=ρX\mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\psi\rangle\langle\psi\rvert\bigr)={\small\begin{pmatrix}0.50&0\\0&0.50\end{pmatrix}}=\rho_{\mathsf{X}}

Purifications are useful because they allow mixed-state problems to be studied through a larger pure state, where tools such as entanglement and unitary evolution can often be applied more directly.

Existence of purifications

Let X\mathsf{X} be a quantum system and let ρ\rho be a density matrix describing a state of X\mathsf{X}. By definition, ρ\rho can be written as a of pure states, for some probability vector (p0,…,pn−1)(p_0,\ldots,p_{n-1}) and some state vectors ∣ϕ0⟩,…,∣ϕn−1⟩\lvert\phi_0\rangle,\ldots,\lvert\phi_{n-1}\rangle of X\mathsf{X}:

ρ=∑a=0n−1pa ∣ϕa⟩⟨ϕa∣.\displaystyle\rho=\sum_{a=0}^{n-1}p_a\,\lvert\phi_a\rangle\langle\phi_a\rvert.

This decomposition already tells us how to construct a purification. Introduce an auxiliary system Y\mathsf{Y} whose classical states are labelled 0,…,n−10,\ldots,n-1, and pair each term of the mixture with a distinct classical state ∣a⟩\lvert a\rangle of that system, so that the probability pap_a becomes the square of an amplitude:

∣ψ⟩=∑a=0n−1pa ∣ϕa⟩⊗∣a⟩.\displaystyle\lvert\psi\rangle=\sum_{a=0}^{n-1}\sqrt{p_a}\,\lvert\phi_a\rangle\otimes\lvert a\rangle.

To see that ∣ψ⟩\lvert\psi\rangle is indeed a purification of ρ\rho, trace out Y\mathsf{Y}. The cross terms vanish because Tr(∣a⟩⟨b∣)=δab\mathrm{Tr}(\lvert a\rangle\langle b\rvert)=\delta_{ab}, which is one when a=ba=b and zero otherwise, so only the diagonal terms survive and the original mixture comes back:

TrY(∣ψ⟩⟨ψ∣)=TrY((∑a=0n−1pa ∣ϕa⟩⊗∣a⟩)(∑b=0n−1pb ⟨ϕb∣⊗⟨b∣))the state vector and its conjugate, written out=∑a,b=0n−1papb ∣ϕa⟩⟨ϕb∣ Tr(∣a⟩⟨b∣)each term splits, and the trace falls on Y=∑a,b=0n−1papb ∣ϕa⟩⟨ϕb∣ δabTr(∣a⟩⟨b∣)=⟨b∣a⟩=δab=∑a=0n−1pa ∣ϕa⟩⟨ϕa∣only the terms with a=b survive=ρthe mixture we started from\displaystyle\begin{aligned} \mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\psi\rangle\langle\psi\rvert\bigr) &=\mathrm{Tr}_{\mathsf{Y}}\Bigl(\Bigl(\sum_{a=0}^{n-1}\sqrt{p_a}\,\lvert\phi_a\rangle\otimes\lvert a\rangle\Bigr)\Bigl(\sum_{b=0}^{n-1}\sqrt{p_b}\,\langle\phi_b\rvert\otimes\langle b\rvert\Bigr)\Bigr) &&\quad\textcolor{#94a3b8}{\text{the state vector and its conjugate, written out}}\\[4pt] &=\sum_{a,b=0}^{n-1}\sqrt{p_ap_b}\,\lvert\phi_a\rangle\langle\phi_b\rvert\,\mathrm{Tr}\bigl(\lvert a\rangle\langle b\rvert\bigr) &&\quad\textcolor{#94a3b8}{\text{each term splits, and the trace falls on }\mathsf{Y}}\\[4pt] &=\sum_{a,b=0}^{n-1}\sqrt{p_ap_b}\,\lvert\phi_a\rangle\langle\phi_b\rvert\,\delta_{ab} &&\quad\textcolor{#94a3b8}{\mathrm{Tr}\bigl(\lvert a\rangle\langle b\rvert\bigr)=\langle b\rvert a\rangle=\delta_{ab}}\\[4pt] &=\sum_{a=0}^{n-1}p_a\,\lvert\phi_a\rangle\langle\phi_a\rvert &&\quad\textcolor{#94a3b8}{\text{only the terms with }a=b\text{ survive}}\\[4pt] &=\rho &&\quad\textcolor{#94a3b8}{\text{the mixture we started from}} \end{aligned}

Finally, every density matrix has at least one such decomposition — its — with at most as many terms as the dimension of X\mathsf{X}. Therefore every state of X\mathsf{X} has a purification, provided that Y\mathsf{Y} has at least as many classical states as X\mathsf{X} does.

Schmidt decomposition

A pure state of a bipartite system can contain correlations between its two subsystems. The Schmidt decomposition provides a useful way to make these correlations explicit by expressing the state as a sum of paired states of the two subsystems.

Every state vector ∣ψ⟩\lvert\psi\rangle of a bipartite system (X,Y)(\mathsf{X},\mathsf{Y}) can be written in the form:

∣ψ⟩=∑a=0r−1pa ∣xa⟩⊗∣ya⟩.\displaystyle\lvert\psi\rangle=\sum_{a=0}^{r-1}\sqrt{p_a}\,\lvert x_a\rangle\otimes\lvert y_a\rangle.
  • The coefficients p0,…,pr−1p_0,\ldots,p_{r-1} are strictly positive and satisfy ∑apa=1\sum_a p_a=1.
  • The sets {∣x0⟩,…,∣xr−1⟩}\{\lvert x_0\rangle,\ldots,\lvert x_{r-1}\rangle\} and {∣y0⟩,…,∣yr−1⟩}\{\lvert y_0\rangle,\ldots,\lvert y_{r-1}\rangle\} are . They do not necessarily span the full state spaces of X\mathsf{X} and Y\mathsf{Y}, since only the subspaces involved in ∣ψ⟩\lvert\psi\rangle are needed.
  • The number of terms rr is the Schmidt rank, and it can be no larger than the dimension of either system. A Schmidt rank of one means the state is a product state, while a larger rank indicates correlations between X\mathsf{X} and Y\mathsf{Y}.

Constructing the decomposition

To find the Schmidt decomposition, first extract the coefficients and basis states on one side, then use the original state to recover the corresponding states on the other side.

  1. Compute the of the reduced state ρX=TrY(∣ψ⟩⟨ψ∣)\rho_{\mathsf{X}}=\mathrm{Tr}_{\mathsf{Y}}(\lvert\psi\rangle\langle\psi\rvert).

    Keep only the rr strictly positive eigenvalues pap_a and their corresponding eigenvectors ∣xa⟩\lvert x_a\rangle:

    ρX=∑a=0r−1pa ∣xa⟩⟨xa∣.\displaystyle\rho_{\mathsf{X}}=\sum_{a=0}^{r-1}p_a\,\lvert x_a\rangle\langle x_a\rvert.
  2. For each aa, project ∣ψ⟩\lvert\psi\rangle onto ∣xa⟩\lvert x_a\rangle and normalize the resulting state on Y\mathsf{Y}:

    ∣ya⟩=(⟨xa∣⊗I)∣ψ⟩pa.\displaystyle\lvert y_a\rangle=\frac{(\langle x_a\rvert\otimes I)\lvert\psi\rangle}{\sqrt{p_a}}.
The vectors ∣ya⟩\lvert y_a\rangle are orthonormal by construction

Take the inner product of two vectors ∣ya⟩\lvert y_a\rangle and ∣yb⟩\lvert y_b\rangle using their definition above. The resulting expression can be written in terms of the reduced state ρX\rho_{\mathsf{X}}, whose eigenvectors ∣xa⟩\lvert x_a\rangle are orthonormal:

⟨yb∣ya⟩=((⟨xb∣⊗I)∣ψ⟩pb)†((⟨xa∣⊗I)∣ψ⟩pa)each vector by its definition=⟨ψ∣(∣xb⟩⊗I)(⟨xa∣⊗I)∣ψ⟩papb(⟨xb∣⊗I)†=∣xb⟩⊗I=⟨ψ∣(∣xb⟩⟨xa∣⊗I)∣ψ⟩papb(A⊗I)(B⊗I)=AB⊗I=Tr(∣xb⟩⟨xa∣ ρX)papb⟨ψ∣(A⊗I)∣ψ⟩=Tr(AρX)=⟨xa∣ρX∣xb⟩papbTr(∣xb⟩⟨xa∣M)=⟨xa∣M∣xb⟩=pb ⟨xa∣xb⟩papbρX∣xb⟩=pb∣xb⟩=δab\displaystyle\begin{aligned} \langle y_b\rvert y_a\rangle &=\Biggl(\frac{(\langle x_b\rvert\otimes I)\lvert\psi\rangle}{\sqrt{p_b}}\Biggr)^{\dagger}\Biggl(\frac{(\langle x_a\rvert\otimes I)\lvert\psi\rangle}{\sqrt{p_a}}\Biggr) &&\quad\textcolor{#94a3b8}{\text{each vector by its definition}}\\[4pt] &=\frac{\langle\psi\rvert(\lvert x_b\rangle\otimes I)(\langle x_a\rvert\otimes I)\lvert\psi\rangle}{\sqrt{p_ap_b}} &&\quad\textcolor{#94a3b8}{(\langle x_b\rvert\otimes I)^{\dagger}=\lvert x_b\rangle\otimes I}\\[4pt] &=\frac{\langle\psi\rvert(\lvert x_b\rangle\langle x_a\rvert\otimes I)\lvert\psi\rangle}{\sqrt{p_ap_b}} &&\quad\textcolor{#94a3b8}{(A\otimes I)(B\otimes I)=AB\otimes I}\\[4pt] &=\frac{\mathrm{Tr}\bigl(\lvert x_b\rangle\langle x_a\rvert\,\rho_{\mathsf{X}}\bigr)}{\sqrt{p_ap_b}} &&\quad\textcolor{#94a3b8}{\langle\psi\rvert(A\otimes I)\lvert\psi\rangle=\mathrm{Tr}(A\rho_{\mathsf{X}})}\\[4pt] &=\frac{\langle x_a\rvert\rho_{\mathsf{X}}\lvert x_b\rangle}{\sqrt{p_ap_b}} &&\quad\textcolor{#94a3b8}{\mathrm{Tr}\bigl(\lvert x_b\rangle\langle x_a\rvert M\bigr)=\langle x_a\rvert M\lvert x_b\rangle}\\[4pt] &=\frac{p_b\,\langle x_a\rvert x_b\rangle}{\sqrt{p_ap_b}} &&\quad\textcolor{#94a3b8}{\rho_{\mathsf{X}}\lvert x_b\rangle=p_b\lvert x_b\rangle}\\[4pt] &=\delta_{ab} \end{aligned}
State

the state of the pair we start with

∣ψ⟩=12 ∣0⟩⏟β⊗∣0⟩⏟fixed+12 ∣+⟩⏟γ⊗∣1⟩⏟fixed⏞X⊗Y\lvert\psi\rangle=\overbrace{\tfrac{1}{\sqrt2}\,\underbrace{\lvert0\rangle}_{\beta}\otimes\underbrace{\lvert0\rangle}_{\text{fixed}}+\tfrac{1}{\sqrt2}\,\underbrace{\lvert{+}\rangle}_{\gamma}\otimes\underbrace{\lvert1\rangle}_{\text{fixed}}}^{\textstyle\mathsf{X}\otimes\mathsf{Y}}
βγ|0⟩|1⟩first qubit, X|0⟩|1⟩second qubit, Y
0°
45°
45°

applying Schmidt decomposition

  1. trace out Y\mathsf{Y} to get the reduced state:

    ρX=TrY(∣ψ⟩⟨ψ∣)\rho_{\mathsf{X}}=\mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\psi\rangle\langle\psi\rvert\bigr)
    ρX=TrY((12 ∣β⟩⊗∣0⟩+12 ∣γ⟩⊗∣1⟩)ρX=TrY((12 ⟨β∣⊗⟨0∣+12 ⟨γ∣⊗⟨1∣))\begin{aligned}&\phantom{\rho_{\mathsf{X}}}{}=\mathrm{Tr}_{\mathsf{Y}}\Bigl(\bigl(\tfrac{1}{\sqrt2}\,\lvert\beta\rangle\otimes\lvert0\rangle+\tfrac{1}{\sqrt2}\,\lvert\gamma\rangle\otimes\lvert1\rangle\bigr)\\[2pt]&\phantom{\rho_{\mathsf{X}}}\phantom{{}=\mathrm{Tr}_{\mathsf{Y}}\Bigl(}\bigl(\tfrac{1}{\sqrt2}\,\langle\beta\rvert\otimes\langle0\rvert+\tfrac{1}{\sqrt2}\,\langle\gamma\rvert\otimes\langle1\rvert\bigr)\Bigr)\end{aligned}
    ρX=TrY(0.50 ∣β⟩⟨β∣⊗∣0⟩⟨0∣+0.50 ∣β⟩⟨γ∣⊗∣0⟩⟨1∣ρX=TrY(+0.50 ∣γ⟩⟨β∣⊗∣1⟩⟨0∣+0.50 ∣γ⟩⟨γ∣⊗∣1⟩⟨1∣)\begin{aligned}&\phantom{\rho_{\mathsf{X}}}{}=\mathrm{Tr}_{\mathsf{Y}}\Bigl(0.50\,\lvert\beta\rangle\langle\beta\rvert\otimes\lvert0\rangle\langle0\rvert+0.50\,\lvert\beta\rangle\langle\gamma\rvert\otimes\lvert0\rangle\langle1\rvert\\[2pt]&\phantom{\rho_{\mathsf{X}}}\phantom{{}=\mathrm{Tr}_{\mathsf{Y}}\Bigl(}{}+0.50\,\lvert\gamma\rangle\langle\beta\rvert\otimes\lvert1\rangle\langle0\rvert+0.50\,\lvert\gamma\rangle\langle\gamma\rvert\otimes\lvert1\rangle\langle1\rvert\Bigr)\end{aligned}
    ρX=0.50 ∣β⟩⟨β∣ Tr(∣0⟩⟨0∣)+0.50 ∣β⟩⟨γ∣ Tr(∣0⟩⟨1∣)ρX=+0.50 ∣γ⟩⟨β∣ Tr(∣1⟩⟨0∣)+0.50 ∣γ⟩⟨γ∣ Tr(∣1⟩⟨1∣)\begin{aligned}&\phantom{\rho_{\mathsf{X}}}{}={}0.50\,\lvert\beta\rangle\langle\beta\rvert\,\mathrm{Tr}\bigl(\lvert0\rangle\langle0\rvert\bigr)+0.50\,\lvert\beta\rangle\langle\gamma\rvert\,\mathrm{Tr}\bigl(\lvert0\rangle\langle1\rvert\bigr)\\[2pt]&\phantom{\rho_{\mathsf{X}}}\phantom{{}={}}{}+0.50\,\lvert\gamma\rangle\langle\beta\rvert\,\mathrm{Tr}\bigl(\lvert1\rangle\langle0\rvert\bigr)+0.50\,\lvert\gamma\rangle\langle\gamma\rvert\,\mathrm{Tr}\bigl(\lvert1\rangle\langle1\rvert\bigr)\end{aligned}
    ρX=0.50 ∣β⟩⟨β∣ ⟨0∣0⟩+0.50 ∣β⟩⟨γ∣ ⟨1∣0⟩ρX=+0.50 ∣γ⟩⟨β∣ ⟨0∣1⟩+0.50 ∣γ⟩⟨γ∣ ⟨1∣1⟩\begin{aligned}&\phantom{\rho_{\mathsf{X}}}{}={}0.50\,\lvert\beta\rangle\langle\beta\rvert\,\langle0\vert0\rangle+0.50\,\lvert\beta\rangle\langle\gamma\rvert\,\langle1\vert0\rangle\\[2pt]&\phantom{\rho_{\mathsf{X}}}\phantom{{}={}}{}+0.50\,\lvert\gamma\rangle\langle\beta\rvert\,\langle0\vert1\rangle+0.50\,\lvert\gamma\rangle\langle\gamma\rvert\,\langle1\vert1\rangle\end{aligned}
    ρX=0.50 ∣β⟩⟨β∣+0.50 ∣γ⟩⟨γ∣\phantom{\rho_{\mathsf{X}}}=0.50\,\lvert\beta\rangle\langle\beta\rvert+0.50\,\lvert\gamma\rangle\langle\gamma\rvert
    ρX=0.50 (1.000.000.000.00)+0.50 (0.500.500.500.50)\phantom{\rho_{\mathsf{X}}}=0.50\,\begin{pmatrix}1.00&0.00\\0.00&0.00\end{pmatrix}+0.50\,\begin{pmatrix}0.50&0.50\\0.50&0.50\end{pmatrix}
    ρX=(0.750.250.250.25)\phantom{\rho_{\mathsf{X}}}=\begin{pmatrix}0.75&0.25\\0.25&0.25\end{pmatrix}

    diagonalise ρX\rho_{\mathsf{X}} to get the perpendicular pair:

    det⁡(ρX−λI)=0\det\bigl(\rho_{\mathsf{X}}-\lambda I\bigr)=0
    det⁡(ρX−λI)=λ2−λ+0.12\phantom{\det\bigl(\rho_{\mathsf{X}}-\lambda I\bigr)}=\lambda^{2}-\lambda+0.12
    det⁡(ρX−λI)⇒  λ=0.85,  0.15\phantom{\det\bigl(\rho_{\mathsf{X}}-\lambda I\bigr)}\Rightarrow\;\lambda=0.85,\;0.15

    solve (ρX−paI)∣xa⟩=0\bigl(\rho_{\mathsf{X}}-p_a I\bigr)\lvert x_a\rangle=0 for each direction:

    p0=0.85∣x0⟩=0.92 ∣0⟩+0.38 ∣1⟩p1=0.15∣x1⟩=−0.38 ∣0⟩+0.92 ∣1⟩\begin{aligned}p_0&=0.85 & \textcolor{#7c3aed}{\lvert x_0\rangle}&=0.92\,\lvert0\rangle+0.38\,\lvert1\rangle\\p_1&=0.15 & \textcolor{#0284c7}{\lvert x_1\rangle}&=-0.38\,\lvert0\rangle+0.92\,\lvert1\rangle\end{aligned}
  2. recover the matching vectors on Y\mathsf{Y}:

    ∣ya⟩=(⟨xa∣⊗I)∣ψ⟩pa\lvert y_a\rangle=\dfrac{(\langle x_a\rvert\otimes I)\lvert\psi\rangle}{\sqrt{p_a}}
    ∣y0⟩=0.65 ∣0⟩+0.65 ∣1⟩0.92=0.71 ∣0⟩+0.71 ∣1⟩∣y1⟩=−0.27 ∣0⟩+0.27 ∣1⟩0.38=−0.71 ∣0⟩+0.71 ∣1⟩\begin{aligned}\textcolor{#7c3aed}{\lvert y_0\rangle}&=\dfrac{0.65\,\lvert0\rangle+0.65\,\lvert1\rangle}{0.92}=0.71\,\lvert0\rangle+0.71\,\lvert1\rangle\\\textcolor{#0284c7}{\lvert y_1\rangle}&=\dfrac{-0.27\,\lvert0\rangle+0.27\,\lvert1\rangle}{0.38}=-0.71\,\lvert0\rangle+0.71\,\lvert1\rangle\end{aligned}

the same state in Schmidt form

∣ψ⟩=0.92 ∣x0⟩⊗∣y0⟩⏟p0=0.85+0.38 ∣x1⟩⊗∣y1⟩⏟p1=0.15⏞X⊗Y\lvert\psi\rangle=\overbrace{\underbrace{0.92\,\textcolor{#7c3aed}{\lvert x_0\rangle\otimes\lvert y_0\rangle}}_{p_0=0.85}+\underbrace{0.38\,\textcolor{#0284c7}{\lvert x_1\rangle\otimes\lvert y_1\rangle}}_{p_1=0.15}}^{\textstyle\mathsf{X}\otimes\mathsf{Y}}
βγx₀x₁|0⟩|1⟩first qubit, Xy₀y₁|0⟩|1⟩second qubit, Y
Correlation71%
product stateevenly shared

For two systems, the Schmidt decomposition gives a particularly clean picture: the joint state is written as a sum of matching orthonormal directions on X\mathsf{X} and Y\mathsf{Y}, with the coefficients pa\sqrt{p_a} showing how much weight each pair carries.

The same idea extends beyond two qubits. For any bipartite state, even when X\mathsf{X} and Y\mathsf{Y} are larger systems, the state can still be decomposed into matching orthonormal sets with one coefficient for each pair. With more than two qubits, what matters is how the qubits are divided into the two sides of the bipartition. For example, three qubits can be split into one qubit in X\mathsf{X} and two qubits in Y\mathsf{Y}, and the same decomposition applies to that split.

Unitary equivalence of purifications

A contains more information than the density matrix it represents: the density matrix describes system X\mathsf{X}, while the purification also specifies how X\mathsf{X} is correlated with an auxiliary system Y\mathsf{Y}. Different purifications can therefore look different, but the difference lies entirely in the choice of states on Y\mathsf{Y}. Any two purifications of the same state on X\mathsf{X} are related by a unitary acting only on Y\mathsf{Y}.

Let ∣ψ⟩\lvert\psi\rangle and ∣ϕ⟩\lvert\phi\rangle be two pure states of the composite system (X,Y)(\mathsf{X},\mathsf{Y}) with identical reduced states:

TrY(∣ψ⟩⟨ψ∣)=ρ=TrY(∣ϕ⟩⟨ϕ∣).\displaystyle \mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\psi\rangle\langle\psi\rvert\bigr)=\rho=\mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\phi\rangle\langle\phi\rvert\bigr).

Choose a of ρ\rho:

ρ=∑a=0r−1pa ∣xa⟩⟨xa∣.\displaystyle \rho=\sum_{a=0}^{r-1}p_a\,\lvert x_a\rangle\langle x_a\rvert.

The then gives both purifications in terms of the same eigenvalues and the same orthonormal vectors on X\mathsf{X}:

∣ψ⟩=∑a=0r−1pa ∣xa⟩⊗∣ya⟩,\displaystyle \lvert\psi\rangle=\sum_{a=0}^{r-1}\sqrt{p_a}\,\lvert x_a\rangle\otimes\lvert y_a\rangle,
∣ϕ⟩=∑a=0r−1pa ∣xa⟩⊗∣za⟩.\displaystyle \lvert\phi\rangle=\sum_{a=0}^{r-1}\sqrt{p_a}\,\lvert x_a\rangle\otimes\lvert z_a\rangle.

Thus, the only difference between the two purifications is the choice of the states {∣ya⟩}\{\lvert y_a\rangle\} and {∣za⟩}\{\lvert z_a\rangle\} on Y\mathsf{Y}.

These sets may not span all of Y\mathsf{Y}. Add normalized vectors orthogonal to all the vectors already present until each set contains enough vectors to span the whole space Y\mathsf{Y}. This produces two full orthonormal bases:

{∣y0⟩,…,∣yr−1⟩,∣yr⟩,…}\displaystyle \{\lvert y_0\rangle,\ldots,\lvert y_{r-1}\rangle,\lvert y_r\rangle,\ldots\}
{∣z0⟩,…,∣zr−1⟩,∣zr⟩,…}.\displaystyle \{\lvert z_0\rangle,\ldots,\lvert z_{r-1}\rangle,\lvert z_r\rangle,\ldots\}.

Now define UU by mapping each vector in the first basis to the corresponding vector in the second:

Because a basis determines every vector in the space, this defines UU on all of Y\mathsf{Y}. For any state ∣v⟩=∑jcj∣yj⟩\lvert v\rangle=\sum_j c_j\lvert y_j\rangle, the map gives

U∣v⟩=∑jcj∣zj⟩\displaystyle U\lvert v\rangle=\sum_j c_j\lvert z_j\rangle

Since both {∣yj⟩}\{\lvert y_j\rangle\} and {∣zj⟩}\{\lvert z_j\rangle\} are orthonormal, all cross terms vanish in the inner products, leaving only the squared magnitudes of the coefficients:

⟨v∣v⟩=∑j∣cj∣2,\displaystyle \langle v\vert v\rangle=\sum_j\lvert c_j\rvert^2,
⟨Uv∣Uv⟩=(∑jcj‾⟨zj∣)(∑kck∣zk⟩)=∑j∣cj∣2.\displaystyle \langle Uv\vert Uv\rangle=\Bigl(\sum_j\overline{c_j}\langle z_j\rvert\Bigr)\Bigl(\sum_k c_k\lvert z_k\rangle\Bigr)=\sum_j\lvert c_j\rvert^2.

Thus UU preserves norms—and, by the same reasoning, all inner products. This is precisely the defining property of a unitary operator. In particular,

U∣ya⟩=∣za⟩for every a=0,…,r−1.\displaystyle U\lvert y_a\rangle=\lvert z_a\rangle\quad\text{for every }a=0,\ldots,r-1.

Applying UU to the auxiliary system Y\mathsf{Y} transforms one purification into the other:

(IX⊗U)∣ψ⟩=(IX⊗U)∑a=0r−1pa ∣xa⟩⊗∣ya⟩=∑a=0r−1pa (IX∣xa⟩)⊗(U∣ya⟩)=∑a=0r−1pa ∣xa⟩⊗U∣ya⟩=∑a=0r−1pa ∣xa⟩⊗∣za⟩=∣ϕ⟩.\displaystyle \begin{aligned}(I_{\mathsf{X}}\otimes U)\lvert\psi\rangle&=(I_{\mathsf{X}}\otimes U)\sum_{a=0}^{r-1}\sqrt{p_a}\,\lvert x_a\rangle\otimes\lvert y_a\rangle\\[6pt]&=\sum_{a=0}^{r-1}\sqrt{p_a}\,\bigl(I_{\mathsf{X}}\lvert x_a\rangle\bigr)\otimes\bigl(U\lvert y_a\rangle\bigr)\\[6pt]&=\sum_{a=0}^{r-1}\sqrt{p_a}\,\lvert x_a\rangle\otimes U\lvert y_a\rangle\\[6pt]&=\sum_{a=0}^{r-1}\sqrt{p_a}\,\lvert x_a\rangle\otimes\lvert z_a\rangle\\[6pt]&=\lvert\phi\rangle.\end{aligned}

Note that the unitary is not necessarily unique. The purification only contains the vectors {∣ya⟩}\{\lvert y_a\rangle\}, so UU is fixed only by how it maps those vectors to {∣za⟩}\{\lvert z_a\rangle\}. Its action on the remaining directions of Y\mathsf{Y} can be chosen in different ways without changing the result.

So, unitary equivalence of purifications means that any purification of ρ\rho can be transformed into any other purification of ρ\rho by applying a suitable unitary to the auxiliary system Y\mathsf{Y} alone.

Superdense coding as unitary equivalence

provides a concrete example of the unitary equivalence of purifications.

Alice holds qubit A\mathsf{A}, and Bob holds qubit B\mathsf{B}. They share an entangled pair, initially in the Bell state:

∣ϕ+⟩=12(∣00⟩+∣11⟩).\displaystyle \lvert\phi^{+}\rangle=\tfrac{1}{\sqrt2}\bigl(\lvert00\rangle+\lvert11\rangle\bigr).

To encode two classical bits, Alice applies a unitary to her qubit, choosing one of four operations according to the value being encoded. The resulting state is one of the four Bell states:

∣ϕ+⟩,∣ϕ−⟩,∣ψ+⟩,∣ψ−⟩.\displaystyle \lvert\phi^{+}\rangle,\quad\lvert\phi^{-}\rangle,\quad\lvert\psi^{+}\rangle,\quad\lvert\psi^{-}\rangle.

From Bob’s perspective, however, these states are identical. Tracing out Alice’s system gives

TrA(∣ϕ+⟩⟨ϕ+∣)\mathrm{Tr}_{\mathsf{A}}\bigl(\lvert\phi^{+}\rangle\langle\phi^{+}\rvert\bigr)=12I,=\tfrac12 I,

and the same calculation holds for all four Bell states:

TrA(∣ϕ+⟩⟨ϕ+∣)\mathrm{Tr}_{\mathsf{A}}\bigl(\lvert\phi^{+}\rangle\langle\phi^{+}\rvert\bigr)==TrA(∣ϕ−⟩⟨ϕ−∣)\mathrm{Tr}_{\mathsf{A}}\bigl(\lvert\phi^{-}\rangle\langle\phi^{-}\rvert\bigr)==TrA(∣ψ+⟩⟨ψ+∣)\mathrm{Tr}_{\mathsf{A}}\bigl(\lvert\psi^{+}\rangle\langle\psi^{+}\rvert\bigr)==TrA(∣ψ−⟩⟨ψ−∣)\mathrm{Tr}_{\mathsf{A}}\bigl(\lvert\psi^{-}\rangle\langle\psi^{-}\rvert\bigr)==12I.\tfrac12 I.

Thus, the four Bell states are different of the same density matrix on B\mathsf{B}. By , any two of them are related by a unitary acting on the purifying system A\mathsf{A}.

For superdense coding, these unitaries are the Pauli operations used to encode the two classical bits:

(I⊗IB)∣ϕ+⟩=∣ϕ+⟩\displaystyle (I\otimes I_{\mathsf{B}})\lvert\phi^{+}\rangle=\lvert\phi^{+}\rangle
(Z⊗IB)∣ϕ+⟩=∣ϕ−⟩\displaystyle (Z\otimes I_{\mathsf{B}})\lvert\phi^{+}\rangle=\lvert\phi^{-}\rangle
(X⊗IB)∣ϕ+⟩=∣ψ+⟩\displaystyle (X\otimes I_{\mathsf{B}})\lvert\phi^{+}\rangle=\lvert\psi^{+}\rangle
(XZ⊗IB)∣ϕ+⟩=−∣ψ−⟩.\displaystyle (XZ\otimes I_{\mathsf{B}})\lvert\phi^{+}\rangle=-\lvert\psi^{-}\rangle.

This gives the structural reason behind the encoding step: because the four Bell states are purifications of the same state on B\mathsf{B}, a unitary on Alice’s system alone can move between them. Bob’s remains 12I\tfrac12 I throughout, so his qubit contains no information about which Bell state was chosen until Alice’s encoded qubit is received.

Hughston-Jozsa-Wootters theorem

Suppose X\mathsf{X} and Y\mathsf{Y} are systems and ∣ϕ⟩\lvert\phi\rangle is a quantum state vector of (X,Y)(\mathsf{X},\mathsf{Y}). Let NN be a positive integer, let (p0,…,pN−1)(p_0,\ldots,p_{N-1}) be a , and let ∣ψ0⟩,…,∣ψN−1⟩\lvert\psi_0\rangle,\ldots,\lvert\psi_{N-1}\rangle be quantum state vectors of X\mathsf{X} such that

TrY(∣ϕ⟩⟨ϕ∣)=∑a=0N−1pa∣ψa⟩⟨ψa∣.\displaystyle \mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\phi\rangle\langle\phi\rvert\bigr)=\sum_{a=0}^{N-1}p_a\lvert\psi_a\rangle\langle\psi_a\rvert.

There exists a {P0,…,PN−1}\{P_0,\ldots,P_{N-1}\} of Y\mathsf{Y} such that these statements are true when Y\mathsf{Y} is measured while (X,Y)(\mathsf{X},\mathsf{Y}) is in the state ∣ϕ⟩\lvert\phi\rangle:

  • Each measurement outcome a∈{0,…,N−1}a\in\{0,\ldots,N-1\} appears with probability pap_a.
  • Conditioned on obtaining the outcome aa, the state of X\mathsf{X} becomes ∣ψa⟩\lvert\psi_a\rangle.

An ensemble is a collection of pure states of X\mathsf{X}, together with the probabilities of preparing them, written as {(pa,∣ψa⟩)}\{(p_a,\lvert\psi_a\rangle)\}. It represents the density matrix ρ=∑apa∣ψa⟩⟨ψa∣\rho=\sum_{a}p_a\lvert\psi_a\rangle\langle\psi_a\rvert.

Different ensembles can represent the same state of X\mathsf{X}, meaning they give the same density matrix ρ\rho. Thus a mixed state of X\mathsf{X} does not have a unique decomposition into pure states. The HJW theorem says that, given a of ρ\rho on (X,Y)(\mathsf{X},\mathsf{Y}), every such ensemble can be realised by a suitable measurement on the purifying system Y\mathsf{Y}. The measurement outcome aa tells us that X\mathsf{X} is in the corresponding pure state ∣ψa⟩\lvert\psi_a\rangle, with probability pap_a. Before the outcome is known, the state of X\mathsf{X} is still described by the same density matrix ρ\rho, regardless of which measurement is chosen on Y\mathsf{Y}.

From an ensemble to a measurement

The hypothesis gives two descriptions of the same state of X\mathsf{X}:

∑a=0N−1pa∣ψa⟩⟨ψa∣=ρ=TrY(∣ϕ⟩⟨ϕ∣).\displaystyle \sum_{a=0}^{N-1}p_a\lvert\psi_a\rangle\langle\psi_a\rvert=\rho=\mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\phi\rangle\langle\phi\rvert\bigr).

The goal is to construct a measurement on Y\mathsf{Y} whose outcome aa occurs with probability pap_a and leaves X\mathsf{X} in the corresponding pure state ∣ψa⟩\lvert\psi_a\rangle.

1. Record the ensemble in a new system

We first need a way to keep track of which pure state in the ensemble was selected. Introduce a new system Z\mathsf{Z} with orthonormal states ∣0⟩, …, ∣N−1⟩\lvert0\rangle,\ \ldots,\ \lvert N-1\rangle, using one state ∣a⟩Z\lvert a\rangle_{\mathsf{Z}} as a label for each pure state ∣ψa⟩X\lvert\psi_a\rangle_{\mathsf{X}}. Now put the label together with the corresponding state of X\mathsf{X}:

∣γ1⟩=∑a=0N−1pa ∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z.\displaystyle \lvert\gamma_1\rangle=\sum_{a=0}^{N-1}\sqrt{p_a}\,\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}}.

The sum is easiest to read one term at a time. Fix a single aa, and the term

pa ∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z\displaystyle \sqrt{p_a}\,\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\textcolor{#4f46e5}{\lvert a\rangle_{\mathsf{Z}}}

is a product of three factors, one for each system:

  • ∣ψa⟩X\lvert\psi_a\rangle_{\mathsf{X}} is the member of the ensemble this term carries, one of the pure states ρ\rho decomposes into.
  • ∣a⟩Z\lvert a\rangle_{\mathsf{Z}} is the label recording which member that is.
  • ∣0⟩Y\lvert0\rangle_{\mathsf{Y}} holds the purifying system in a fixed state. It does not depend on aa and takes no part in the record, so it is along for the ride.

In front of them sits pa\sqrt{p_a}, which is an amplitude rather than a probability. Measuring Z\mathsf{Z} in its own basis gives the term carrying the label aa probability ∣pa∣2=pa\lvert\sqrt{p_a}\rvert^2=p_a, which is exactly the weight that term has in the ensemble.

So ∣γ1⟩\lvert\gamma_1\rangle pairs each pure state ∣ψa⟩\lvert\psi_a\rangle of X\mathsf{X} with its own label ∣a⟩\lvert a\rangle in Z\mathsf{Z}, and gives that pairing probability pap_a.

The important question is whether this larger state still represents the same state ρ\rho on X\mathsf{X}. It does. If we ignore both Y\mathsf{Y} and the record Z\mathsf{Z}, we should recover exactly the original mixed state.

To see this, start from the density matrix of ∣γ1⟩\lvert\gamma_1\rangle. Multiplying the sum by its own adjoint pairs every term with every other, so the result is a double sum indexed by aa and bb:

∣γ1⟩⟨γ1∣=∑a,b=0N−1papb ∣ψa⟩⟨ψb∣⊗∣0⟩⟨0∣⊗∣a⟩⟨b∣.\displaystyle \lvert\gamma_1\rangle\langle\gamma_1\rvert=\sum_{a,b=0}^{N-1}\sqrt{p_a p_b}\,\lvert\psi_a\rangle\langle\psi_b\rvert\otimes\lvert0\rangle\langle0\rvert\otimes\lvert a\rangle\langle b\rvert.

Every factor is now separated by system, so the trace over Y\mathsf{Y} and Z\mathsf{Z} passes straight through the X\mathsf{X} factor and closes the other two into numbers, using Tr(∣u⟩⟨v∣)=⟨v∣u⟩\mathrm{Tr}\bigl(\lvert u\rangle\langle v\rvert\bigr)=\langle v\vert u\rangle:

TrYZ(∣γ1⟩⟨γ1∣)=∑a,b=0N−1papb ∣ψa⟩⟨ψb∣  ⟨0∣0⟩  ⟨b∣a⟩.\displaystyle \mathrm{Tr}_{\mathsf{YZ}}\bigl(\lvert\gamma_1\rangle\langle\gamma_1\rvert\bigr)=\sum_{a,b=0}^{N-1}\sqrt{p_a p_b}\,\lvert\psi_a\rangle\langle\psi_b\rvert\;\langle0\vert0\rangle\;\langle b\vert a\rangle.

Both numbers are easy to read off. Y\mathsf{Y} carries the same ∣0⟩\lvert0\rangle in every term, so ⟨0∣0⟩=1\langle0\vert0\rangle=1, and the labels are orthonormal, so ⟨b∣a⟩=δab\langle b\vert a\rangle=\delta_{ab} is 11 when a=ba=b and 00 otherwise. Every off-diagonal term therefore disappears and the double sum collapses back to a single one:

TrYZ(∣γ1⟩⟨γ1∣)=∑a,b=0N−1papb ∣ψa⟩⟨ψb∣  ⟨0∣0⟩  ⟨b∣a⟩=∑a,b=0N−1papb ∣ψa⟩⟨ψb∣  δab=∑a=0N−1papa ∣ψa⟩⟨ψa∣=∑a=0N−1pa∣ψa⟩⟨ψa∣=ρ.\displaystyle \begin{aligned}\mathrm{Tr}_{\mathsf{YZ}}\bigl(\lvert\gamma_1\rangle\langle\gamma_1\rvert\bigr)&=\sum_{a,b=0}^{N-1}\sqrt{p_a p_b}\,\lvert\psi_a\rangle\langle\psi_b\rvert\;\langle0\vert0\rangle\;\langle b\vert a\rangle\\[4pt]&=\sum_{a,b=0}^{N-1}\sqrt{p_a p_b}\,\lvert\psi_a\rangle\langle\psi_b\rvert\;\delta_{ab}\\[4pt]&=\sum_{a=0}^{N-1}\sqrt{p_a p_a}\,\lvert\psi_a\rangle\langle\psi_a\rvert\\[4pt]&=\sum_{a=0}^{N-1}p_a\lvert\psi_a\rangle\langle\psi_a\rvert\\[4pt]&=\rho.\end{aligned}

What survives is the ensemble decomposition we started from, which is just ρ\rho written out. So ∣γ1⟩\lvert\gamma_1\rangle is a purification of ρ\rho: it is a pure state of the larger system (X,Y,Z)(\mathsf{X},\mathsf{Y},\mathsf{Z}) whose reduced state on X\mathsf{X} is ρ\rho.

The role of Z\mathsf{Z} is therefore very concrete: it stores a coherent record of which member of the ensemble is associated with each term, while the overall state seen from X\mathsf{X} remains the same ρ\rho.

2. Connect the record to the given purification

We now have a purification ∣γ1⟩\lvert\gamma_1\rangle that contains the ensemble we want, and we already had the given purification ∣ϕ⟩\lvert\phi\rangle of ρ\rho. What remains is to get from one to the other without touching X\mathsf{X}.

Start with ∣ϕ⟩\lvert\phi\rangle and append the same auxiliary system Z\mathsf{Z}, but leave its record blank by putting it in the fixed state ∣0⟩\lvert0\rangle:

∣γ0⟩=∣ϕ⟩XY⊗∣0⟩Z.\displaystyle \lvert\gamma_0\rangle=\lvert\phi\rangle_{\mathsf{XY}}\otimes\lvert0\rangle_{\mathsf{Z}}.

This is another purification of ρ\rho. Nothing has changed on X\mathsf{X}, and Z\mathsf{Z} is only an extra system sitting in a fixed state, so tracing both away gives back what ∣ϕ⟩\lvert\phi\rangle gave:

TrYZ(∣γ0⟩⟨γ0∣)=TrY(∣ϕ⟩⟨ϕ∣)=ρ.\displaystyle \begin{aligned}\mathrm{Tr}_{\mathsf{YZ}}\bigl(\lvert\gamma_0\rangle\langle\gamma_0\rvert\bigr)&=\mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\phi\rangle\langle\phi\rvert\bigr)\\[4pt]&=\rho.\end{aligned}

So there are now two purifications of exactly the same state of X\mathsf{X}:

∣γ0⟩=∣ϕ⟩XY⊗∣0⟩Z\displaystyle \lvert\gamma_0\rangle=\lvert\phi\rangle_{\mathsf{XY}}\otimes\lvert0\rangle_{\mathsf{Z}}
∣γ1⟩=∑a=0N−1pa ∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z\displaystyle \lvert\gamma_1\rangle=\sum_{a=0}^{N-1}\sqrt{p_a}\,\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}}

In both, X\mathsf{X} has the same reduced state ρ\rho. Only the purifying systems (Y,Z)(\mathsf{Y},\mathsf{Z}) are arranged differently. The now says that there is a unitary UU acting only on (Y,Z)(\mathsf{Y},\mathsf{Z}) that turns one into the other:

(IX⊗U)∣γ0⟩=∣γ1⟩.\displaystyle (I_{\mathsf{X}}\otimes U)\lvert\gamma_0\rangle=\lvert\gamma_1\rangle.

This is the key step. Starting from a blank record, a unitary on (Y,Z)(\mathsf{Y},\mathsf{Z}) builds exactly the correlations needed to write the ensemble down, and it leaves X\mathsf{X} untouched. Afterwards the state of Z\mathsf{Z} says which ∣ψa⟩\lvert\psi_a\rangle goes with which term.

A different ensemble of the same ρ\rho gives a different ∣γ1⟩\lvert\gamma_1\rangle, and so a different UU. This is how the different decompositions of ρ\rho will turn into different measurements on Y\mathsf{Y}.

3. Read the record

We now have everything needed to read the ensemble stored in Z\mathsf{Z}: start with ∣ϕ⟩\lvert\phi\rangle, append Z\mathsf{Z} in ∣0⟩\lvert0\rangle to reach ∣γ0⟩\lvert\gamma_0\rangle, and apply UU to (Y,Z)(\mathsf{Y},\mathsf{Z}) to reach ∣γ1⟩\lvert\gamma_1\rangle.

∣ϕ⟩\lvert\phi\rangleZYX∣γ0⟩\lvert\gamma_0\rangle∣γ1⟩\lvert\gamma_1\rangleaa (with probability pap_a)

After the unitary, the state is

∣γ1⟩=∑a=0N−1pa ∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z.\displaystyle \lvert\gamma_1\rangle=\sum_{a=0}^{N-1}\sqrt{p_a}\,\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}}.

The states ∣a⟩Z\lvert a\rangle_{\mathsf{Z}} are the labels stored for the ensemble, so measuring Z\mathsf{Z} in its own basis asks a single question: which label is present?

Suppose the measurement gives outcome aa. Projecting onto ∣a⟩Z\lvert a\rangle_{\mathsf{Z}} keeps only the term carrying that label:

(IX⊗IY⊗∣a⟩⟨a∣)∣γ1⟩=pa  ∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z.\displaystyle \bigl(I_{\mathsf{X}}\otimes I_{\mathsf{Y}}\otimes\lvert a\rangle\langle a\rvert\bigr)\lvert\gamma_1\rangle=\sqrt{p_a}\;\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}}.

The surviving vector has squared norm ∥pa  ∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z∥2=pa\bigl\lVert\sqrt{p_a}\;\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}}\bigr\rVert^2=p_a, so the outcome aa occurs with probability pap_a.

The post-measurement state is obtained by dividing by its norm:

1pa(pa  ∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z)=∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z.\displaystyle \begin{aligned}&\frac{1}{\sqrt{p_a}}\bigl(\sqrt{p_a}\;\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}}\bigr)\\[4pt]&=\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}}.\end{aligned}

Once the outcome aa is known, the post-measurement state is

∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z.\displaystyle \lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}}.

This is a product state of the composite system (X,Y,Z)(\mathsf{X},\mathsf{Y},\mathsf{Z}): it factors into a state of X\mathsf{X}, a state of Y\mathsf{Y}, and a state of Z\mathsf{Z}. In particular, X\mathsf{X} is not entangled with Y\mathsf{Y} or Z\mathsf{Z}, and its state is exactly ∣ψa⟩\lvert\psi_a\rangle.

System Y\mathsf{Y} is always in the same state ∣0⟩Y\lvert0\rangle_{\mathsf{Y}}, regardless of the measurement outcome. It therefore carries no information about aa and can be discarded.

Thus, measuring Z\mathsf{Z} produces exactly the desired ensemble: outcome aa occurs with probability pap_a and leaves X\mathsf{X} in the state ∣ψa⟩\lvert\psi_a\rangle.

There is one problem, though. The theorem asks for a measurement of Y\mathsf{Y}, not of the newly introduced system Z\mathsf{Z}. The final step is to show that this whole detour through Z\mathsf{Z} can be rewritten as a measurement on Y\mathsf{Y} alone.

4. Turn the construction into a measurement of Y\mathsf{Y}

The construction so far gives everything the theorem requires except one point: we measure Z\mathsf{Z}, but the theorem asks for a measurement on Y\mathsf{Y}. We therefore need to express the same procedure as a measurement on Y\mathsf{Y}, with Z\mathsf{Z} removed from the final description.

Z\mathsf{Z} is only an auxiliary system. It starts in the fixed state ∣0⟩Z\lvert0\rangle_{\mathsf{Z}}, interacts with Y\mathsf{Y} through the unitary UU, and is then measured. The only information we keep from Z\mathsf{Z} is the measurement outcome aa. The three steps can therefore be combined into a single operation on Y\mathsf{Y}.

Recall where UU came from. The states ∣γ0⟩=∣ϕ⟩XY⊗∣0⟩Z\lvert\gamma_0\rangle=\lvert\phi\rangle_{\mathsf{XY}}\otimes\lvert0\rangle_{\mathsf{Z}} and ∣γ1⟩=∑apa ∣ψa⟩X⊗∣0⟩Y⊗∣a⟩Z\lvert\gamma_1\rangle=\sum_a\sqrt{p_a}\,\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}} are purifications of the same state on X\mathsf{X}. By the , there is a unitary UU acting only on (Y,Z)(\mathsf{Y},\mathsf{Z}) such that (IX⊗U)∣γ0⟩=∣γ1⟩(I_{\mathsf{X}}\otimes U)\lvert\gamma_0\rangle=\lvert\gamma_1\rangle. We now use this same UU to describe the corresponding measurement on Y\mathsf{Y}.

Take an arbitrary normalised state ∣v⟩Y\lvert v\rangle_{\mathsf{Y}} and temporarily leave X\mathsf{X} out. Appending the auxiliary system Z\mathsf{Z} in its fixed state gives ∣v⟩Y⊗∣0⟩Z\lvert v\rangle_{\mathsf{Y}}\otimes\lvert0\rangle_{\mathsf{Z}}, and applying UU gives U(∣v⟩Y⊗∣0⟩Z)U\bigl(\lvert v\rangle_{\mathsf{Y}}\otimes\lvert0\rangle_{\mathsf{Z}}\bigr).

Because {∣a⟩Z}\{\lvert a\rangle_{\mathsf{Z}}\} is the basis in which Z\mathsf{Z} is measured, this state can be expanded in that basis: U(∣v⟩Y⊗∣0⟩Z)=∑a∣va⟩Y⊗∣a⟩ZU\bigl(\lvert v\rangle_{\mathsf{Y}}\otimes\lvert0\rangle_{\mathsf{Z}}\bigr)=\sum_a\lvert v_a\rangle_{\mathsf{Y}}\otimes\lvert a\rangle_{\mathsf{Z}}. Here ∣va⟩Y\lvert v_a\rangle_{\mathsf{Y}} is simply the vector of Y\mathsf{Y} that accompanies the basis state ∣a⟩Z\lvert a\rangle_{\mathsf{Z}}. No additional assumption is being made. This is just the expansion of a joint state in the basis of Z\mathsf{Z}, and the vectors ∣va⟩\lvert v_a\rangle need not be normalised.

Now measure Z\mathsf{Z}. For outcome aa, the probability is ∥∣va⟩∥2\bigl\lVert\lvert v_a\rangle\bigr\rVert^2, and the corresponding normalised state of Y\mathsf{Y} is ∣va⟩ / ∥∣va⟩∥\lvert v_a\rangle\,\big/\,\bigl\lVert\lvert v_a\rangle\bigr\rVert. Thus ∣va⟩\lvert v_a\rangle contains both pieces of information associated with outcome aa: its squared norm gives the probability, while its normalised direction gives the resulting state of Y\mathsf{Y}.

The vector ∣va⟩\lvert v_a\rangle depends linearly on the input ∣v⟩\lvert v\rangle, because every step used to obtain it is linear. We can therefore describe this dependence by a linear operator on Y\mathsf{Y}. To extract the aa-labelled component, apply ⟨a∣\langle a\rvert to the Z\mathsf{Z} system:

(IY⊗⟨a∣) U (∣v⟩Y⊗∣0⟩Z)=∣va⟩Y.\displaystyle (I_{\mathsf{Y}}\otimes\langle a\rvert)\,U\,\bigl(\lvert v\rangle_{\mathsf{Y}}\otimes\lvert0\rangle_{\mathsf{Z}}\bigr)=\lvert v_a\rangle_{\mathsf{Y}}.

The operation that appends ∣0⟩Z\lvert0\rangle_{\mathsf{Z}} can be written as IY⊗∣0⟩I_{\mathsf{Y}}\otimes\lvert0\rangle. Here ∣0⟩\lvert0\rangle is read as an operator rather than as a state: it takes a number cc to the vector c ∣0⟩Zc\,\lvert0\rangle_{\mathsf{Z}}, so it maps the one-dimensional space of numbers into Z\mathsf{Z}. Tensoring it with IYI_{\mathsf{Y}} gives an operator that takes a state of Y\mathsf{Y} to a state of (Y,Z)(\mathsf{Y},\mathsf{Z}): IYI_{\mathsf{Y}} carries ∣v⟩\lvert v\rangle through unchanged, and ∣0⟩\lvert0\rangle supplies the new factor, so (IY⊗∣0⟩)∣v⟩Y=∣v⟩Y⊗∣0⟩Z(I_{\mathsf{Y}}\otimes\lvert0\rangle)\lvert v\rangle_{\mathsf{Y}}=\lvert v\rangle_{\mathsf{Y}}\otimes\lvert0\rangle_{\mathsf{Z}}. Therefore

(IY⊗⟨a∣) U (IY⊗∣0⟩) ∣v⟩Y=∣va⟩Y.\displaystyle (I_{\mathsf{Y}}\otimes\langle a\rvert)\,U\,(I_{\mathsf{Y}}\otimes\lvert0\rangle)\,\lvert v\rangle_{\mathsf{Y}}=\lvert v_a\rangle_{\mathsf{Y}}.

Since this holds for every input ∣v⟩\lvert v\rangle, define

Ma=(IY⊗⟨a∣) U (IY⊗∣0⟩).\displaystyle M_a=(I_{\mathsf{Y}}\otimes\langle a\rvert)\,U\,(I_{\mathsf{Y}}\otimes\lvert0\rangle).

The three factors correspond exactly to the three steps of the original construction:

  • IY⊗∣0⟩I_{\mathsf{Y}}\otimes\lvert0\rangle: append Z\mathsf{Z} in the fixed state ∣0⟩\lvert0\rangle.
  • UU: apply the unitary to (Y,Z)(\mathsf{Y},\mathsf{Z}).
  • IY⊗⟨a∣I_{\mathsf{Y}}\otimes\langle a\rvert: extract the component corresponding to outcome aa.

The auxiliary system Z\mathsf{Z} has now disappeared from the input and output of MaM_a. It remains only inside the definition of the operator.

For an input ∣v⟩\lvert v\rangle, the probability of outcome aa is therefore ∥Ma∣v⟩∥2=⟨v∣Ma†Ma∣v⟩\bigl\lVert M_a\lvert v\rangle\bigr\rVert^2=\langle v\rvert M_a^{\dagger}M_a\lvert v\rangle, so define Pa=Ma†MaP_a=M_a^{\dagger}M_a. Each PaP_a is positive semidefinite because it has the form Ma†MaM_a^{\dagger}M_a. It remains to check that the probabilities sum to one, which is equivalent to ∑aPa=IY\sum_a P_a=I_{\mathsf{Y}}. Using the definition of MaM_a:

∑aPa=∑a(IY⊗⟨0∣) U† (IY⊗∣a⟩⟨a∣) U (IY⊗∣0⟩)=(IY⊗⟨0∣) U†(IY⊗∑a∣a⟩⟨a∣)U (IY⊗∣0⟩)only the middle factor depends on a=(IY⊗⟨0∣) U† (IY⊗IZ) U (IY⊗∣0⟩)completeness: ∑a∣a⟩⟨a∣=IZ=(IY⊗⟨0∣) U† IYZ U (IY⊗∣0⟩)IY⊗IZ is the identity on (Y,Z)=(IY⊗⟨0∣) U†U (IY⊗∣0⟩)multiplying by the identity changes nothing=(IY⊗⟨0∣) IYZ (IY⊗∣0⟩)U is unitary: U†U=IYZ=(IY⊗⟨0∣)(IY⊗∣0⟩)again the identity changes nothing=IY⊗⟨0∣0⟩each system multiplies on its own: IYIY=IY on Y, ⟨0∣∣0⟩ on Z=IY∣0⟩Z is a unit vector, so ⟨0∣0⟩=1 and IY⊗1=IY\displaystyle \begin{aligned} \sum_a P_a &=\sum_a(I_{\mathsf{Y}}\otimes\langle0\rvert)\,U^{\dagger}\,(I_{\mathsf{Y}}\otimes\lvert a\rangle\langle a\rvert)\,U\,(I_{\mathsf{Y}}\otimes\lvert0\rangle)\\[4pt] &=(I_{\mathsf{Y}}\otimes\langle0\rvert)\,U^{\dagger}\Bigl(I_{\mathsf{Y}}\otimes\sum_a\lvert a\rangle\langle a\rvert\Bigr)U\,(I_{\mathsf{Y}}\otimes\lvert0\rangle) &&\quad\textcolor{#94a3b8}{\text{only the middle factor depends on }a}\\[4pt] &=(I_{\mathsf{Y}}\otimes\langle0\rvert)\,U^{\dagger}\,(I_{\mathsf{Y}}\otimes I_{\mathsf{Z}})\,U\,(I_{\mathsf{Y}}\otimes\lvert0\rangle) &&\quad\textcolor{#94a3b8}{\text{completeness: }\sum_a\lvert a\rangle\langle a\rvert=I_{\mathsf{Z}}}\\[4pt] &=(I_{\mathsf{Y}}\otimes\langle0\rvert)\,U^{\dagger}\,I_{\mathsf{YZ}}\,U\,(I_{\mathsf{Y}}\otimes\lvert0\rangle) &&\quad\textcolor{#94a3b8}{I_{\mathsf{Y}}\otimes I_{\mathsf{Z}}\text{ is the identity on }(\mathsf{Y},\mathsf{Z})}\\[4pt] &=(I_{\mathsf{Y}}\otimes\langle0\rvert)\,U^{\dagger}U\,(I_{\mathsf{Y}}\otimes\lvert0\rangle) &&\quad\textcolor{#94a3b8}{\text{multiplying by the identity changes nothing}}\\[4pt] &=(I_{\mathsf{Y}}\otimes\langle0\rvert)\,I_{\mathsf{YZ}}\,(I_{\mathsf{Y}}\otimes\lvert0\rangle) &&\quad\textcolor{#94a3b8}{U\text{ is unitary: }U^{\dagger}U=I_{\mathsf{YZ}}}\\[4pt] &=(I_{\mathsf{Y}}\otimes\langle0\rvert)(I_{\mathsf{Y}}\otimes\lvert0\rangle) &&\quad\textcolor{#94a3b8}{\text{again the identity changes nothing}}\\[4pt] &=I_{\mathsf{Y}}\otimes\langle0\vert0\rangle &&\quad\textcolor{#94a3b8}{\text{each system multiplies on its own: }I_{\mathsf{Y}}I_{\mathsf{Y}}=I_{\mathsf{Y}}\text{ on }\mathsf{Y},\ \langle0\rvert\lvert0\rangle\text{ on }\mathsf{Z}}\\[4pt] &=I_{\mathsf{Y}} &&\quad\textcolor{#94a3b8}{\lvert0\rangle_{\mathsf{Z}}\text{ is a unit vector, so }\langle0\vert0\rangle=1\text{ and }I_{\mathsf{Y}}\otimes1=I_{\mathsf{Y}}} \end{aligned}

Thus {Pa}\{P_a\} is a valid measurement on Y\mathsf{Y}.

It remains to check that this measurement produces the desired ensemble when applied to the original purification ∣ϕ⟩XY\lvert\phi\rangle_{\mathsf{XY}}. Apply MaM_a to its Y\mathsf{Y} part, with X\mathsf{X} carried along untouched:

(IX⊗Ma)∣ϕ⟩XY=(IX⊗(IY⊗⟨a∣) U (IY⊗∣0⟩))∣ϕ⟩XYthe definition of Ma=(IX⊗IY⊗⟨a∣) (IX⊗U) (IX⊗IY⊗∣0⟩)∣ϕ⟩XYsplitting the product, each factor with its own IX=(IX⊗IY⊗⟨a∣) (IX⊗U) (∣ϕ⟩XY⊗∣0⟩Z)the identities carry ∣ϕ⟩XY through, ∣0⟩ supplies the Z factor=(IX⊗IY⊗⟨a∣) (IX⊗U)∣γ0⟩∣γ0⟩=∣ϕ⟩XY⊗∣0⟩Z=(IX⊗IY⊗⟨a∣)∣γ1⟩(IX⊗U)∣γ0⟩=∣γ1⟩=(IX⊗IY⊗⟨a∣)∑bpb ∣ψb⟩X⊗∣0⟩Y⊗∣b⟩Z∣γ1⟩ written out=∑bpb ∣ψb⟩X⊗∣0⟩Y ⟨a∣b⟩⟨a∣ meets the Z label of every term=pa ∣ψa⟩X⊗∣0⟩Y⟨a∣b⟩=δab, so only b=a survives\displaystyle \begin{aligned} (I_{\mathsf{X}}\otimes M_a)\lvert\phi\rangle_{\mathsf{XY}} &=\bigl(I_{\mathsf{X}}\otimes(I_{\mathsf{Y}}\otimes\langle a\rvert)\,U\,(I_{\mathsf{Y}}\otimes\lvert0\rangle)\bigr)\lvert\phi\rangle_{\mathsf{XY}} &&\quad\textcolor{#94a3b8}{\text{the definition of }M_a}\\[4pt] &=(I_{\mathsf{X}}\otimes I_{\mathsf{Y}}\otimes\langle a\rvert)\,(I_{\mathsf{X}}\otimes U)\,(I_{\mathsf{X}}\otimes I_{\mathsf{Y}}\otimes\lvert0\rangle)\lvert\phi\rangle_{\mathsf{XY}} &&\quad\textcolor{#94a3b8}{\text{splitting the product, each factor with its own }I_{\mathsf{X}}}\\[4pt] &=(I_{\mathsf{X}}\otimes I_{\mathsf{Y}}\otimes\langle a\rvert)\,(I_{\mathsf{X}}\otimes U)\,\bigl(\lvert\phi\rangle_{\mathsf{XY}}\otimes\lvert0\rangle_{\mathsf{Z}}\bigr) &&\quad\textcolor{#94a3b8}{\text{the identities carry }\lvert\phi\rangle_{\mathsf{XY}}\text{ through, }\lvert0\rangle\text{ supplies the }\mathsf{Z}\text{ factor}}\\[4pt] &=(I_{\mathsf{X}}\otimes I_{\mathsf{Y}}\otimes\langle a\rvert)\,(I_{\mathsf{X}}\otimes U)\lvert\gamma_0\rangle &&\quad\textcolor{#94a3b8}{\lvert\gamma_0\rangle=\lvert\phi\rangle_{\mathsf{XY}}\otimes\lvert0\rangle_{\mathsf{Z}}}\\[4pt] &=(I_{\mathsf{X}}\otimes I_{\mathsf{Y}}\otimes\langle a\rvert)\lvert\gamma_1\rangle &&\quad\textcolor{#94a3b8}{(I_{\mathsf{X}}\otimes U)\lvert\gamma_0\rangle=\lvert\gamma_1\rangle}\\[4pt] &=(I_{\mathsf{X}}\otimes I_{\mathsf{Y}}\otimes\langle a\rvert)\sum_b\sqrt{p_b}\,\lvert\psi_b\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\otimes\lvert b\rangle_{\mathsf{Z}} &&\quad\textcolor{#94a3b8}{\lvert\gamma_1\rangle\text{ written out}}\\[4pt] &=\sum_b\sqrt{p_b}\,\lvert\psi_b\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}\,\langle a\vert b\rangle &&\quad\textcolor{#94a3b8}{\langle a\rvert\text{ meets the }\mathsf{Z}\text{ label of every term}}\\[4pt] &=\sqrt{p_a}\,\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}} &&\quad\textcolor{#94a3b8}{\langle a\vert b\rangle=\delta_{ab}\text{, so only }b=a\text{ survives}} \end{aligned}

The result (IX⊗Ma)∣ϕ⟩XY=pa ∣ψa⟩X⊗∣0⟩Y(I_{\mathsf{X}}\otimes M_a)\lvert\phi\rangle_{\mathsf{XY}}=\sqrt{p_a}\,\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}} describes the state of X\mathsf{X} and Y\mathsf{Y} when outcome aa occurs. Its squared norm is pap_a, so outcome aa occurs with probability pap_a. After normalisation, the state is ∣ψa⟩X⊗∣0⟩Y\lvert\psi_a\rangle_{\mathsf{X}}\otimes\lvert0\rangle_{\mathsf{Y}}. Thus, outcome aa leaves X\mathsf{X} in the pure state ∣ψa⟩\lvert\psi_a\rangle, while Y\mathsf{Y} is left in the fixed state ∣0⟩\lvert0\rangle and is no longer entangled with X\mathsf{X}.

Therefore, measuring Y\mathsf{Y} with the POVM {Pa}\{P_a\} produces outcome aa with probability pap_a and prepares X\mathsf{X} in the corresponding state ∣ψa⟩\lvert\psi_a\rangle. The auxiliary system Z\mathsf{Z} was introduced as a construction tool: by adding Z\mathsf{Z}, applying the unitary UU, and measuring Z\mathsf{Z}, the proof derives the measurement operators MaM_a and hence the POVM elements PaP_a acting directly on Y\mathsf{Y}.

The theorem guarantees the existence of the required unitary UU, but it does not by itself provide a physical circuit for implementing UU. Constructing such a circuit is a separate problem. The HJW theorem merely establishes the mathematical fact that the desired measurement on Y\mathsf{Y} exists — what a bummer! DM me if you have read this far and we could rant about this together.

In this way, every decomposition of a density matrix into pure states can be realised by measuring a purifying system. Different decompositions of the same mixed state correspond to different measurements on that purifying system.

Fidelity

The fidelity of two quantum states measures their similarity (overlap). It ranges from 00 to 11, with 11 for identical states and 00 for perfectly distinguishable states. For two states given by density matrices ρ\rho and σ\sigma it is defined as:

F(ρ,σ)=Trρ σρ.\displaystyle F(\rho,\sigma)=\mathrm{Tr}\sqrt{\sqrt{\rho}\,\sigma\sqrt{\rho}}.

Two distinct matrix square roots appear here. Both square roots must exist for the formula to be well defined:

  • ρ\sqrt{\rho}: ρ\rho is positive semidefinite by definition, so this square root exists. It is obtained from the of ρ\rho by taking the square root of each eigenvalue.
  • ρ σρ\sqrt{\sqrt{\rho}\,\sigma\sqrt{\rho}}: since σ\sigma is positive semidefinite as well, it splits as σ=σσ\sigma=\sqrt{\sigma}\sqrt{\sigma}. The adjoint reverses a product, and ρ\sqrt{\rho} and σ\sqrt{\sigma} are Hermitian, so (σρ)†=ρσ(\sqrt{\sigma}\sqrt{\rho})^{\dagger}=\sqrt{\rho}\sqrt{\sigma}. Writing M=σρM=\sqrt{\sigma}\sqrt{\rho}, the matrix under the root is ρ σρ=ρσσρ=M†M\sqrt{\rho}\,\sigma\sqrt{\rho}=\sqrt{\rho}\sqrt{\sigma}\sqrt{\sigma}\sqrt{\rho}=M^{\dagger}M. Any matrix of the form M†MM^{\dagger}M is positive semidefinite, so this square root exists as well.

So the trace in the fidelity formula is a well-defined nonnegative number.

Being positive semidefinite, ρ σρ\sqrt{\rho}\,\sigma\sqrt{\rho} has a spectral decomposition with nonnegative eigenvalues λk\lambda_k, and its square root is taken eigenvalue by eigenvalue:

ρ σρ=∑k=0n−1λk∣ϕk⟩⟨ϕk∣⟹ρ σρ=∑k=0n−1λk ∣ϕk⟩⟨ϕk∣.\displaystyle \sqrt{\rho}\,\sigma\sqrt{\rho}=\sum_{k=0}^{n-1}\lambda_k\lvert\phi_k\rangle\langle\phi_k\rvert \quad\Longrightarrow\quad \sqrt{\sqrt{\rho}\,\sigma\sqrt{\rho}}=\sum_{k=0}^{n-1}\sqrt{\lambda_k}\,\lvert\phi_k\rangle\langle\phi_k\rvert.

The trace of a spectral decomposition is the sum of its eigenvalues, so the fidelity is the sum of the square roots of the λk\lambda_k. This gives one useful form of the fidelity. Several equivalent formulas are useful, and they are shown side by side:

  • F(ρ,σ)=∑k=0n−1λk\displaystyle F(\rho,\sigma)=\sum_{k=0}^{n-1}\sqrt{\lambda_k}

    the sum of the square roots of the eigenvalues of ρ σρ\sqrt{\rho}\,\sigma\sqrt{\rho}. This is the direct computational form: diagonalising the matrix reduces the fidelity to a sum of numbers.

  • F(ρ,σ)=∥ρσ∥1=∥σρ∥1\displaystyle F(\rho,\sigma)=\bigl\lVert\sqrt{\rho}\sqrt{\sigma}\bigr\rVert_1=\bigl\lVert\sqrt{\sigma}\sqrt{\rho}\bigr\rVert_1

    a trace-norm form that makes the symmetry between ρ\rho and σ\sigma explicit. It follows from the definition of the trace norm, ∥M∥1=TrM†M\lVert M\rVert_1=\mathrm{Tr}\sqrt{M^{\dagger}M}, by taking M=σρM=\sqrt{\sigma}\sqrt{\rho}.

  • F(ρ,σ)=max⁡U unitary∣Tr(ρσ U)∣\displaystyle F(\rho,\sigma)=\max_{U\ \text{unitary}}\bigl\lvert\mathrm{Tr}\bigl(\sqrt{\rho}\sqrt{\sigma}\,U\bigr)\bigr\rvert

    a variational form: fidelity is the largest possible overlap after applying a unitary. It follows from the variational characterization ∥M∥1=max⁡U∣Tr(MU)∣\lVert M\rVert_1=\max_U\lvert\mathrm{Tr}(MU)\rvert.

There are simpler formulas when at least one of the states is pure:

  • F(∣ϕ⟩⟨ϕ∣,∣ψ⟩⟨ψ∣)=∣⟨ϕ∣ψ⟩∣\displaystyle F\bigl(\lvert\phi\rangle\langle\phi\rvert,\lvert\psi\rangle\langle\psi\rvert\bigr)=\lvert\langle\phi\vert\psi\rangle\rvert

    for two pure states, the absolute value of their inner product.

  • F(∣ϕ⟩⟨ϕ∣,σ)=⟨ϕ∣σ∣ϕ⟩\displaystyle F\bigl(\lvert\phi\rangle\langle\phi\rvert,\sigma\bigr)=\sqrt{\langle\phi\rvert\sigma\lvert\phi\rangle}

    for one pure state and one arbitrary state, the square root of the probability of obtaining ∣ϕ⟩\lvert\phi\rangle in a measurement containing it as an outcome.

Example: fidelity between two states of a single qubit

Two states of a single qubit appear as vectors in the . Their fidelity depends only on their radii rρr_\rho, rσr_\sigma and the angle θ\theta between them. The radii measure mixedness: 00 is completely mixed and 11 is pure. A common rotation changes nothing, so ρ\rho is fixed upright and only their relative angle varies.

|0⟩|1⟩|+⟩|−⟩
θ\theta
ρ\rho
σ\sigma
F=1  +  0.14⏟alignment  +  0.40⏟shared mixedness2  ≈  0.88F=\sqrt{\frac{1\;+\;\underbrace{0.14}_{\text{alignment}}\;+\;\underbrace{0.40}_{\text{shared mixedness}}}{2}}\;\approx\;\textcolor{#0f766e}{0.88}
Fidelity0.88
0 · distinguishableidentical · 1
75°
0.85
0.65

Properties of fidelity

The following properties capture the key ways fidelity behaves. Some follow directly from the , others require a proof.

  1. For any two density matrices ρ\rho and σ\sigma we have 0≤F(ρ,σ)≤10\le F(\rho,\sigma)\le 1.

    • F(ρ,σ)=0F(\rho,\sigma)=0 if and only if ρ\rho and σ\sigma have orthogonal images, which is the case in which one measurement tells them apart every time.
    • F(ρ,σ)=1F(\rho,\sigma)=1 if and only if ρ=σ\rho=\sigma.
  2. The fidelity is symmetric: F(ρ,σ)=F(σ,ρ)F(\rho,\sigma)=F(\sigma,\rho).
  3. The fidelity is multiplicative for product states:

    F(ρ1⊗⋯⊗ρm, σ1⊗⋯⊗σm)=F(ρ1,σ1)⋯F(ρm,σm).\displaystyle F(\rho_1\otimes\cdots\otimes\rho_m,\ \sigma_1\otimes\cdots\otimes\sigma_m)=F(\rho_1,\sigma_1)\cdots F(\rho_m,\sigma_m).

    Thus, even states with fidelity close to 11 become nearly perfectly distinguishable when many copies are compared: a fidelity below 11 raised to the mmth power approaches 00 as mm grows.

  4. Fidelity cannot decrease under a Φ\Phi:

    F(ρ,σ)≤F(Φ(ρ),Φ(σ)).\displaystyle F(\rho,\sigma)\le F\bigl(\Phi(\rho),\Phi(\sigma)\bigr).

    No physical process can make two states more distinguishable. Applying the same channel to both states can only increase their fidelity, including operations such as discarding part of the system. Thus, once two states are close, further processing cannot make them more distinguishable.

  5. Fidelity and trace distance are closely related:

    1−12∥ρ−σ∥1 ≤ F(ρ,σ) ≤ 1−14∥ρ−σ∥1 2.\displaystyle 1-\tfrac12\lVert\rho-\sigma\rVert_1\ \le\ F(\rho,\sigma)\ \le\ \sqrt{1-\tfrac14\lVert\rho-\sigma\rVert_1^{\,2}}.

    The trace distance x=12∥ρ−σ∥1x=\tfrac12\lVert\rho-\sigma\rVert_1 quantifies how well the two states can be . Fidelity and trace distance do not determine each other exactly, but each bounds the other.

    0101
    F(ρ,σ)F(\rho,\sigma)
    xx
    1−x2\sqrt{1-x^2}
    1−x1-x
    0.40
    0.40
    x=0.40⟹0.60  ≤  F(ρ,σ)  ≤  0.92x=0.40\qquad\Longrightarrow\qquad 0.60\;\le\;F(\rho,\sigma)\;\le\;0.92

Uhlmann’s theorem

A mixed state has no single state vector, so its similarity to another mixed state cannot be read from an ordinary inner product. It can, however, be represented by a pure state on a larger system through a . Uhlmann’s theorem turns this freedom into a geometric interpretation of : choose the purifications of the two states that align as closely as possible, then take the absolute value of their inner product.

The fidelity of two states is the largest overlap that any two purifications of them can have.

More precisely, let ρ\rho and σ\sigma be states of a system X\mathsf{X}, and let Y\mathsf{Y} be an auxiliary system with dimension at least that of X\mathsf{X}. Then

F(ρ,σ)=max⁡{∣⟨ϕ∣ψ⟩∣ : TrY(∣ϕ⟩⟨ϕ∣)=ρ, TrY(∣ψ⟩⟨ψ∣)=σ}.\displaystyle F(\rho,\sigma)=\max\Bigl\{\lvert\langle\phi\vert\psi\rangle\rvert\ :\ \mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\phi\rangle\langle\phi\rvert\bigr)=\rho,\ \mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\psi\rangle\langle\psi\rvert\bigr)=\sigma\Bigr\}.

Here ∣ϕ⟩\lvert\phi\rangle and ∣ψ⟩\lvert\psi\rangle are pure states of X⊗Y\mathsf{X}\otimes\mathsf{Y}. The two partial-trace conditions say that they reduce to ρ\rho and σ\sigma on X\mathsf{X}, so the maximum runs over all pairs of purifications of ρ\rho and σ\sigma on the same larger system. The absolute value makes the comparison independent of a global phase. The theorem says both that no pair has a larger overlap and that an optimal pair exists.

To prove the claim, start by choosing of the two density matrices:

ρ=∑a=0n−1pa ∣ua⟩⟨ua∣\displaystyle \rho=\sum_{a=0}^{n-1}p_a\,\lvert u_a\rangle\langle u_a\rvert
σ=∑b=0n−1qb ∣vb⟩⟨vb∣.\displaystyle \sigma=\sum_{b=0}^{n-1}q_b\,\lvert v_b\rangle\langle v_b\rvert.

The vectors ∣ua⟩\lvert u_a\rangle and ∣vb⟩\lvert v_b\rangle are orthonormal eigenvectors, while pap_a and qbq_b are their nonnegative eigenvalues. Because each density matrix has trace 11, both sets of eigenvalues sum to 11.

The next goal is to construct pure states on X⊗Y\mathsf{X}\otimes\mathsf{Y} whose reduced states on X\mathsf{X} are ρ\rho and σ\sigma. The eigenvalues are probabilities, whereas the coefficients of a pure state are amplitudes. We therefore use pa\sqrt{p_a} and qb\sqrt{q_b} as coefficients, so forming the density matrix recovers the required weights pap_a and qbq_b.

For each eigenvector on X\mathsf{X}, introduce an orthonormal record with the same index in the new auxiliary system Y\mathsf{Y}. Thus ∣ua‾⟩Y\lvert\overline{u_a}\rangle_{\mathsf{Y}} is a new vector in Y\mathsf{Y}, not a rotation or a second occurrence of ∣ua⟩X\lvert u_a\rangle_{\mathsf{X}} in the original state. It records which aa-term is present. The same construction is used for the bb-terms of σ\sigma:

∣ϕ0⟩=∑a=0n−1pa ∣ua⟩X⊗∣ua‾⟩Y\displaystyle \lvert\phi_0\rangle=\sum_{a=0}^{n-1}\sqrt{p_a}\,\lvert u_a\rangle_{\mathsf{X}}\otimes\lvert\overline{u_a}\rangle_{\mathsf{Y}}
∣ψ0⟩=∑b=0n−1qb ∣vb⟩X⊗∣vb‾⟩Y.\displaystyle \lvert\psi_0\rangle=\sum_{b=0}^{n-1}\sqrt{q_b}\,\lvert v_b\rangle_{\mathsf{X}}\otimes\lvert\overline{v_b}\rangle_{\mathsf{Y}}.

The subscript 00 is a label, not an index being summed over. It marks this particular pair, the purifications built directly from the eigenbases, and distinguishes them from the arbitrary purifications ∣ϕ⟩\lvert\phi\rangle and ∣ψ⟩\lvert\psi\rangle in the statement of the theorem. The whole point of what follows is that these two are only one choice among many.

Tracing out Y\mathsf{Y} now verifies the construction:

TrY(∣ϕ0⟩⟨ϕ0∣)=∑a,a′papa′ ∣ua⟩⟨ua′∣ ⟨ua′‾∣ua‾⟩expand the pure-state density matrix=∑a,a′papa′ ∣ua⟩⟨ua′∣ δa′athe records in Y are orthonormal=∑apa ∣ua⟩⟨ua∣=ρonly the terms with a′=a remain.\displaystyle \begin{aligned} \mathrm{Tr}_{\mathsf{Y}}\bigl(\lvert\phi_0\rangle\langle\phi_0\rvert\bigr) &=\sum_{a,a'}\sqrt{p_a p_{a'}}\,\lvert u_a\rangle\langle u_{a'}\rvert\, \langle\overline{u_{a'}}\vert\overline{u_a}\rangle &&\quad\textcolor{#94a3b8}{\text{expand the pure-state density matrix}}\\[4pt] &=\sum_{a,a'}\sqrt{p_a p_{a'}}\,\lvert u_a\rangle\langle u_{a'}\rvert\,\delta_{a'a} &&\quad\textcolor{#94a3b8}{\text{the records in }\mathsf{Y}\text{ are orthonormal}}\\[4pt] &=\sum_a p_a\,\lvert u_a\rangle\langle u_a\rvert=\rho &&\quad\textcolor{#94a3b8}{\text{only the terms with }a'=a\text{ remain}.} \end{aligned}

The same calculation with qbq_b and ∣vb⟩\lvert v_b\rangle gives TrY(∣ψ0⟩⟨ψ0∣)=σ\mathrm{Tr}_{\mathsf{Y}}(\lvert\psi_0\rangle\langle\psi_0\rvert)=\sigma. Hence both vectors are purifications of the states we started with.

These are convenient purifications, but they are not the only ones. By the , every purification of either state is obtained from its canonical purification by applying a unitary to the auxiliary system Y\mathsf{Y}. Thus a general pair, and therefore every pair included in the theorem’s maximum, has the following form, where UU and VV are arbitrary unitaries on Y\mathsf{Y}:

∣ϕ⟩=∑a=0n−1pa ∣ua⟩X⊗U∣ua‾⟩Y\displaystyle \lvert\phi\rangle=\sum_{a=0}^{n-1}\sqrt{p_a}\,\lvert u_a\rangle_{\mathsf{X}}\otimes U\lvert\overline{u_a}\rangle_{\mathsf{Y}}
∣ψ⟩=∑b=0n−1qb ∣vb⟩X⊗V∣vb‾⟩Y.\displaystyle \lvert\psi\rangle=\sum_{b=0}^{n-1}\sqrt{q_b}\,\lvert v_b\rangle_{\mathsf{X}}\otimes V\lvert\overline{v_b}\rangle_{\mathsf{Y}}.

Now expand the overlap of this pair. This will put the two states into the factors ρ\sqrt{\rho} and σ\sqrt{\sigma} and isolate the unitary choices:

Each pair of purifications is represented by a choice of UU and VV. Maximising their overlap gives:

This equality shows that the largest possible overlap between purifications of ρ\rho and σ\sigma is exactly their fidelity. The unitary QQ that gives the maximum in the trace-norm characterisation can be written as Q=VTU‾Q=V^{\mathsf{T}}\overline{U} for suitable unitaries UU and VV. These unitaries therefore define a pair of purifications whose overlap equals the fidelity. Thus the maximum is achieved by an actual pair of purifications, proving Uhlmann’s theorem.

Gentle measurement lemma

A measurement can change the state it acts on. The gentle measurement lemma says that one particular measurement update leaves the state close to its starting point when the observed outcome was already very likely. Here, closeness is measured by .

Let ρ\rho be a state of X\mathsf{X}, and let {P0,…,Pm−1}\{P_0,\ldots,P_{m-1}\} be a , so its effects are positive semidefinite and sum to II. Write p=Tr(P0ρ)p=\mathrm{Tr}(P_0\rho) for the probability of outcome 00, and suppose that outcome is highly likely, in the sense that for some 0<ε<10<\varepsilon<1:

p>1−ε.\displaystyle p>1-\varepsilon.

Use the with operators Pa\sqrt{P_a} supplied by . If outcome 00 appears, the system is left in the conditional state

ρ0=P0 ρP0p.\displaystyle \rho_0=\frac{\sqrt{P_0}\,\rho\sqrt{P_0}}{p}.

The squared fidelity between the original and conditional states is at least the probability of that outcome:

F(ρ,ρ0)2≥p>1−ε.\displaystyle F(\rho,\rho_0)^2\ge p>1-\varepsilon.

For example, take the qubit state ∣ψp⟩=p ∣0⟩+1−p ∣1⟩\lvert\psi_p\rangle=\sqrt p\,\lvert0\rangle+\sqrt{1-p}\,\lvert1\rangle and the measurement with P0=∣0⟩⟨0∣P_0=\lvert0\rangle\langle0\rvert. Outcome 00 arrives with probability pp and leaves ρ0=∣0⟩⟨0∣\rho_0=\lvert0\rangle\langle0\rvert behind.

|0⟩|1⟩
θ\theta
ρ0\rho_0
ρ\rho
F(ρ,ρ0)=∣⟨0∣ψp⟩∣=cos⁡θ=p≈0.95F(\rho,\rho_0)=\lvert\langle0\vert\psi_p\rangle\rvert=\cos\theta=\sqrt p\approx\textcolor{#0f766e}{0.95}
Angle θ\theta between the states18°
F(ρ,ρ0)F(\rho,\rho_0)0.95
F(ρ,ρ0)2F(\rho,\rho_0)^20.90
0.90

Proof

The idea is simple: rewrite the fidelity until it becomes a single trace involving the measurement operator P0P_0. At that point, the condition 0≤P0≤I0\le P_0\le I gives the bound.

The lemma captures a useful tradeoff between information and disturbance. When the outcome 00 is already almost certain, learning that outcome does not significantly change the state. The probability of an unexpected outcome controls how much disturbance can be introduced.

This is why the result is useful later: a state can be checked, decoded, or confirmed while remaining close to the state needed for whatever comes next. The disturbance is tied to the probability that the measurement reveals something unexpected, rather than to the mere act of obtaining information.