Quantum information basics

From classical bits to qubits: state vectors and Dirac notation, measurement, compound systems, quantum circuits, and entanglement, with teleportation, superdense coding and the CHSH game.

Classical information

Classical states

Suppose we have a physical system XX that stores information. In the classical model, the system has a finite set of possible classical states. Call that set Σ\Sigma.

At any given moment, XX is in exactly one of those states. If the current state is written as x(t)x(t), then x(t)∈Σx(t)\in\Sigma.

Bit

Σ={0,1}\Sigma=\{0,1\}

The system is either 0 or 1, never both.

Dice

Σ={1,2,3,4,5,6}\Sigma=\{1,2,3,4,5,6\}

The upward face can be only one value at a time.

Probabilistic states

When we are uncertain about which state XX is in, we describe our knowledge using a probabilistic state. A probability Pr⁡(x=a)≥0\Pr(x=a)\geq 0 is assigned to each a∈Σa\in\Sigma, and all probabilities together sum to 1.

0
1
Pr⁡(x=0)=34Pr⁡(x=1)=14\Pr(x{=}0)=\tfrac{3}{4} \qquad \Pr(x{=}1)=\tfrac{1}{4}

Probability vectors

The same distribution can be written as a probability vector - a column of numbers, one per state in Σ\Sigma, in a fixed order:

(3414)←  Pr⁡(x=0), the probability the system is in state 0←  Pr⁡(x=1), the probability the system is in state 1\begin{pmatrix}\tfrac{3}{4}\\[4pt]\tfrac{1}{4}\end{pmatrix}\begin{matrix}\leftarrow\;\Pr(x{=}0)\text{, the probability the system is in state }0\\[4pt]\leftarrow\;\Pr(x{=}1)\text{, the probability the system is in state }1\end{matrix}

Dirac notation

The same probability vector can be written with Dirac notation. Suppose the elements of the state set are ordered as Σ=(a1,…,a∣Σ∣)\Sigma=(a_1,\ldots,a_{|\Sigma|}). For any state a∈Σa\in\Sigma, ∣a⟩\lvert a\rangle is the column vector having a 1 in the entry corresponding to aa in that ordering, with 0 for all other entries.

Standard basis vectors:

∣0⟩=(10)∣1⟩=(01)\lvert 0\rangle=\begin{pmatrix}1\\0\end{pmatrix}\qquad \lvert 1\rangle=\begin{pmatrix}0\\1\end{pmatrix}

Any vector is then a combination of standard basis vectors. In our example, 34\tfrac{3}{4} is the probability of ∣0⟩\lvert 0\rangle and 14\tfrac{1}{4} is the probability of ∣1⟩\lvert 1\rangle:

(3414)=34∣0⟩+14∣1⟩\begin{pmatrix}\tfrac{3}{4}\\[4pt]\tfrac{1}{4}\end{pmatrix}=\tfrac{3}{4}\lvert 0\rangle+\tfrac{1}{4}\lvert 1\rangle

These notation versions all describe the same distribution. The probability statements, the column vector, and the Dirac notation are just different tools for writing the same information, each useful in a different setting.

Deterministic operations

A deterministic operation has no chance involved: each input state has exactly one output state. We can write it as a function from the state set to itself:

f:Σ→Σ,a↦f(a)f:\Sigma\to\Sigma,\qquad a\mapsto f(a)

In Dirac notation, each state aa is represented by a basis vector ∣a⟩\lvert a\rangle. The operation becomes a matrix MfM_f that acts on those basis vectors:

Mf∣a⟩=∣f(a)⟩for every a∈ΣM_f\lvert a\rangle=\lvert f(a)\rangle\qquad\text{for every }a\in\Sigma
  • feed in the basis vector ∣a⟩\lvert a\rangle for state aa
  • get out the basis vector ∣f(a)⟩\lvert f(a)\rangle for state f(a)f(a)

The matrix entries are defined by:

(Mf)b,a={1,b=f(a)0,b≠f(a)(M_f)_{b,a}=\begin{cases}1,&b=f(a)\\0,&b\neq f(a)\end{cases}

This says:

  • column aa describes what happens to input state aa
  • the only 11 in that column appears in the row for the output state f(a)f(a)
  • all other entries are 0

For example, if Σ={0,1,2}\Sigma=\{0,1,2\} and f(0)=1, f(1)=1, f(2)=0f(0)=1,\ f(1)=1,\ f(2)=0, then the operation always sends, with no uncertainty involved:

  • 0↦10\mapsto1
  • 1↦11\mapsto1
  • 2↦02\mapsto0

This function can be represented by the matrix MfM_f. Its columns are the possible input states, and its rows are the possible output states. For each input aa, column aa contains a single 11 in row f(a)f(a), with every other entry equal to 00:

Mf=M_f =
(001110000)\begin{pmatrix}0&0&1\\1&1&0\\0&0&0\end{pmatrix}
input 0
←
input 1
←
input 2
←
←row 0: output slot for state 0
←row 1: output slot for state 1
←row 2: output slot for state 2

Matrix-vector updates

Instead of tracking individual state transitions, we can apply the matrix to the entire probability distribution at once. If vv is the current probability vector, the updated distribution is:

v′=Mfvv'=M_fv

Consider a probability distribution where input 0 has probability 12\tfrac{1}{2}, input 1 has probability 13\tfrac{1}{3}, and input 2 has probability 16\tfrac{1}{6}. The vector vv represents this distribution. When MfM_f is applied, the probability associated with input 2 moves to output 0, the probabilities of inputs 0 and 1 are merged at output 1, and output 2 receives no probability mass:

v=(121316)Mfv=(001110000)(121316)=(1612+130)=(16560)v=\begin{pmatrix}\tfrac{1}{2}\\[4pt]\tfrac{1}{3}\\[4pt]\tfrac{1}{6}\end{pmatrix}\qquad M_fv= \begin{pmatrix} 0&0&1\\ 1&1&0\\ 0&0&0 \end{pmatrix} \begin{pmatrix} \tfrac{1}{2}\\[4pt] \tfrac{1}{3}\\[4pt] \tfrac{1}{6} \end{pmatrix} = \begin{pmatrix} \tfrac{1}{6}\\[4pt] \tfrac{1}{2}+\tfrac{1}{3}\\[4pt] 0 \end{pmatrix} = \begin{pmatrix} \tfrac{1}{6}\\[4pt] \tfrac{5}{6}\\[4pt] 0 \end{pmatrix}

We could update the probabilities by following each input separately and moving its probability mass to the corresponding output. The matrix MfM_f packages all of these transfers into a single matrix-vector multiplication, producing the updated probability distribution in one step.

One-bit matrices

For a single bit Σ={0,1}\Sigma=\{0,1\}, there are four possible deterministic operations:

Set to 0

0,1↦00,1\mapsto0
M0=(1100)M_0=\begin{pmatrix}1&1\\0&0\end{pmatrix}

Identity

0↦0,1↦10\mapsto0,\quad1\mapsto1
I=(1001)I=\begin{pmatrix}1&0\\0&1\end{pmatrix}

NOT / bit flip

0↦1,1↦00\mapsto1,\quad1\mapsto0
X=(0110)X=\begin{pmatrix}0&1\\1&0\end{pmatrix}

Set to 1

0,1↦10,1\mapsto1
M1=(0011)M_1=\begin{pmatrix}0&0\\1&1\end{pmatrix}

Bras and inner products

In Dirac notation, column vectors are called kets and are written as ∣0⟩\lvert 0\rangle and ∣1⟩\lvert 1\rangle. Their row-vector counterparts are called bras and are written with the bracket facing the opposite direction:

⟨bra∣ket⟩⟨0∣=(10)⟨1∣=(01)\langle\text{bra}\mid\text{ket}\rangle\qquad \langle 0\rvert=\begin{pmatrix}1&0\end{pmatrix}\qquad \langle 1\rvert=\begin{pmatrix}0&1\end{pmatrix}

Multiplying a row vector by a column vector produces a single number. This operation is called the inner product, or dot product. It measures how much two vectors overlap:

(r1r2⋯rn)(c1c2⋮cn)=r1c1+r2c2+⋯+rncn=∑i=1nrici\begin{pmatrix}r_1&r_2&\cdots&r_n\end{pmatrix}\begin{pmatrix}c_1\\c_2\\\vdots\\c_n\end{pmatrix}=r_1c_1+r_2c_2+\cdots+r_nc_n=\sum_{i=1}^{n}r_ic_i

For basis states, kets and bras contain a single 1 and zeros everywhere else. If we multiply matching states, the 1s line up. If we multiply different states, the 1s occur in different positions, so every term in the sum is zero:

⟨0∣0⟩=(10)(10)=1⋅1+0⋅0=1⟨0∣1⟩=(10)(01)=1⋅0+0⋅1=0\langle 0\vert0\rangle=\begin{pmatrix}1&0\end{pmatrix}\begin{pmatrix}1\\0\end{pmatrix}=1\cdot1+0\cdot0=1\qquad \langle 0\vert1\rangle=\begin{pmatrix}1&0\end{pmatrix}\begin{pmatrix}0\\1\end{pmatrix}=1\cdot0+0\cdot1=0

So multiplying a bra by a ket, denoted as ⟨a∣b⟩\langle a\vert b\rangle, acts like an equality test: it returns 1 when the states are the same and 0 when they are different:

⟨a∣b⟩={1,a=b0,a≠b\langle a\vert b\rangle=\begin{cases}1,&a=b\\0,&a\neq b\end{cases}

Ket-bra products

If row-by-column multiplication gives a single number, then column-by-row multiplication gives a matrix. This operation is called an outer product:

(c1c2⋮cm)(r1r2⋯rn)=(c1r1c1r2⋯c1rnc2r1c2r2⋯c2rn⋮⋮⋱⋮cmr1cmr2⋯cmrn)\begin{pmatrix}c_1\\c_2\\\vdots\\c_m\end{pmatrix}\begin{pmatrix}r_1&r_2&\cdots&r_n\end{pmatrix} = \begin{pmatrix} c_1r_1&c_1r_2&\cdots&c_1r_n\\ c_2r_1&c_2r_2&\cdots&c_2r_n\\ \vdots&\vdots&\ddots&\vdots\\ c_mr_1&c_mr_2&\cdots&c_mr_n \end{pmatrix}

For one-bit basis states, a ket-bra product creates a matrix with a single 1 and zeros everywhere else:

∣0⟩⟨0∣=(10)(10)=(1000)∣0⟩⟨1∣=(10)(01)=(0100)\lvert0\rangle\langle0\rvert= \begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix} = \begin{pmatrix}1&0\\0&0\end{pmatrix} \qquad \lvert0\rangle\langle1\rvert= \begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix} = \begin{pmatrix}0&1\\0&0\end{pmatrix}

Suppose we are given a deterministic operation f:Σ→Σf:\Sigma\to\Sigma where, for each state b∈Σb\in\Sigma, the function specifies an output state f(b)f(b), representing this individual rule b↦f(b)b\mapsto f(b) with the ket-bra operator ∣f(b)⟩⟨b∣\lvert f(b)\rangle\langle b\rvert, where the bra ⟨b∣\langle b\rvert identifies the input state ∣b⟩\lvert b\rangle while the ket ∣f(b)⟩\lvert f(b)\rangle specifies the state that should be produced, so that adding one such term for every possible input state yields a matrix that implements the entire deterministic operation:

Mf=∑b∈Σ∣f(b)⟩⟨b∣M_f=\sum_{b\in\Sigma}\lvert f(b)\rangle\langle b\rvert

For the function f(0)=1, f(1)=1, f(2)=0f(0)=1,\ f(1)=1,\ f(2)=0, the deterministic rules are 0↦1, 1↦1, 2↦00\mapsto1,\ 1\mapsto1,\ 2\mapsto0, and each rule contributes one ket-bra term: the rule 0↦10\mapsto1 becomes ∣1⟩⟨0∣\lvert1\rangle\langle0\rvert, the rule 1↦11\mapsto1 becomes ∣1⟩⟨1∣\lvert1\rangle\langle1\rvert, and the rule 2↦02\mapsto0 becomes ∣0⟩⟨2∣\lvert0\rangle\langle2\rvert — note that while we are used to reading input on the left and output on the right, in ket-bra notation the output comes first: ∣1⟩⟨0∣\lvert1\rangle\langle0\rvert means "produce ∣1⟩\lvert1\rangle when you see ∣0⟩\lvert0\rangle."

Adding these terms gives:

Mf=∣1⟩⟨0∣+∣1⟩⟨1∣+∣0⟩⟨2∣M_f=\lvert1\rangle\langle0\rvert+\lvert1\rangle\langle1\rvert+\lvert0\rangle\langle2\rvert

Applying it to the basis state ∣2⟩\lvert2\rangle, the first two inner products vanish and only the matching term survives:

Mf∣2⟩=∣1⟩⟨0∣2⟩⏟0+∣1⟩⟨1∣2⟩⏟0+∣0⟩⟨2∣2⟩⏟1=∣0⟩M_f\lvert2\rangle=\lvert1\rangle\underbrace{\langle0\vert2\rangle}_{0}+\lvert1\rangle\underbrace{\langle1\vert2\rangle}_{0}+\lvert0\rangle\underbrace{\langle2\vert2\rangle}_{1}=\lvert0\rangle

Each of the three ket-bra terms in MfM_f is an ordinary matrix:

∣1⟩⟨0∣=(000100000)∣1⟩⟨1∣=(000010000)∣0⟩⟨2∣=(001000000)\lvert1\rangle\langle0\rvert= \begin{pmatrix}0&0&0\\1&0&0\\0&0&0\end{pmatrix} \qquad \lvert1\rangle\langle1\rvert= \begin{pmatrix}0&0&0\\0&1&0\\0&0&0\end{pmatrix} \qquad \lvert0\rangle\langle2\rvert= \begin{pmatrix}0&0&1\\0&0&0\\0&0&0\end{pmatrix}

Adding them together gives exactly the deterministic-operation matrix constructed earlier:

Mf=(000100000)+(000010000)+(001000000)=(001110000)M_f= \begin{pmatrix}0&0&0\\1&0&0\\0&0&0\end{pmatrix} + \begin{pmatrix}0&0&0\\0&1&0\\0&0&0\end{pmatrix} + \begin{pmatrix}0&0&1\\0&0&0\\0&0&0\end{pmatrix} = \begin{pmatrix}0&0&1\\1&1&0\\0&0&0\end{pmatrix}

The action of this matrix becomes clear when it is applied to a basis state ∣a⟩\lvert a\rangle. In each term, the inner product ⟨b∣a⟩\langle b\vert a\rangle acts as an equality test: it equals 1 when b=ab=a and 0 otherwise. As a result, every term in the sum vanishes except the one corresponding to b=ab=a, leaving:

Mf∣a⟩=(∑b∈Σ∣f(b)⟩⟨b∣)∣a⟩=∑b∈Σ∣f(b)⟩⟨b∣a⟩=∣f(a)⟩M_f\lvert a\rangle= \left(\sum_{b\in\Sigma}\lvert f(b)\rangle\langle b\rvert\right)\lvert a\rangle =\sum_{b\in\Sigma}\lvert f(b)\rangle\langle b\vert a\rangle =\lvert f(a)\rangle

Thus the matrix sends each basis state ∣a⟩\lvert a\rangle to the state specified by the function, ∣f(a)⟩\lvert f(a)\rangle.

Probabilistic operations

A deterministic operation maps each input state to exactly one output state. A probabilistic operation is more general: each input state produces a probability distribution over output states. The operation is described by a matrix MM where entry Mb,aM_{b,a} gives the probability that input state aa produces output state bb.

For MM to represent a valid probabilistic operation, two conditions must hold. All entries must be nonnegative real numbers:

Mb,a≥0for all a,b∈ΣM_{b,a}\geq 0\qquad\text{for all }a,b\in\Sigma

And the entries in each column must sum to 1 — each column is a probability vector describing the output distribution for one input state:

∑b∈ΣMb,a=1for all a∈Σ\sum_{b\in\Sigma}M_{b,a}=1\qquad\text{for all }a\in\Sigma

A matrix satisfying both conditions is called a stochastic matrix. Every deterministic-operation matrix is a special case: its columns each contain a single 1 and zeros elsewhere, which is a valid probability vector.

Composing operations

When two operations are applied in sequence — first M1M_1, then M2M_2 — the result is M2(M1v)M_2(M_1 v). Because matrix multiplication is associative, this equals (M2M1)v(M_2 M_1)v: the composition is itself a matrix, and the product of two stochastic matrices is stochastic.

M2(M1v)=(M2M1)vM_2(M_1 v)=(M_2 M_1)v

Matrix multiplication is not commutative, however: M2M1M_2 M_1 and M1M2M_1 M_2 generally produce different results. The order in which operations are applied matters.

Using the one-bit matrices introduced earlier, we can see this directly. Applying M0M_0 (set to 0) and then XX (NOT) always produces 1 — equivalent to M1M_1. Doing the same two operations in the opposite order always produces 0 — equivalent to M0M_0 itself:

XM0=(0110)(1100)=(0011)=M1X M_0=\begin{pmatrix}0&1\\1&0\end{pmatrix}\begin{pmatrix}1&1\\0&0\end{pmatrix}=\begin{pmatrix}0&0\\1&1\end{pmatrix}=M_1
M0X=(1100)(0110)=(1100)=M0M_0 X=\begin{pmatrix}1&1\\0&0\end{pmatrix}\begin{pmatrix}0&1\\1&0\end{pmatrix}=\begin{pmatrix}1&1\\0&0\end{pmatrix}=M_0

Quantum information

Two levels of description

Quantum information can be described at two levels of generality. The simplified picture — kets and unitary matrices — is enough for pure states and reversible operations. The general picture adds density matrices and a broader class of measurements and operations, covering mixed states, noise, and measurement.

SimplifiedGeneral
Quantum statesKets — state vectorsDensity matrices
OperationsUnitary matricesMore general class of measurements and operations

Quantum states

A quantum state of a system is represented by a column vector whose indices are placed in correspondence with the classical states of that system:

∣ψ⟩=(α1⋮αn)\lvert\psi\rangle=\begin{pmatrix}\alpha_1\\\vdots\\\alpha_n\end{pmatrix}

A valid quantum state vector must satisfy two conditions:

  • Entries are complex numbers — amplitudes, not probabilities: αk∈C\alpha_k\in\mathbb{C}
  • The sum of squared absolute values of the entries equals 1: ∑k=1n∣αk∣2=1\sum_{k=1}^{n}|\alpha_k|^2=1

The second condition uses the Euclidean norm — the generalisation of vector length to complex numbers:

∥ψ∥=∑k=1n∣αk∣2\|\psi\|=\sqrt{\sum_{k=1}^{n}|\alpha_k|^2}

Real vector

v=(34)v=\begin{pmatrix}3\\4\end{pmatrix}
∥v∥=32+42=25=5\begin{aligned}\|v\|&=\sqrt{3^2+4^2}\\&=\sqrt{25}=5\end{aligned}

Complex vector

v=(32i2)v=\begin{pmatrix}\tfrac{3}{\sqrt{2}}\\[4pt]\tfrac{i}{\sqrt{2}}\end{pmatrix}
∥v∥=∣32∣2+∣i2∣2=92+12=5\|v\|=\sqrt{\left|\tfrac{3}{\sqrt{2}}\right|^2+\left|\tfrac{i}{\sqrt{2}}\right|^2}=\sqrt{\tfrac{9}{2}+\tfrac{1}{2}}=\sqrt{5}

Quantum state vectors are unit vectors — ∥ψ∥=1\|\psi\|=1. The amplitudes encode probability amplitudes: the probability of observing the system in classical state kk is ∣αk∣2|\alpha_k|^2.

Qubit states

A qubit is a quantum system whose classical state set is Σ={0,1}\Sigma=\{0,1\}, so its state vector has two entries. The basis states ∣0⟩\lvert0\rangle and ∣1⟩\lvert1\rangle look like ordinary integer vectors, but their entries are complex numbers that happen to have zero imaginary part:

∣0⟩=(10)∣1⟩=(01)\lvert0\rangle=\begin{pmatrix}1\\0\end{pmatrix}\qquad\lvert1\rangle=\begin{pmatrix}0\\1\end{pmatrix}

A general qubit state is any normalised complex combination of these two basis states:

∣ψ⟩=α∣0⟩+β∣1⟩,α,β∈C,∣α∣2+∣β∣2=1\lvert\psi\rangle=\alpha\lvert0\rangle+\beta\lvert1\rangle,\qquad\alpha,\beta\in\mathbb{C},\qquad|\alpha|^2+|\beta|^2=1

Two particularly important named states are the plus and minus states. They assign equal probability to the basis states ∣0⟩\lvert0\rangle and ∣1⟩\lvert1\rangle — 50% and 50% — but differ in the relative sign between their amplitudes:

∣+⟩=12∣0⟩+12∣1⟩∣−⟩=12∣0⟩−12∣1⟩\lvert+\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{\sqrt{2}}\lvert1\rangle\qquad\lvert-\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle-\tfrac{1}{\sqrt{2}}\lvert1\rangle

Most qubit states have no special name. Any normalised choice of α\alpha and β\beta is valid — for example:

∣ϕ⟩=1+2i3∣0⟩−23∣1⟩∣1+2i3∣2+∣23∣2=59+49=1✓\lvert\phi\rangle=\tfrac{1+2i}{3}\lvert0\rangle-\tfrac{2}{3}\lvert1\rangle\qquad\left|\tfrac{1+2i}{3}\right|^2+\left|\tfrac{2}{3}\right|^2=\tfrac{5}{9}+\tfrac{4}{9}=1\checkmark

Conjugate transpose

Every ket ∣ψ⟩\lvert\psi\rangle has a corresponding bra ⟨ψ∣\langle\psi\rvert obtained by the conjugate transpose — written with a †† (dagger):

⟨ψ∣=∣ψ⟩†\langle\psi\rvert=\lvert\psi\rangle^\dagger

Two steps: transpose the column vector into a row, then replace each entry with its complex conjugate — a+bi  ↦  a−bia+bi\;\mapsto\;a-bi. For the qubit state ∣ϕ⟩=1+2i3∣0⟩−23∣1⟩\lvert\phi\rangle=\tfrac{1+2i}{3}\lvert0\rangle-\tfrac{2}{3}\lvert1\rangle:

⟨ϕ∣=∣ϕ⟩†=1−2i3⟨0∣−23⟨1∣\langle\phi\rvert=\lvert\phi\rangle^\dagger=\frac{1-2i}{3}\langle0\rvert-\frac{2}{3}\langle1\rvert

The real amplitude −23-\tfrac{2}{3} is unchanged — conjugating a real number leaves it the same. Only the complex entry 1+2i3\tfrac{1+2i}{3} flips its imaginary part to give 1−2i3\tfrac{1-2i}{3}.


Why do we conjugate?

A complex number is not just a number—it can also be viewed as a point (or vector) in the complex plane.

z=x+iyz = x + iy

Its length is simply the Euclidean distance from the origin:

∣z∣=x2+y2\lvert z\rvert=\sqrt{x^2+y^2}

The challenge is to compute this length using only complex arithmetic. Notice what happens if we multiply zz by its conjugate:

z‾ z=(x−iy)(x+iy)=x2 + ixy − ixy +y2=x2+y2\overline{z}\,z=(x-iy)(x+iy)=x^2\,\textcolor{#dc2626}{\cancel{+\,ixy}}\,\textcolor{#dc2626}{\cancel{-\,ixy}}\,+y^2=x^2+y^2

The imaginary terms cancel, leaving exactly the square of the Euclidean length:

z‾ z=(x2+y2)2=∣z∣2\overline{z}\,z=\left(\sqrt{x^2+y^2}\right)^2=\lvert z\rvert^2
ReImxyzzzz0
z = 1.20 + 0.60i
z = 1.20 − 0.60i
zz= (1.20)² + (0.60)²= 1.44 + 0.36= 1.80  (real)
|z| = √1.80 = 1.34
Drag z around the plane. Its conjugate z is the mirror image across the real axis, and the product zz always lands on the real axis at |z|².

This is why the complex conjugate appears. It isn't an arbitrary rule—it is the operation that recovers the ordinary Euclidean length of a complex number.

The inner product extends this same idea to vectors by applying the conjugate transpose to every amplitude:

⟨ψ∣ψ⟩=∑kαk‾ αk=∑k∣αk∣2\langle\psi\vert\psi\rangle=\sum_{k}\overline{\alpha_k}\,\alpha_k=\sum_{k}\lvert\alpha_k\rvert^2

Measurements

Measuring a quantum state extracts classical information from it. In a standard basis measurement, the possible outcomes are the classical states — the same states that label the entries of the state vector. For a state ∣ψ⟩=∑a∈Σαa∣a⟩\lvert\psi\rangle=\sum_{a\in\Sigma}\alpha_a\lvert a\rangle, the probability of obtaining outcome aa is the squared absolute value of the corresponding amplitude:

Pr⁡(outcome=a)=∣αa∣2\Pr(\text{outcome}=a)=|\alpha_a|^2

For example, measuring the state ∣+⟩=12∣0⟩+12∣1⟩\lvert{+}\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{\sqrt{2}}\lvert1\rangle gives each outcome with equal probability:

Pr⁡(outcome=0)=∣12∣2=12Pr⁡(outcome=1)=∣12∣2=12\Pr(\text{outcome}=0)=\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}\qquad\Pr(\text{outcome}=1)=\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}

For the state with complex amplitudes:

∣ψ⟩=1+2i3∣0⟩−23∣1⟩\lvert\psi\rangle=\frac{1+2i}{3}\lvert0\rangle-\frac{2}{3}\lvert1\rangle
Pr⁡(outcome=0)=∣1+2i3∣2=12+229=59Pr⁡(outcome=1)=∣23∣2=49\Pr(\text{outcome}=0)=\left|\frac{1+2i}{3}\right|^2=\frac{1^2+2^2}{9}=\frac{5}{9}\qquad\Pr(\text{outcome}=1)=\left|\frac{2}{3}\right|^2=\frac{4}{9}

Measurement also changes the state. Once outcome aa is observed, the quantum state collapses to the corresponding basis state ∣a⟩\lvert a\rangle. A second measurement on the collapsed state will always return the same outcome — this is called the collapse of the quantum state.

The same logic applies to ordinary probability. Flip a coin and let it land face-up on the table. Before you look, each side has probability 12\tfrac{1}{2}. The moment you look, the uncertainty is gone — the coin is showing heads, and re-checking it a second or third time still shows heads with certainty. What changes is not the coin but your state of knowledge about it. Quantum collapse works the same way from the outside: once you have a measurement result, subsequent measurements on that same state are no longer uncertain.

After the measurement, regardless of the pre-measurement state, the system is in a definite classical state ∣0⟩\lvert0\rangle or ∣1⟩\lvert1\rangle. This places a fundamental limit on how much classical information can be extracted from a quantum state in a single measurement.

Unitary operations

Quantum operations are represented by unitary matrices — a different constraint from the stochastic matrices of classical probabilistic operations. A square matrix UU is unitary if its conjugate transpose is also its inverse:

U†U=I=UU†U^\dagger U = I = UU^\dagger

This has two equivalent restatements:

  • The inverse is simply the conjugate transpose — U−1=U†U^{-1}=U^\dagger— so inverting a unitary is cheap.
  • Unitary matrices preserve the Euclidean norm: ∥U∣ψ⟩∥=∥∣ψ⟩∥\|U\lvert\psi\rangle\|=\|\lvert\psi\rangle\|.

To see why the norm is preserved, first notice that for any column vector, multiplying it by its own conjugate transpose gives the squared norm — the conjugates pair with each entry to produce αˉkαk=∣αk∣2\bar\alpha_k\alpha_k=|\alpha_k|^2:

v†v=(αˉ1⋯αˉn)(α1⋮αn)=∣α1∣2+⋯+∣αn∣2=∥v∥2v^\dagger v=\begin{pmatrix}\bar\alpha_1&\cdots&\bar\alpha_n\end{pmatrix}\begin{pmatrix}\alpha_1\\\vdots\\\alpha_n\end{pmatrix}=|\alpha_1|^2+\cdots+|\alpha_n|^2=\|v\|^2

Or, in Dirac notation:

⟨ψ∣ψ⟩=αˉ1α1+⋯+αˉnαn=∣α1∣2+⋯+∣αn∣2=∥∣ψ⟩∥2\langle\psi\vert\psi\rangle=\bar\alpha_1\alpha_1+\cdots+\bar\alpha_n\alpha_n=|\alpha_1|^2+\cdots+|\alpha_n|^2=\|\lvert\psi\rangle\|^2

Since v†v=∥v∥2v^\dagger v=\|v\|^2 holds for any vector, we can apply it to v=U∣ψ⟩v=U\lvert\psi\rangle. The rule (AB)†=B†A†(AB)^\dagger=B^\dagger A^\dagger says the conjugate transpose of a product reverses the order — so:

(U∣ψ⟩)†=∣ψ⟩†U†=⟨ψ∣U†(U\lvert\psi\rangle)^\dagger=\lvert\psi\rangle^\dagger U^\dagger=\langle\psi\rvert U^\dagger

Substituting into v†vv^\dagger v and applying U†U=IU^\dagger U=I:

∥U∣ψ⟩∥2=⟨ψ∣U†U∣ψ⟩=⟨ψ∣I∣ψ⟩=⟨ψ∣ψ⟩=∥∣ψ⟩∥2\|U\lvert\psi\rangle\|^2=\langle\psi\rvert U^\dagger U\lvert\psi\rangle=\langle\psi\rvert I\lvert\psi\rangle=\langle\psi\vert\psi\rangle=\|\lvert\psi\rangle\|^2

Taking square roots gives ∥U∣ψ⟩∥=∥∣ψ⟩∥\|U\lvert\psi\rangle\|=\|\lvert\psi\rangle\|. As a numeric check, applying σx\sigma_x to the familiar state:

∣ϕ⟩=1+2i3∣0⟩−23∣1⟩\lvert\phi\rangle=\tfrac{1+2i}{3}\lvert0\rangle-\tfrac{2}{3}\lvert1\rangle

It simply swaps the two entries:

σx∣ϕ⟩=(0110)(1+2i3−23)=(−231+2i3)\sigma_x\lvert\phi\rangle=\begin{pmatrix}0&1\\1&0\end{pmatrix}\begin{pmatrix}\tfrac{1+2i}{3}\\[4pt]-\tfrac{2}{3}\end{pmatrix}=\begin{pmatrix}-\tfrac{2}{3}\\[4pt]\tfrac{1+2i}{3}\end{pmatrix}
∥∣ϕ⟩∥=59+49=1∥σx∣ϕ⟩∥=49+59=1\|\lvert\phi\rangle\|=\sqrt{\tfrac{5}{9}+\tfrac{4}{9}}=1\qquad\|\sigma_x\lvert\phi\rangle\|=\sqrt{\tfrac{4}{9}+\tfrac{5}{9}}=1

Geometrically, a unitary transformation is the complex-number analogue of a rotation: it changes the direction of the vector, but never its length. Since quantum state vectors are unit vectors, a unitary operation always maps valid quantum states to valid quantum states — it can never take a state outside the unit sphere.

To check if a matrix is unitary, multiply it by its conjugate transpose and see if the result is the identity.

For σx\sigma_x, which is real and symmetric so σx†=σx\sigma_x^\dagger=\sigma_x:

σx†σx=(0110)(0110)=(1001)=I✓\begin{aligned} \sigma_x^\dagger\sigma_x &=\begin{pmatrix}0&1\\1&0\end{pmatrix}\begin{pmatrix}0&1\\1&0\end{pmatrix}\\ &=\begin{pmatrix}1&0\\0&1\end{pmatrix}=I\checkmark \end{aligned}

Qubit unitary operations

1. Pauli operations

The four Pauli matrices are the most common single-qubit unitary operations:

Identity

I=(1001)I=\begin{pmatrix}1&0\\0&1\end{pmatrix}

Bit flip

σx=(0110)\sigma_x=\begin{pmatrix}0&1\\1&0\end{pmatrix}

Phase + bit flip

σy=(0−ii0)\sigma_y=\begin{pmatrix}0&-i\\i&0\end{pmatrix}

Phase flip

σz=(100−1)\sigma_z=\begin{pmatrix}1&0\\0&-1\end{pmatrix}

The Pauli matrices also happen to be Hermitian — a matrix is Hermitian if it equals its own conjugate transpose, M†=MM^\dagger=M. For a real symmetric matrix this just means symmetry across the diagonal. For a complex matrix, entries are mirrored across the diagonal and conjugated.

2. Hadamard

The Hadamard gate H, named after the French mathematician Jacques Hadamard, is one of the most important quantum gates. It acts as a bridge between two ways of describing a qubit:

  • The computational basis: ∣0⟩\lvert0\rangle and ∣1⟩\lvert1\rangle
  • The superposition basis: ∣+⟩\lvert+\rangle and ∣−⟩\lvert-\rangle
H=(121212−12)=12(111−1)H=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix}

From 0, 1 to +, -:

H∣0⟩=(121212−12)(10)=(1212)=∣+⟩H\lvert0\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}1\\0\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}\end{pmatrix}=\lvert+\rangle
H∣1⟩=(121212−12)(01)=(12−12)=∣−⟩H\lvert1\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}0\\1\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[6pt]-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\lvert-\rangle

From +, - back to 0, 1:

H∣+⟩=(121212−12)(1212)=(10)=∣0⟩H\lvert+\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}\end{pmatrix}=\begin{pmatrix}1\\0\end{pmatrix}=\lvert0\rangle
H∣−⟩=(121212−12)(12−12)=(01)=∣1⟩H\lvert-\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[6pt]-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\begin{pmatrix}0\\1\end{pmatrix}=\lvert1\rangle

Applying a Hadamard gate transforms a computational basis state into an equal superposition of ∣0⟩\lvert0\rangle and ∣1⟩\lvert1\rangle, meaning that a measurement would find each outcome with equal probability:

H∣0⟩=∣+⟩=∣0⟩+∣1⟩2,H∣1⟩=∣−⟩=∣0⟩−∣1⟩2H\lvert0\rangle=\lvert+\rangle=\frac{\lvert0\rangle+\lvert1\rangle}{\sqrt{2}},\qquad H\lvert1\rangle=\lvert-\rangle=\frac{\lvert0\rangle-\lvert1\rangle}{\sqrt{2}}

You can think of the Hadamard gate as a quantum "basis changer" that lets us move between definite states and equal-probability superpositions, making it a fundamental building block of many quantum algorithms.

3. Phase gates

Phase gates leave ∣0⟩\lvert0\rangle unchanged and rotate ∣1⟩\lvert1\rangle by a complex phase. They do not change measurement probabilities on their own, but they shift the relative phase between amplitudes, which affects how states interfere in a larger circuit.

The general phase rotation gate RϕR_\phi applies a phase eiϕe^{i\phi} to ∣1⟩\lvert1\rangle while leaving ∣0⟩\lvert0\rangle alone. ϕ\phi is a real number, making iϕi\phi always purely imaginary. This means eiϕe^{i\phi} always lies on the complex unit circle — ∣eiϕ∣=1|e^{i\phi}|=1 — so it is a pure phase that rotates without scaling:

Rϕ=(100eiϕ)R_\phi=\begin{pmatrix}1&0\\0&e^{i\phi}\end{pmatrix}

Two standard choices give the S gate and the T gate:

S=Rπ/2=(100i)T=Rπ/4=(100eiπ/4)S=R_{\pi/2}=\begin{pmatrix}1&0\\0&i\end{pmatrix}\qquad T=R_{\pi/4}=\begin{pmatrix}1&0\\0&e^{i\pi/4}\end{pmatrix}

S applies a quarter-turn phase (90°) and satisfies S2=ZS^2=Z, where Z is the Pauli phase-flip gate from section 1: Z=(100−1)Z=\begin{pmatrix}1&0\\0&-1\end{pmatrix}. T applies an eighth-turn phase (45°) and satisfies T2=ST^2=S and T4=ZT^4=Z. Together with H, T generates a gate set that can approximate any single-qubit unitary to arbitrary precision.

Applied to the basis states:

S∣0⟩=∣0⟩,S∣1⟩=i∣1⟩S\lvert0\rangle=\lvert0\rangle,\quad S\lvert1\rangle=i\lvert1\rangle
T∣0⟩=∣0⟩,T∣1⟩=eiπ/4∣1⟩T\lvert0\rangle=\lvert0\rangle,\quad T\lvert1\rangle=e^{i\pi/4}\lvert1\rangle

Composing unitary operations

Gates compose by matrix multiplication. In the product the rightmost matrix acts first — the state travels right to left through the sequence:

∣ψout⟩=U3(3)  U2(2)  U1(1)∣ψin⟩\lvert\psi_\text{out}\rangle=\overset{(3)}{U_3}\;\overset{(2)}{U_2}\;\overset{(1)}{U_1}\lvert\psi_\text{in}\rangle

For example, in HSH\textcolor{#2563eb}{H}\textcolor{#d97706}{S}\textcolor{#7c3aed}{H} the rightmost gate (purple H) acts first, then S, then the leftmost H (blue). Expanding step by step:

HSH=(121212−12)(100i)(121212−12)=(121212−12)(1212i2−i2)=12(1+i1−i1−i1+i)=X\textcolor{#2563eb}{H}\textcolor{#d97706}{S}\textcolor{#7c3aed}{H}=\textcolor{#2563eb}{\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}}\textcolor{#d97706}{\begin{pmatrix}1&0\\0&i\end{pmatrix}}\textcolor{#7c3aed}{\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}}=\textcolor{#2563eb}{\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}}\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{i}{\sqrt{2}}&-\tfrac{i}{\sqrt{2}}\end{pmatrix}=\frac{1}{2}\begin{pmatrix}1+i&1-i\\1-i&1+i\end{pmatrix}=\sqrt{X}

This matrix is called the square root of NOT (written X\sqrt{X}), because applying it twice gives the Pauli X (NOT) gate. To see why, insert HH=IHH=I in the middle and use S2=ZS^2=Z:

(HSH)2=HSHH⏟ISH=HS2H=HZH=X(HSH)^2=HS\underbrace{HH}_{I}SH=HS^2H=HZH=X

In ordinary arithmetic, squaring something makes it "more of the same" — you would never expect a number squared to flip a sign. But unitary matrices can have complex eigenvalues such as ii, and i2=−1i^2=-1 introduces the sign change that turns a partial rotation into a full logical inversion. This is one of the ways quantum gates behave fundamentally differently from classical boolean operations.

As another example, consider HTH. T is a rotation around the Z axis by π/4\pi/4. Placing H on both sides redirects that same rotation onto the X axis:

HTH=12(1+eiπ/41−eiπ/41−eiπ/41+eiπ/4)=eiπ/8(cos⁡π8−isin⁡π8−isin⁡π8cos⁡π8)=eiπ/8Rx ⁣(π4)HTH=\frac{1}{2}\begin{pmatrix}1+e^{i\pi/4}&1-e^{i\pi/4}\\1-e^{i\pi/4}&1+e^{i\pi/4}\end{pmatrix}=e^{i\pi/8}\begin{pmatrix}\cos\tfrac{\pi}{8}&-i\sin\tfrac{\pi}{8}\\-i\sin\tfrac{\pi}{8}&\cos\tfrac{\pi}{8}\end{pmatrix}=e^{i\pi/8}R_x\!\left(\tfrac{\pi}{4}\right)

So T and HTH are rotations around two non-parallel axes — Z and X — each by π/4\pi/4. Combining rotations around any two non-parallel axes generates all rotations of the sphere, so any single-qubit unitary can be reached by some finite sequence of T and HTH to any desired precision.

Multiple systems: classical

Classical states

Suppose we have two systems:

  • XX with classical state set Σ\Sigma.
  • YY with classical state set Γ\Gamma.

Together they form a compound system, written (X,Y)(X,Y) or simply XYXY. At any moment the compound system is in exactly one state - a pair (a,b)(a,b) where a∈Σa\in\Sigma is the state of XX and b∈Γb\in\Gamma is the state of YY.

The full set of possible states of XYXY is the Cartesian product:

Σ×Γ={(a,b):a∈Σ, b∈Γ}\Sigma\times\Gamma=\{(a,b):a\in\Sigma,\,b\in\Gamma\}

For example, if both XX and YY are bits so Σ=Γ={0,1}\Sigma=\Gamma=\{0,1\}, the compound system has four possible states:

Σ×Γ={(0,0),(0,1),(1,0),(1,1)}\Sigma\times\Gamma=\{(0,0),(0,1),(1,0),(1,1)\}

General formula: if we combine nn classical systems with state sets Σ1,…,Σn\Sigma_1,\ldots,\Sigma_n, their compound state set is:

Σ1×⋯×Σn={(a1,…,an):ai∈Σi for i=1,…,n}\Sigma_1\times\cdots\times\Sigma_n=\{(a_1,\ldots,a_n):a_i\in\Sigma_i\text{ for }i=1,\ldots,n\}

Convention notes

  • When we list a Cartesian product, we usually use lexicographic order: compare the first coordinate, then the second, and so on.
  • This assumes each state set already has an order. For bits, we use 0<10<1.
  • Significance decreases from left to right: the leftmost coordinate is the most significant, and the rightmost coordinate changes fastest.

Example: for three bits, Σ1=Σ2=Σ3={0,1}\Sigma_1=\Sigma_2=\Sigma_3=\{0,1\}. In lexicographic order:

Σ1×Σ2×Σ3={(0,0,0),(0,0,1),(0,1,0),(0,1,1),(1,0,0),(1,0,1),(1,1,0),(1,1,1)}\Sigma_1\times\Sigma_2\times\Sigma_3=\{(0,0,0),(0,0,1),(0,1,0),(0,1,1),(1,0,0),(1,0,1),(1,1,0),(1,1,1)\}

Probabilistic states

For a compound classical system, a probabilistic state assigns one probability to each state in the Cartesian product. If xx is the current state of XX and yy is the current state of YY, then:

Pr⁡((x,y)=(a,b))≥0for all (a,b)∈Σ×Γ,∑(a,b)∈Σ×ΓPr⁡((x,y)=(a,b))=1\Pr\bigl((x,y)=(a,b)\bigr)\geq 0\quad\text{for all }(a,b)\in\Sigma\times\Gamma,\qquad\sum_{(a,b)\in\Sigma\times\Gamma}\Pr\bigl((x,y)=(a,b)\bigr)=1

For example, for the system of two bits, where Σ=Γ={0,1}\Sigma=\Gamma=\{0,1\} and the possible compound states are (0,0),(0,1),(1,0),(1,1)(0,0),(0,1),(1,0),(1,1), one of multiple possible probabilistic states can be:

Pr⁡((x,y)=(0,0))=12Pr⁡((x,y)=(0,1))=0Pr⁡((x,y)=(1,0))=0Pr⁡((x,y)=(1,1))=12\begin{aligned} \Pr\bigl((x,y)=(0,0)\bigr)&=\tfrac{1}{2}\\[4pt] \Pr\bigl((x,y)=(0,1)\bigr)&=0\\[4pt] \Pr\bigl((x,y)=(1,0)\bigr)&=0\\[4pt] \Pr\bigl((x,y)=(1,1)\bigr)&=\tfrac{1}{2} \end{aligned}

The two-bit system is equally likely to be in state (0,0)(0,0) or state (1,1)(1,1), and has probability zero of being in the other two states. In vector form (using lexicographic order):

u=(120012)←  probability associated with state 00←  probability associated with state 01←  probability associated with state 10←  probability associated with state 11u=\begin{pmatrix} \tfrac{1}{2}\\[4pt] 0\\[4pt] 0\\[4pt] \tfrac{1}{2} \end{pmatrix} \begin{matrix} \leftarrow\;\text{probability associated with state }00\\[4pt] \leftarrow\;\text{probability associated with state }01\\[4pt] \leftarrow\;\text{probability associated with state }10\\[4pt] \leftarrow\;\text{probability associated with state }11 \end{matrix}

For another two-bit system, the probability state might look like:

v=(14141414)←  probability associated with state 00←  probability associated with state 01←  probability associated with state 10←  probability associated with state 11v=\begin{pmatrix} \tfrac{1}{4}\\[4pt] \tfrac{1}{4}\\[4pt] \tfrac{1}{4}\\[4pt] \tfrac{1}{4} \end{pmatrix} \begin{matrix} \leftarrow\;\text{probability associated with state }00\\[4pt] \leftarrow\;\text{probability associated with state }01\\[4pt] \leftarrow\;\text{probability associated with state }10\\[4pt] \leftarrow\;\text{probability associated with state }11 \end{matrix}

For a given probabilistic state of (X,Y)(X,Y), we say XX and YY are independent if Pr⁡((x,y)=(a,b))=Pr⁡(x=a) Pr⁡(y=b)\Pr\bigl((x,y)=(a,b)\bigr)=\Pr(x=a)\,\Pr(y=b) for all a∈Σa\in\Sigma and b∈Γb\in\Gamma.

Let's check uu and vv against this rule — by looking for a contradiction.

u — suppose the rule holds:

u=(120012)← ① Pr⁡(x=0)>0← ③ =0, so Pr⁡(x=0)=0 or Pr⁡(y=1)=0 — contradicts ①② ✗←← ② Pr⁡(y=1)>0u=\begin{pmatrix}\tfrac{1}{2}\\[4pt]0\\[4pt]0\\[4pt]\tfrac{1}{2}\end{pmatrix}\begin{array}{l}\leftarrow\text{ ① }\Pr(x{=}0)>0\\[4pt]\leftarrow\text{ ③ }=0,\text{ so }\Pr(x{=}0){=}0\text{ or }\Pr(y{=}1){=}0\text{ — contradicts ①② ✗}\\[4pt]\phantom{\leftarrow}\\[4pt]\leftarrow\text{ ② }\Pr(y{=}1)>0\end{array}

Contradiction — XX and YY are not independent under uu.

v — no contradiction arises:

v=(14141414)←  14=12⋅12=Pr⁡(x=0) Pr⁡(y=0)  ✓←  14=12⋅12=Pr⁡(x=0) Pr⁡(y=1)  ✓←  14=12⋅12=Pr⁡(x=1) Pr⁡(y=0)  ✓←  14=12⋅12=Pr⁡(x=1) Pr⁡(y=1)  ✓v=\begin{pmatrix}\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\end{pmatrix}\begin{matrix}\leftarrow\;\tfrac{1}{4}=\tfrac{1}{2}\cdot\tfrac{1}{2}=\Pr(x{=}0)\,\Pr(y{=}0)\;\checkmark\\[4pt]\leftarrow\;\tfrac{1}{4}=\tfrac{1}{2}\cdot\tfrac{1}{2}=\Pr(x{=}0)\,\Pr(y{=}1)\;\checkmark\\[4pt]\leftarrow\;\tfrac{1}{4}=\tfrac{1}{2}\cdot\tfrac{1}{2}=\Pr(x{=}1)\,\Pr(y{=}0)\;\checkmark\\[4pt]\leftarrow\;\tfrac{1}{4}=\tfrac{1}{2}\cdot\tfrac{1}{2}=\Pr(x{=}1)\,\Pr(y{=}1)\;\checkmark\end{matrix}

XX and YY are independent under vv.

Correlation is, in a sense, a lack of independence: when two systems are not independent, the state of one carries information about the state of the other.

Dirac notation

The same rule for independence can be written in Dirac notation. Suppose that a probabilistic state of (X,Y)(X,Y) is expressed as a vector:

∣π⟩=∑(a,b)∈Σ×Γpab ∣ab⟩|\pi\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|ab\rangle

The systems XX and YY are independent if there exist probability vectors

∣ϕ⟩=∑a∈Σqa ∣a⟩and∣ψ⟩=∑b∈Γrb ∣b⟩|\phi\rangle=\sum_{a\in\Sigma}q_a\,|a\rangle\quad\text{and}\quad|\psi\rangle=\sum_{b\in\Gamma}r_b\,|b\rangle

such that pab=qarbp_{ab}=q_a r_b for all a∈Σa\in\Sigma and b∈Γb\in\Gamma.

  • ∣π⟩|\pi\rangle — the probabilistic state of the compound system (X,Y)(X,Y)
  • pabp_{ab} — the probability assigned to outcome (a,b)(a,b)
  • ∣ab⟩|ab\rangle — a basis vector labelling the outcome X=a, Y=bX{=}a,\,Y{=}bfor X,Y∈{0,1}X,Y\in\{0,1\} these are the four standard basis vectors:
∣00⟩=(1000),∣01⟩=(0100),∣10⟩=(0010),∣11⟩=(0001)|00\rangle=\begin{pmatrix}1\\0\\0\\0\end{pmatrix},\quad|01\rangle=\begin{pmatrix}0\\1\\0\\0\end{pmatrix},\quad|10\rangle=\begin{pmatrix}0\\0\\1\\0\end{pmatrix},\quad|11\rangle=\begin{pmatrix}0\\0\\0\\1\end{pmatrix}

Returning to uu and vv from the earlier example. For uu, its column vector alongside its ∣π⟩|\pi\rangle expansion:

u=(120012)∣π⟩=12 ∣00⟩+0⋅∣01⟩+0⋅∣10⟩+12 ∣11⟩u=\begin{pmatrix}\textcolor{blue}{\tfrac{1}{2}}\\[4pt]0\\[4pt]0\\[4pt]\textcolor{orange}{\tfrac{1}{2}}\end{pmatrix}\qquad|\pi\rangle=\textcolor{blue}{\tfrac{1}{2}}\,|00\rangle+0\cdot|01\rangle+0\cdot|10\rangle+\textcolor{orange}{\tfrac{1}{2}}\,|11\rangle

For vv, its column vector alongside its ∣π⟩|\pi\rangle expansion:

v=(14141414)∣π⟩=14 ∣00⟩+14 ∣01⟩+14 ∣10⟩+14 ∣11⟩v=\begin{pmatrix}\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\end{pmatrix}\qquad|\pi\rangle=\tfrac{1}{4}\,|00\rangle+\tfrac{1}{4}\,|01\rangle+\tfrac{1}{4}\,|10\rangle+\tfrac{1}{4}\,|11\rangle

Since XX and YY are independent under vv, there exist probability vectors ∣ϕ⟩|\phi\rangle and ∣ψ⟩|\psi\rangle such that each coefficient in ∣π⟩|\pi\rangle is a product of one factor from each:

∣ϕ⟩=12 ∣0⟩+12 ∣1⟩∣ψ⟩=12 ∣0⟩+12 ∣1⟩|\phi\rangle=\textcolor{blue}{\tfrac{1}{2}}\,|0\rangle+\textcolor{teal}{\tfrac{1}{2}}\,|1\rangle\qquad|\psi\rangle=\textcolor{orange}{\tfrac{1}{2}}\,|0\rangle+\textcolor{violet}{\tfrac{1}{2}}\,|1\rangle
∣π⟩=12⋅12 ∣00⟩+12⋅12 ∣01⟩+12⋅12 ∣10⟩+12⋅12 ∣11⟩|\pi\rangle=\textcolor{blue}{\tfrac{1}{2}}\cdot\textcolor{orange}{\tfrac{1}{2}}\,|00\rangle+\textcolor{blue}{\tfrac{1}{2}}\cdot\textcolor{violet}{\tfrac{1}{2}}\,|01\rangle+\textcolor{teal}{\tfrac{1}{2}}\cdot\textcolor{orange}{\tfrac{1}{2}}\,|10\rangle+\textcolor{teal}{\tfrac{1}{2}}\cdot\textcolor{violet}{\tfrac{1}{2}}\,|11\rangle

Tensor products of vectors

When working with multiple systems, we need a way to combine their vector spaces into a larger one. The mathematical operation that does this is the tensor product.

Given two vectors, ∣ϕ⟩|\phi\rangle and ∣ψ⟩|\psi\rangle, their tensor product, written ∣ϕ⟩⊗∣ψ⟩|\phi\rangle\otimes|\psi\rangle, represents the combined state of both systems. Every basis state of the first vector is paired with every basis state of the second, and the corresponding coefficients are multiplied.

∣ϕ⟩=∑a∈Σαa ∣a⟩and∣ψ⟩=∑b∈Γβb ∣b⟩|\phi\rangle=\sum_{a\in\Sigma}\alpha_a\,|a\rangle\quad\text{and}\quad|\psi\rangle=\sum_{b\in\Gamma}\beta_b\,|b\rangle
∣ϕ⟩⊗∣ψ⟩=∑(a,b)∈Σ×Γαaβb ∣ab⟩|\phi\rangle\otimes|\psi\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}\alpha_a\beta_b\,|ab\rangle
  • ∣ϕ⟩|\phi\rangle — probability vector for system XX, with coefficients αa\alpha_a
  • ∣ψ⟩|\psi\rangle — probability vector for system YY, with coefficients βb\beta_b
  • αaβb\alpha_a\beta_b — coefficient of the compound basis state ∣ab⟩|ab\rangle in ∣ϕ⟩⊗∣ψ⟩|\phi\rangle\otimes|\psi\rangle
  • Σ×Γ\Sigma\times\Gamma — the Cartesian product of the two state sets; the sum runs over all possible pairs (a,b)(a,b)

Inner product form

Equivalently, the vector ∣π⟩=∣ϕ⟩⊗∣ψ⟩|\pi\rangle=|\phi\rangle\otimes|\psi\rangle is defined by this condition:

⟨ab∣π⟩=⟨a∣ϕ⟩⟨b∣ψ⟩(for all a∈Σ and b∈Γ)\langle ab|\pi\rangle=\langle a|\phi\rangle\langle b|\psi\rangle\qquad(\text{for all }a\in\Sigma\text{ and }b\in\Gamma)

Here, ⟨a∣ϕ⟩\langle a|\phi\rangle is a single number — the coefficient (amplitude) of basis state ∣a⟩|a\rangle in ∣ϕ⟩|\phi\rangle; and ⟨b∣ψ⟩\langle b|\psi\rangle is also a single number — the coefficient of basis state ∣b⟩|b\rangle in ∣ψ⟩|\psi\rangle:

⟨a∣ϕ⟩=αa⟨b∣ψ⟩=βb\langle a|\phi\rangle=\alpha_a\qquad\langle b|\psi\rangle=\beta_b

So the condition says: the coefficient of the combined basis state ∣ab⟩|ab\rangle in ∣π⟩|\pi\rangle is the product of the coefficients of ∣a⟩|a\rangle and ∣b⟩|b\rangle separately. ⟨ab∣π⟩\langle ab|\pi\rangle means: take the vector ∣π⟩|\pi\rangle and ask what its coefficient is along basis state ∣ab⟩|ab\rangle.

For example, suppose

∣π⟩=0.3 ∣00⟩+0.4 ∣01⟩+0.1 ∣10⟩+0.2 ∣11⟩|\pi\rangle=0.3\,|00\rangle+\textcolor{teal}{0.4}\,|01\rangle+0.1\,|10\rangle+0.2\,|11\rangle

Then ⟨01∣π⟩=0.4\langle \textcolor{teal}{01}|\pi\rangle=\textcolor{teal}{0.4}. It is just extracting one coefficient.

Now suppose

∣ϕ⟩=12 ∣0⟩+12 ∣1⟩and∣ψ⟩=35 ∣0⟩+45 ∣1⟩|\phi\rangle=\textcolor{blue}{\tfrac{1}{\sqrt{2}}}\,|0\rangle+\tfrac{1}{\sqrt{2}}\,|1\rangle\qquad\text{and}\qquad|\psi\rangle=\tfrac{3}{5}\,|0\rangle+\textcolor{orange}{\tfrac{4}{5}}\,|1\rangle

Then:

⟨0∣ϕ⟩=12,⟨1∣ψ⟩=45\langle \textcolor{blue}{0}|\phi\rangle=\textcolor{blue}{\tfrac{1}{\sqrt{2}}},\qquad\langle \textcolor{orange}{1}|\psi\rangle=\textcolor{orange}{\tfrac{4}{5}}

Therefore:

⟨01∣π⟩=⟨0∣ϕ⟩⟨1∣ψ⟩=12⋅45\langle \textcolor{blue}{0}\textcolor{orange}{1}|\pi\rangle=\langle \textcolor{blue}{0}|\phi\rangle\langle \textcolor{orange}{1}|\psi\rangle=\textcolor{blue}{\tfrac{1}{\sqrt{2}}}\cdot\textcolor{orange}{\tfrac{4}{5}}

The amplitude of the combined outcome 0101 equals the product of the amplitudes of the individual outcomes 00 and 11.

Column vector form

The tensor product can be viewed as a "multiply every entry by every entry" operation. Each coefficient of the first vector is paired with every coefficient of the second, producing a larger vector that represents all possible combinations of the two systems. As a result, dimensions multiply: if ∣ϕ⟩|\phi\rangle has mm entries and ∣ψ⟩|\psi\rangle has kk entries, then ∣ϕ⟩⊗∣ψ⟩|\phi\rangle\otimes|\psi\rangle has m⋅km\cdot k entries.

(α1⋮αm)⊗(β1⋮βk)=(α1β1⋮α1βkα2β1⋮α2βk⋮αmβ1⋮αmβk)\begin{pmatrix}\alpha_1\\\vdots\\\alpha_m\end{pmatrix}\otimes\begin{pmatrix}\beta_1\\\vdots\\\beta_k\end{pmatrix}=\begin{pmatrix}\alpha_1\beta_1\\\vdots\\\alpha_1\beta_k\\\alpha_2\beta_1\\\vdots\\\alpha_2\beta_k\\\vdots\\\alpha_m\beta_1\\\vdots\\\alpha_m\beta_k\end{pmatrix}

Tensor product of standard basis vectors

The tensor product of two standard basis vectors is often written by simply combining their labels into a single basis label. Thus, instead of writing ∣a⟩⊗∣b⟩|a\rangle\otimes|b\rangle, we commonly write ∣ab⟩|ab\rangle, which can be viewed as shorthand for the basis vector indexed by the pair (a,b)(a,b). More explicitly, one could write ∣(a,b)⟩|(a,b)\rangle, but in practice the parentheses are usually omitted and the notation ∣a,b⟩|a,b\rangle is preferred. This follows a common mathematical convention of removing symbols that do not add information or eliminate ambiguity. From a mathematician's perspective, once the structure is understood, the parentheses are carrying no real content and can be safely discarded. Of course, for anyone still getting comfortable with Dirac notation, the more explicit form ∣(a,b)⟩|(a,b)\rangle can be a useful stepping stone — it makes the two-label structure impossible to miss, and once that structure feels natural, dropping the parentheses costs nothing.

For example, for two-bit systems where Σ=Γ={0,1}\Sigma=\Gamma=\{0,1\}, each of the four standard basis vectors of the compound system arises as a tensor product:

∣0⟩⊗∣0⟩=∣00⟩,∣0⟩⊗∣1⟩=∣01⟩,∣1⟩⊗∣0⟩=∣10⟩,∣1⟩⊗∣1⟩=∣11⟩|0\rangle\otimes|0\rangle=|00\rangle,\quad|0\rangle\otimes|1\rangle=|01\rangle,\quad|1\rangle\otimes|0\rangle=|10\rangle,\quad|1\rangle\otimes|1\rangle=|11\rangle

To see this concretely, take ∣0⟩⊗∣1⟩|0\rangle\otimes|1\rangle and apply the column-vector rule — multiply every entry of the first vector by every entry of the second:

∣0⟩⊗∣1⟩=(10)⊗(01)=(1⋅01⋅10⋅00⋅1)=(0100)=∣01⟩|0\rangle\otimes|1\rangle=\begin{pmatrix}1\\0\end{pmatrix}\otimes\begin{pmatrix}0\\1\end{pmatrix}=\begin{pmatrix}1\cdot0\\1\cdot1\\0\cdot0\\0\cdot1\end{pmatrix}=\begin{pmatrix}0\\1\\0\\0\end{pmatrix}=|01\rangle

The result is exactly the standard basis vector ∣01⟩|01\rangle — confirming that the label shorthand and the column-vector computation agree.

The same shorthand is used for basis bras: ⟨a∣⊗⟨b∣=⟨ab∣\langle a|\otimes\langle b|=\langle ab|. For two-bit systems, ⟨0∣⊗⟨0∣=⟨00∣\langle 0|\otimes\langle 0|=\langle 00|, ⟨0∣⊗⟨1∣=⟨01∣\langle 0|\otimes\langle 1|=\langle 01|, ⟨1∣⊗⟨0∣=⟨10∣\langle 1|\otimes\langle 0|=\langle 10|, and ⟨1∣⊗⟨1∣=⟨11∣\langle 1|\otimes\langle 1|=\langle 11|.

Properties of tensor product

The tensor product is bilinear — it preserves the familiar rules of linearity in both of its arguments. You can distribute over addition and pull out scalar factors from either side independently.

First argument

(∣ϕ1⟩+∣ϕ2⟩)⊗∣ψ⟩=∣ϕ1⟩⊗∣ψ⟩+∣ϕ2⟩⊗∣ψ⟩(|\phi_1\rangle+|\phi_2\rangle)\otimes|\psi\rangle=|\phi_1\rangle\otimes|\psi\rangle+|\phi_2\rangle\otimes|\psi\rangle
(c ∣ϕ⟩)⊗∣ψ⟩=c (∣ϕ⟩⊗∣ψ⟩)(c\,|\phi\rangle)\otimes|\psi\rangle=c\,(|\phi\rangle\otimes|\psi\rangle)

For example, take ∣ϕ1⟩=∣0⟩|\phi_1\rangle=|0\rangle, ∣ϕ2⟩=∣1⟩|\phi_2\rangle=|1\rangle, ∣ψ⟩=∣0⟩|\psi\rangle=|0\rangle:

(∣0⟩+∣1⟩)⊗∣0⟩=(11)⊗(10)=(1010)=∣00⟩+∣10⟩(|0\rangle+|1\rangle)\otimes|0\rangle=\begin{pmatrix}1\\1\end{pmatrix}\otimes\begin{pmatrix}1\\0\end{pmatrix}=\begin{pmatrix}1\\0\\1\\0\end{pmatrix}=|00\rangle+|10\rangle

For scalar multiplication, take c=3c=3, ∣ϕ⟩=∣0⟩|\phi\rangle=|0\rangle, ∣ψ⟩=∣1⟩|\psi\rangle=|1\rangle:

(3 ∣0⟩)⊗∣1⟩=(30)⊗(01)=(0300)=3 ∣01⟩(3\,|0\rangle)\otimes|1\rangle=\begin{pmatrix}3\\0\end{pmatrix}\otimes\begin{pmatrix}0\\1\end{pmatrix}=\begin{pmatrix}0\\3\\0\\0\end{pmatrix}=3\,|01\rangle

Second argument

∣ϕ⟩⊗(∣ψ1⟩+∣ψ2⟩)=∣ϕ⟩⊗∣ψ1⟩+∣ϕ⟩⊗∣ψ2⟩|\phi\rangle\otimes(|\psi_1\rangle+|\psi_2\rangle)=|\phi\rangle\otimes|\psi_1\rangle+|\phi\rangle\otimes|\psi_2\rangle
∣ϕ⟩⊗(c ∣ψ⟩)=c (∣ϕ⟩⊗∣ψ⟩)|\phi\rangle\otimes(c\,|\psi\rangle)=c\,(|\phi\rangle\otimes|\psi\rangle)

For example, take ∣ϕ⟩=∣1⟩|\phi\rangle=|1\rangle, ∣ψ1⟩=∣0⟩|\psi_1\rangle=|0\rangle, ∣ψ2⟩=∣1⟩|\psi_2\rangle=|1\rangle:

∣1⟩⊗(∣0⟩+∣1⟩)=(01)⊗(11)=(0011)=∣10⟩+∣11⟩|1\rangle\otimes(|0\rangle+|1\rangle)=\begin{pmatrix}0\\1\end{pmatrix}\otimes\begin{pmatrix}1\\1\end{pmatrix}=\begin{pmatrix}0\\0\\1\\1\end{pmatrix}=|10\rangle+|11\rangle

For scalar multiplication, take c=2c=2, ∣ϕ⟩=∣0⟩|\phi\rangle=|0\rangle, ∣ψ⟩=∣1⟩|\psi\rangle=|1\rangle:

∣0⟩⊗(2 ∣1⟩)=(10)⊗(02)=(0200)=2 ∣01⟩|0\rangle\otimes(2\,|1\rangle)=\begin{pmatrix}1\\0\end{pmatrix}\otimes\begin{pmatrix}0\\2\end{pmatrix}=\begin{pmatrix}0\\2\\0\\0\end{pmatrix}=2\,|01\rangle

Multiple systems (multilinearity)

Tensor products generalize to three or more systems. If ∣ϕ1⟩,…,∣ϕn⟩|\phi_1\rangle,\ldots,|\phi_n\rangle are vectors, their tensor product ∣ψ⟩=∣ϕ1⟩⊗⋯⊗∣ϕn⟩|\psi\rangle=|\phi_1\rangle\otimes\cdots\otimes|\phi_n\rangle is defined by the equation ⟨a1⋯an∣ψ⟩=⟨a1∣ϕ1⟩⋯⟨an∣ϕn⟩\langle a_1\cdots a_n|\psi\rangle=\langle a_1|\phi_1\rangle\cdots\langle a_n|\phi_n\rangle.

The bra ⟨a1⋯an∣\langle a_1\cdots a_n| is shorthand for ⟨a1∣⊗⟨a2∣⊗⋯⊗⟨an∣\langle a_1|\otimes\langle a_2|\otimes\cdots\otimes\langle a_n| (tensor symbols are omitted for concise notation), so ⟨a1⋯an∣ψ⟩\langle a_1\cdots a_n|\psi\rangle means applying that product bra to the tensor-product state. The equation says the larger inner product splits into matching ordinary overlaps:

(⟨a1∣⊗⋯⊗⟨an∣)(∣ϕ1⟩⊗⋯⊗∣ϕn⟩)=⟨a1∣ϕ1⟩⋯⟨an∣ϕn⟩(\langle a_1|\otimes\cdots\otimes\langle a_n|)(|\phi_1\rangle\otimes\cdots\otimes|\phi_n\rangle)=\langle a_1|\phi_1\rangle\cdots\langle a_n|\phi_n\rangle

For example, take n=3n=3 with ∣ϕ1⟩=∣0⟩, ∣ϕ2⟩=35 ∣0⟩+45 ∣1⟩, ∣ϕ3⟩=∣1⟩|\phi_1\rangle=|0\rangle,\ |\phi_2\rangle=\tfrac{3}{5}\,|0\rangle+\tfrac{4}{5}\,|1\rangle,\ |\phi_3\rangle=|1\rangle and let ∣ψ⟩=∣ϕ1⟩⊗∣ϕ2⟩⊗∣ϕ3⟩|\psi\rangle=|\phi_1\rangle\otimes|\phi_2\rangle\otimes|\phi_3\rangle. To read off the amplitude of the specific outcome (a1,a2,a3)=(0,1,1)(a_1,a_2,a_3)=(0,1,1), apply the formula directly — no need to expand the full tensor product first:

⟨011∣ψ⟩=⟨0∣ϕ1⟩⋅⟨1∣ϕ2⟩⋅⟨1∣ϕ3⟩=⟨0∣0⟩⋅⟨1∣(35∣0⟩+45∣1⟩)⋅⟨1∣1⟩=1⋅45⋅1=45\langle \textcolor{blue}{0}\textcolor{violet}{1}\textcolor{orange}{1}|\psi\rangle=\textcolor{blue}{\langle 0|\phi_1\rangle}\cdot\textcolor{violet}{\langle 1|\phi_2\rangle}\cdot\textcolor{orange}{\langle 1|\phi_3\rangle}=\textcolor{blue}{\langle 0|0\rangle}\cdot\textcolor{violet}{\langle 1|\bigl(\tfrac{3}{5}|0\rangle+\tfrac{4}{5}|1\rangle\bigr)}\cdot\textcolor{orange}{\langle 1|1\rangle}=\textcolor{blue}{1}\cdot\textcolor{violet}{\tfrac{4}{5}}\cdot\textcolor{orange}{1}=\tfrac{4}{5}

The same value can be found by expanding the tensor product first, but that is a bit more work because we build the combined state and then apply the bra:

∣ψ⟩=∣0⟩⊗(35∣0⟩+45∣1⟩)⊗∣1⟩=35 ∣001⟩+45 ∣011⟩⟨011∣ψ⟩=⟨011∣(35 ∣001⟩+45 ∣011⟩)=35⋅0+45⋅1=45\begin{aligned}|\psi\rangle&=|0\rangle\otimes\bigl(\tfrac{3}{5}|0\rangle+\tfrac{4}{5}|1\rangle\bigr)\otimes|1\rangle=\textcolor{teal}{\tfrac{3}{5}}\,|001\rangle+\textcolor{violet}{\tfrac{4}{5}}\,|011\rangle\\\langle \textcolor{blue}{0}\textcolor{violet}{1}\textcolor{orange}{1}|\psi\rangle&=\langle \textcolor{blue}{0}\textcolor{violet}{1}\textcolor{orange}{1}|\bigl(\textcolor{teal}{\tfrac{3}{5}}\,|001\rangle+\textcolor{violet}{\tfrac{4}{5}}\,|011\rangle\bigr)=\textcolor{teal}{\tfrac{3}{5}}\cdot0+\textcolor{violet}{\tfrac{4}{5}}\cdot1=\tfrac{4}{5}\end{aligned}

Any other outcome follows the same pattern. For instance, (a1,a2,a3)=(0,0,1)(a_1,a_2,a_3)=(0,0,1):

⟨001∣ψ⟩=⟨0∣ϕ1⟩⋅⟨0∣ϕ2⟩⋅⟨1∣ϕ3⟩=1⋅35⋅1=35\langle 001|\psi\rangle=\langle 0|\phi_1\rangle\cdot\langle 0|\phi_2\rangle\cdot\langle 1|\phi_3\rangle=1\cdot\tfrac{3}{5}\cdot1=\tfrac{3}{5}

The direct formula is the cleanest way to get one amplitude, but the recursive view is useful when we want the whole combined vector. It peels off the last factor:

∣ϕ1⟩⊗⋯⊗∣ϕn⟩=(∣ϕ1⟩⊗⋯⊗∣ϕn−1⟩)⊗∣ϕn⟩|\phi_1\rangle\otimes\cdots\otimes|\phi_n\rangle=\bigl(|\phi_1\rangle\otimes\cdots\otimes|\phi_{n-1}\rangle\bigr)\otimes|\phi_n\rangle

Using the same example, keep the two coefficients from ∣ϕ2⟩|\phi_2\rangle visible while peeling off ∣ϕ3⟩=∣1⟩|\phi_3\rangle=|1\rangle:

∣ψ⟩=(∣ϕ1⟩⊗∣ϕ2⟩)⊗∣ϕ3⟩=((10)⊗(3545))⊗(01)=(354500)⊗(01)=(0350450000)=35 ∣001⟩+45 ∣011⟩\begin{aligned} |\psi\rangle &=\bigl(|\phi_1\rangle\otimes|\phi_2\rangle\bigr)\otimes\textcolor{orange}{|\phi_3\rangle}\\ &=\left(\begin{pmatrix}1\\0\end{pmatrix}\otimes\begin{pmatrix}\textcolor{teal}{\tfrac{3}{5}}\\\textcolor{violet}{\tfrac{4}{5}}\end{pmatrix}\right)\otimes\textcolor{orange}{\begin{pmatrix}0\\1\end{pmatrix}}\\ &=\begin{pmatrix}\textcolor{teal}{\tfrac{3}{5}}\\\textcolor{violet}{\tfrac{4}{5}}\\0\\0\end{pmatrix}\otimes\textcolor{orange}{\begin{pmatrix}0\\1\end{pmatrix}}\\ &=\begin{pmatrix}0\\\textcolor{teal}{\tfrac{3}{5}}\\0\\\textcolor{violet}{\tfrac{4}{5}}\\0\\0\\0\\0\end{pmatrix}\\ &=\textcolor{teal}{\tfrac{3}{5}}\,|001\rangle+\textcolor{violet}{\tfrac{4}{5}}\,|011\rangle \end{aligned}

Measurement of probabilistic states

Consider the compound system (X,Y)(X,Y) in the probabilistic state: 12∣00⟩+12∣11⟩\tfrac{1}{2}|00\rangle+\tfrac{1}{2}|11\rangle

Measuring all systems

Sampling machine measuring X and Y togethersample (X,Y)XY00|00>

Measuring the entire compound system at once is equivalent to measuring each subsystem independently — provided all systems are measured. The measurement produces a single combined outcome drawn from the joint probability distribution.

For this state the two possible combined outcomes are ∣00⟩|00\rangle and ∣11⟩|11\rangle, each with probability 12\tfrac{1}{2}. Each individual measurement still happens according to those probabilities, but the whole state is measured together — producing exactly one combined result at a time: either ∣00⟩|00\rangle or ∣11⟩|11\rangle.

Measuring some systems

Suppose only a subset of systems is measured — for example, only XX from (X,Y)(X,Y). Then, to find the probability that XX equals some value aa, add up probabilities of all outcomes where X=aX=a:

Pr⁡(X=a)=∑b∈ΓPr⁡((X,Y)=(a,b))\Pr(X=a)=\sum_{b\in\Gamma}\Pr\bigl((X,Y)=(a,b)\bigr)

For the probabilistic state 12∣00⟩+12∣11⟩\tfrac{1}{2}|00\rangle+\tfrac{1}{2}|11\rangle, the table below expands the Dirac notation into the four possible joint outcomes of (X,Y)(X,Y) and their probabilities.

Joint probability table for measuring X from the compound system X Y
XYPr
0012\tfrac{1}{2}
0100
1000
1112\tfrac{1}{2}
Pr⁡(X=0)=Pr⁡((X,Y)=(0,0))+Pr⁡((X,Y)=(0,1))=12+0=12\Pr(X=0)=\Pr\bigl((X,Y)=(0,0)\bigr)+\Pr\bigl((X,Y)=(0,1)\bigr)=\tfrac{1}{2}+0=\tfrac{1}{2}
Pr⁡(X=1)=Pr⁡((X,Y)=(1,0))+Pr⁡((X,Y)=(1,1))=0+12=12\Pr(X=1)=\Pr\bigl((X,Y)=(1,0)\bigr)+\Pr\bigl((X,Y)=(1,1)\bigr)=0+\tfrac{1}{2}=\tfrac{1}{2}

But what happens to our knowledge of YY?

After measuring XX and getting result aa, there can still be uncertainty about the state of YY. If we already measured XX and got aa, what is the probability that Y=bY=b?

Pr⁡(Y=b∣X=a)=Pr⁡((X,Y)=(a,b))Pr⁡(X=a)←  Probability that both X=a and Y=b happen←  Probability that X=a happened\Pr(Y=b\mid X=a)= \begin{aligned} &\frac{\Pr\bigl((X,Y)=(a,b)\bigr)}{\Pr(X=a)} \quad \begin{array}{l} \leftarrow\;\text{Probability that both }X=a\text{ and }Y=b\text{ happen}\\[-1pt] \leftarrow\;\text{Probability that }X=a\text{ happened} \end{array} \end{aligned}

Think of the division as rescaling the probabilities so they add up to 11 again after we focus on only one situation. For example, once we learn that X=1X=1, we ignore all outcomes where X=0X=0. The probabilities of the remaining outcomes no longer add up to 11, because part of the original probability space was removed.

To turn the remaining outcomes into a valid probability distribution again, we divide each remaining probability by the total probability of X=1X=1. This process is called normalization. So the division means: out of the world where X=1X=1 happened, how likely is each remaining outcome?

If we measured X=0X=0

Pr⁡(Y=0∣X=0)=Pr⁡((X,Y)=(0,0))Pr⁡(X=0)=1212=1Pr⁡(Y=1∣X=0)=Pr⁡((X,Y)=(0,1))Pr⁡(X=0)=012=0\begin{aligned} \Pr(Y=0\mid X=0)&=\frac{\Pr\bigl((X,Y)=(0,0)\bigr)}{\Pr(X=0)}=\frac{\tfrac{1}{2}}{\tfrac{1}{2}}=1\\ \Pr(Y=1\mid X=0)&=\frac{\Pr\bigl((X,Y)=(0,1)\bigr)}{\Pr(X=0)}=\frac{0}{\tfrac{1}{2}}=0 \end{aligned}
Measuring X=0: the (1,1) outcomes are greyed out, then (0,0) stretches to fill the bar(X=1,Y=1)Pr=½(X=1,Y=1)Pr=½(X=0,Y=0)Pr=½(X=0,Y=0)Pr=101

If we measured X=1X=1

Pr⁡(Y=0∣X=1)=Pr⁡((X,Y)=(1,0))Pr⁡(X=1)=012=0Pr⁡(Y=1∣X=1)=Pr⁡((X,Y)=(1,1))Pr⁡(X=1)=1212=1\begin{aligned} \Pr(Y=0\mid X=1)&=\frac{\Pr\bigl((X,Y)=(1,0)\bigr)}{\Pr(X=1)}=\frac{0}{\tfrac{1}{2}}=0\\ \Pr(Y=1\mid X=1)&=\frac{\Pr\bigl((X,Y)=(1,1)\bigr)}{\Pr(X=1)}=\frac{\tfrac{1}{2}}{\tfrac{1}{2}}=1 \end{aligned}
Measuring X=1: the (0,0) outcomes are greyed out, then (1,1) stretches to fill the bar(X=0,Y=0)Pr=½(X=0,Y=0)Pr=½(X=1,Y=1)Pr=½(X=1,Y=1)Pr=101

Measuring one system, in Dirac notation

Everything above can be expressed directly on the state vector:

∑(a,b)∈Σ×Γpab ∣ab⟩=∑(a,b)∈Σ×Γpab ∣a⟩⊗∣b⟩=∑a∈Σ∣a⟩⊗(∑b∈Γpab ∣b⟩)\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|ab\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|a\rangle\otimes|b\rangle=\sum_{a\in\Sigma}|a\rangle\otimes\Bigl(\sum_{b\in\Gamma}p_{ab}\,|b\rangle\Bigr)

Reading the chain left to right:

  • ∑(a,b)∈Σ×Γpab ∣ab⟩\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|ab\rangle — a general probabilistic state of (X,Y)(X,Y): the sum over all possible joint outcomes (a,b)(a,b), each weighted by its probability pabp_{ab}.
  • ∑pab ∣a⟩⊗∣b⟩\sum p_{ab}\,|a\rangle\otimes|b\rangle — split every compound basis state into a tensor product, ∣ab⟩=∣a⟩⊗∣b⟩|ab\rangle=|a\rangle\otimes|b\rangle, separating the XX part from the YY part.
  • ∑a∈Σ∣a⟩⊗(∑b∈Γpab ∣b⟩)\sum_{a\in\Sigma}|a\rangle\otimes\bigl(\sum_{b\in\Gamma}p_{ab}\,|b\rangle\bigr)
    — use bilinearity to pull the shared ∣a⟩|a\rangle out of the inner sum. This groups the terms by the value of XX: one branch per value aa, with all of YY's weight collected inside the parentheses.

For example, for the probabilistic state 12∣00⟩+12∣11⟩\tfrac{1}{2}|00\rangle+\tfrac{1}{2}|11\rangle, we can split the compound state ∣00⟩|00\rangle into ∣0⟩⊗∣0⟩|0\rangle\otimes|0\rangle, and ∣11⟩|11\rangle into ∣1⟩⊗∣1⟩|1\rangle\otimes|1\rangle, and group by XX:

12∣00⟩+12∣11⟩=12(∣0⟩⊗∣0⟩)+12(∣1⟩⊗∣1⟩)=\tfrac{1}{2}|00\rangle+\tfrac{1}{2}|11\rangle=\tfrac{1}{2}\bigl(|0\rangle\otimes|0\rangle\bigr)+\tfrac{1}{2}\bigl(|1\rangle\otimes|1\rangle\bigr)=
∣0⟩|0\rangle↑if X=|0⟩
⊗\otimes
(12∣0⟩)\bigl(\tfrac{1}{2}|0\rangle\bigr)↑Y has this distribution
++
∣1⟩|1\rangle↑if X=|1⟩
⊗\otimes
(12∣1⟩)\bigl(\tfrac{1}{2}|1\rangle\bigr)↑Y has this distribution

To get the probability of measuring X=aX=a, add the probabilities of all states that start with aa:

Pr⁡(X=a)=∑b∈Γpab\Pr(X=a)=\sum_{b\in\Gamma}p_{ab}

Where:

  • pabp_{ab} — the probability of the joint outcome (a,b)(a,b). The first index aa is the value of XX, the second index bb is the value of YY — so p01p_{01} means Pr⁡((X,Y)=(0,1))\Pr\bigl((X,Y)=(0,1)\bigr).
  • ∑b∈Γ\sum_{b\in\Gamma} — keep XX fixed at aa and sweep over every possible YY; that is, add up every outcome that starts with aa.

After measuring X=aX=a, all branches with other values of XX disappear. The remaining state of YY is the surviving branch, divided by its total weight so the coefficients add back up to 11:

∑b∈Γpab ∣b⟩Pr⁡(X=a)wherePr⁡(X=a)=∑c∈Γpac\frac{\sum_{b\in\Gamma}p_{ab}\,|b\rangle}{\Pr(X=a)}\qquad\text{where}\qquad\Pr(X=a)=\sum_{c\in\Gamma}p_{ac}

The index in the normalizing sum is just a dummy variable: writing cc instead of bb avoids clashing with the bb in the numerator, but both range over all of Γ\Gamma. What it really collects is every outcome whose first coordinate is aa.

Example

Take the probabilistic state of (X,Y)(X,Y):

112 ∣00⟩+14 ∣01⟩+13 ∣10⟩+13 ∣11⟩\tfrac{1}{12}\,|00\rangle+\tfrac{1}{4}\,|01\rangle+\tfrac{1}{3}\,|10\rangle+\tfrac{1}{3}\,|11\rangle

We measure only XX (the first bit). Grouping by XX as above:

∣0⟩⊗(112 ∣0⟩+14 ∣1⟩)+∣1⟩⊗(13 ∣0⟩+13 ∣1⟩)|0\rangle\otimes\Bigl(\tfrac{1}{12}\,|0\rangle+\tfrac{1}{4}\,|1\rangle\Bigr)+|1\rangle\otimes\Bigl(\tfrac{1}{3}\,|0\rangle+\tfrac{1}{3}\,|1\rangle\Bigr)

If we measured X=0X=0

Pr⁡(X=0)=112+14=13\Pr(X=0)=\tfrac{1}{12}+\tfrac{1}{4}=\tfrac{1}{3}

The probabilistic state of YY becomes (normalize by 13\tfrac{1}{3}):

112 ∣0⟩+14 ∣1⟩13=14 ∣0⟩+34 ∣1⟩\frac{\tfrac{1}{12}\,|0\rangle+\tfrac{1}{4}\,|1\rangle}{\tfrac{1}{3}}=\tfrac{1}{4}\,|0\rangle+\tfrac{3}{4}\,|1\rangle

If we measured X=1X=1

Pr⁡(X=1)=13+13=23\Pr(X=1)=\tfrac{1}{3}+\tfrac{1}{3}=\tfrac{2}{3}

The probabilistic state of YY becomes (normalize by 23\tfrac{2}{3}):

13 ∣0⟩+13 ∣1⟩23=12 ∣0⟩+12 ∣1⟩\frac{\tfrac{1}{3}\,|0\rangle+\tfrac{1}{3}\,|1\rangle}{\tfrac{2}{3}}=\tfrac{1}{2}\,|0\rangle+\tfrac{1}{2}\,|1\rangle

Measuring some: Y

Just like we grouped by XX, we can apply the same principle and group by YY — regroup the same sum by ∣b⟩|b\rangle instead of ∣a⟩|a\rangle:

∑(a,b)∈Σ×Γpab ∣ab⟩=∑(a,b)∈Σ×Γpab ∣a⟩⊗∣b⟩=∑b∈Γ\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|ab\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|a\rangle\otimes|b\rangle=\sum_{b\in\Gamma}
(∑a∈Σpab ∣a⟩)\Bigl(\sum_{a\in\Sigma}p_{ab}\,|a\rangle\Bigr)↑X has this distribution
⊗\otimes
∣b⟩|b\rangle↑if Y=b

To get the probability of measuring Y=bY=b, add the probabilities of all states that end with bb:

Pr⁡(Y=b)=∑a∈Σpab\Pr(Y=b)=\sum_{a\in\Sigma}p_{ab}

After measuring Y=bY=b, all branches with other values of YY disappear, and the remaining state of XX is that branch normalized by its total weight:

∑a∈Σpab ∣a⟩Pr⁡(Y=b)wherePr⁡(Y=b)=∑c∈Σpcb\frac{\sum_{a\in\Sigma}p_{ab}\,|a\rangle}{\Pr(Y=b)}\qquad\text{where}\qquad\Pr(Y=b)=\sum_{c\in\Sigma}p_{cb}

Example

Take the same state, but this time measure only YY (the second bit):

112 ∣00⟩+14 ∣01⟩+13 ∣10⟩+13 ∣11⟩\tfrac{1}{12}\,|00\rangle+\tfrac{1}{4}\,|01\rangle+\tfrac{1}{3}\,|10\rangle+\tfrac{1}{3}\,|11\rangle

Grouping by YY this time:

(112 ∣0⟩+13 ∣1⟩)⊗∣0⟩+(14 ∣0⟩+13 ∣1⟩)⊗∣1⟩\Bigl(\tfrac{1}{12}\,|0\rangle+\tfrac{1}{3}\,|1\rangle\Bigr)\otimes|0\rangle+\Bigl(\tfrac{1}{4}\,|0\rangle+\tfrac{1}{3}\,|1\rangle\Bigr)\otimes|1\rangle

If we measured Y=0Y=0

Pr⁡(Y=0)=112+13=512\Pr(Y=0)=\tfrac{1}{12}+\tfrac{1}{3}=\tfrac{5}{12}

The probabilistic state of XX becomes (normalize by 512\tfrac{5}{12}):

112 ∣0⟩+13 ∣1⟩512=15 ∣0⟩+45 ∣1⟩\frac{\tfrac{1}{12}\,|0\rangle+\tfrac{1}{3}\,|1\rangle}{\tfrac{5}{12}}=\tfrac{1}{5}\,|0\rangle+\tfrac{4}{5}\,|1\rangle

If we measured Y=1Y=1

Pr⁡(Y=1)=14+13=712\Pr(Y=1)=\tfrac{1}{4}+\tfrac{1}{3}=\tfrac{7}{12}

The probabilistic state of XX becomes (normalize by 712\tfrac{7}{12}):

14 ∣0⟩+13 ∣1⟩712=37 ∣0⟩+47 ∣1⟩\frac{\tfrac{1}{4}\,|0\rangle+\tfrac{1}{3}\,|1\rangle}{\tfrac{7}{12}}=\tfrac{3}{7}\,|0\rangle+\tfrac{4}{7}\,|1\rangle

Note: for classical, probabilistic systems these conditional probabilities are easier to handle with simple conditional-probability tables. But later, with quantum states, those simple tables stop working — so Dirac notation it is (I am starting to get used to it, but I understand the struggle, deeply!).

Operations on probabilistic states

Probabilistic operations on compound systems, just like for individual systems, are represented by stochastic matrices. But this time the matrices have rows and columns corresponding to the Cartesian product of the individual systems' classical state sets.

A deterministic example: controlled-NOT

Take a controlled-NOT operation on two bits (X,Y)(X,Y):

  • if x=1x=1 ⇒ apply NOT to yy;
  • if x=0x=0 ⇒ do nothing.

Here XX is the control bit and YY is the target bit. It is deterministic — each input state maps to exactly one output state:

Input → output

∣00⟩↦∣00⟩∣01⟩↦∣01⟩∣10⟩↦∣11⟩∣11⟩↦∣10⟩\begin{aligned}|00\rangle&\mapsto|00\rangle\\|01\rangle&\mapsto|01\rangle\\|10\rangle&\mapsto|11\rangle\\|11\rangle&\mapsto|10\rangle\end{aligned}

Matrix representation

M=(1000010000010010)M=\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}

For example, applying the matrix operation to ∣10⟩|10\rangle flips the target, giving ∣11⟩|11\rangle:

M ∣10⟩=(1000010000010010)(0010)=(0001)=∣11⟩M\,|10\rangle=\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}\begin{pmatrix}0\\0\\1\\0\end{pmatrix}=\begin{pmatrix}0\\0\\0\\1\end{pmatrix}=|11\rangle

Nothing special is going on — the operation is just a matrix acting on the probability vector.

A probabilistic example

A probabilistic operation makes a random choice. For example:

  • with probability 12\tfrac{1}{2} ⇒ set y=xy=x;
  • with probability 12\tfrac{1}{2} ⇒ set x=yx=y.

Each branch is itself a deterministic stochastic matrix. Take the input ∣10⟩|10\rangle to see how both behave:

Set y=xy=x

(1100000000000011)(0010)=(0001)=∣11⟩\begin{pmatrix}1&1&0&0\\0&0&0&0\\0&0&0&0\\0&0&1&1\end{pmatrix}\begin{pmatrix}0\\0\\1\\0\end{pmatrix}=\begin{pmatrix}0\\0\\0\\1\end{pmatrix}=|11\rangle

Set x=yx=y

(1010000000000101)(0010)=(1000)=∣00⟩\begin{pmatrix}1&0&1&0\\0&0&0&0\\0&0&0&0\\0&1&0&1\end{pmatrix}\begin{pmatrix}0\\0\\1\\0\end{pmatrix}=\begin{pmatrix}1\\0\\0\\0\end{pmatrix}=|00\rangle

The overall probabilistic operation is the average of these two matrices, weighted by the probabilities 12\tfrac{1}{2}:

(11212000000000012121)=12(1100000000000011)+12(1010000000000101)\begin{pmatrix}1&\tfrac12&\tfrac12&0\\0&0&0&0\\0&0&0&0\\0&\tfrac12&\tfrac12&1\end{pmatrix}=\tfrac12\begin{pmatrix}1&1&0&0\\0&0&0&0\\0&0&0&0\\0&0&1&1\end{pmatrix}+\tfrac12\begin{pmatrix}1&0&1&0\\0&0&0&0\\0&0&0&0\\0&1&0&1\end{pmatrix}

The resulting matrix does not describe one concrete execution — it describes the distribution over outcomes after the random choice.

Simultaneous operations on a compound system

Suppose instead that two probabilistic operations act on separate systems — MM on XX and NN on YY, each its own stochastic matrix on its own probability vector. If we perform both at the same time, how do we describe their combined effect on the (X,Y)(X,Y) compound system?

Tensor product of matrices

Performing MM on XX and NN on YY simultaneously is described by a single matrix on the compound system — the tensor product M⊗NM\otimes N. Writing each operation in Dirac notation,

M=∑a,b∈Σαab ∣a⟩⟨b∣N=∑c,d∈Γβcd ∣c⟩⟨d∣M=\sum_{a,b\in\Sigma}\alpha_{ab}\,|a\rangle\langle b|\qquad N=\sum_{c,d\in\Gamma}\beta_{cd}\,|c\rangle\langle d|

the tensor product pairs every term of one with every term of the other, multiplying their coefficients:

M⊗N=∑a,b∈Σ  ∑c,d∈Γαab βcd ∣ac⟩⟨bd∣M\otimes N=\sum_{a,b\in\Sigma}\;\sum_{c,d\in\Gamma}\alpha_{ab}\,\beta_{cd}\,|ac\rangle\langle bd|
  • αab\alpha_{ab} — the entry of MM in row aa, column bb
  • βcd\beta_{cd} — the entry of NN in row cc, column dd
  • ∣ac⟩⟨bd∣|ac\rangle\langle bd| — the compound outer product, using the identity
    ∣a⟩⟨b∣⊗∣c⟩⟨d∣=∣ac⟩⟨bd∣|a\rangle\langle b|\otimes|c\rangle\langle d|=|ac\rangle\langle bd|

Example: bit‑flip ⊗\otimes identity

First rewrite each in Dirac notation by reading off its entries: the entry in row aa, column bb is the coefficient of ∣a⟩⟨b∣|a\rangle\langle b|:

MM — bit‑flip on XX

M=(0110)α00=0  ⟶  0 ∣0⟩⟨0∣α01=1  ⟶  1 ∣0⟩⟨1∣α10=1  ⟶  1 ∣1⟩⟨0∣α11=0  ⟶  0 ∣1⟩⟨1∣M=\begin{pmatrix}\textcolor{blue}{0}&\textcolor{orange}{1}\\[2pt]\textcolor{teal}{1}&\textcolor{violet}{0}\end{pmatrix}\qquad\begin{array}{l}\textcolor{blue}{\alpha_{00}=0}\;\longrightarrow\;\textcolor{blue}{0\,|0\rangle\langle 0|}\\[4pt]\textcolor{orange}{\alpha_{01}=1}\;\longrightarrow\;\textcolor{orange}{1\,|0\rangle\langle 1|}\\[4pt]\textcolor{teal}{\alpha_{10}=1}\;\longrightarrow\;\textcolor{teal}{1\,|1\rangle\langle 0|}\\[4pt]\textcolor{violet}{\alpha_{11}=0}\;\longrightarrow\;\textcolor{violet}{0\,|1\rangle\langle 1|}\end{array}

Written out and simplified:

M=0 ∣0⟩⟨0∣+1 ∣0⟩⟨1∣+1 ∣1⟩⟨0∣+0 ∣1⟩⟨1∣=∣0⟩⟨1∣+∣1⟩⟨0∣\begin{aligned}M&=\textcolor{blue}{0}\,|0\rangle\langle 0|+\textcolor{orange}{1}\,|0\rangle\langle 1|+\textcolor{teal}{1}\,|1\rangle\langle 0|+\textcolor{violet}{0}\,|1\rangle\langle 1|\\[4pt]&=\textcolor{orange}{|0\rangle\langle 1|}+\textcolor{teal}{|1\rangle\langle 0|}\end{aligned}

NN — identity on YY

N=(1001)β00=1  ⟶  1 ∣0⟩⟨0∣β01=0  ⟶  0 ∣0⟩⟨1∣β10=0  ⟶  0 ∣1⟩⟨0∣β11=1  ⟶  1 ∣1⟩⟨1∣N=\begin{pmatrix}\textcolor{#dc2626}{1}&\textcolor{#15803d}{0}\\[2pt]\textcolor{#b45309}{0}&\textcolor{#db2777}{1}\end{pmatrix}\qquad\begin{array}{l}\textcolor{#dc2626}{\beta_{00}=1}\;\longrightarrow\;\textcolor{#dc2626}{1\,|0\rangle\langle 0|}\\[4pt]\textcolor{#15803d}{\beta_{01}=0}\;\longrightarrow\;\textcolor{#15803d}{0\,|0\rangle\langle 1|}\\[4pt]\textcolor{#b45309}{\beta_{10}=0}\;\longrightarrow\;\textcolor{#b45309}{0\,|1\rangle\langle 0|}\\[4pt]\textcolor{#db2777}{\beta_{11}=1}\;\longrightarrow\;\textcolor{#db2777}{1\,|1\rangle\langle 1|}\end{array}

Written out and simplified:

N=1 ∣0⟩⟨0∣+0 ∣0⟩⟨1∣+0 ∣1⟩⟨0∣+1 ∣1⟩⟨1∣=∣0⟩⟨0∣+∣1⟩⟨1∣\begin{aligned}N&=\textcolor{#dc2626}{1}\,|0\rangle\langle 0|+\textcolor{#15803d}{0}\,|0\rangle\langle 1|+\textcolor{#b45309}{0}\,|1\rangle\langle 0|+\textcolor{#db2777}{1}\,|1\rangle\langle 1|\\[4pt]&=\textcolor{#dc2626}{|0\rangle\langle 0|}+\textcolor{#db2777}{|1\rangle\langle 1|}\end{aligned}

Method 1 — Kronecker blocks

Replace each entry of MM with the block Mij NM_{ij}\,N:

M⊗N=(0110)⊗(1001)=(0⋅N1⋅N1⋅N0⋅N)=M\otimes N=\begin{pmatrix}\textcolor{blue}{0}&\textcolor{orange}{1}\\\textcolor{teal}{1}&\textcolor{violet}{0}\end{pmatrix}\otimes\begin{pmatrix}1&0\\0&1\end{pmatrix}=\begin{pmatrix}\textcolor{blue}{0\cdot N}&\textcolor{orange}{1\cdot N}\\\textcolor{teal}{1\cdot N}&\textcolor{violet}{0\cdot N}\end{pmatrix}=(0010000110000100)\begin{pmatrix}0&0&1&0\\0&0&0&1\\1&0&0&0\\0&1&0&0\end{pmatrix}

Method 2 — Dirac expansion

Distribute (bilinearity), collapse each term with the identity, then keep the equality chain going by turning each outer product into its matrix:

∣a⟩⟨b∣⊗∣c⟩⟨d∣=∣ac⟩⟨bd∣|a\rangle\langle b|\otimes|c\rangle\langle d|=|ac\rangle\langle bd|
M⊗N=(∣0⟩⟨1∣+∣1⟩⟨0∣)⊗(∣0⟩⟨0∣+∣1⟩⟨1∣)=∣0⟩⟨1∣⊗∣0⟩⟨0∣+∣0⟩⟨1∣⊗∣1⟩⟨1∣+∣1⟩⟨0∣⊗∣0⟩⟨0∣+∣1⟩⟨0∣⊗∣1⟩⟨1∣=∣00⟩⟨10∣+∣01⟩⟨11∣+∣10⟩⟨00∣+∣11⟩⟨01∣=(0010000000000000)+(0000000100000000)+(0000000010000000)+(0000000000000100)=(0010000110000100)\begin{aligned} M\otimes N &=\bigl(|0\rangle\langle 1|+|1\rangle\langle 0|\bigr)\otimes\bigl(|0\rangle\langle 0|+|1\rangle\langle 1|\bigr)\\[4pt] &=|0\rangle\langle 1|\otimes|0\rangle\langle 0|+|0\rangle\langle 1|\otimes|1\rangle\langle 1|+|1\rangle\langle 0|\otimes|0\rangle\langle 0|+|1\rangle\langle 0|\otimes|1\rangle\langle 1|\\[4pt] &=|00\rangle\langle 10|+|01\rangle\langle 11|+|10\rangle\langle 00|+|11\rangle\langle 01|\\[4pt] &=\begin{pmatrix}0&0&1&0\\0&0&0&0\\0&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&1\\0&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&0\\1&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&0\\0&0&0&0\\0&1&0&0\end{pmatrix}\\[4pt] &=\begin{pmatrix}0&0&1&0\\0&0&0&1\\1&0&0&0\\0&1&0&0\end{pmatrix} \end{aligned}

Key idea of Dirac notation

A sandwich ⟨a∣M∣b⟩\langle a|M|b\rangle is just the single entry of MM at row aa, column bb. The bra ⟨a∣\langle a| is the unit row vector that picks out row aa, and the ket ∣b⟩|b\rangle is the unit column vector that picks out column bb — multiplying them on either side of MM collapses the whole matrix down to one number. Indices count from 00, so ⟨0∣M∣0⟩\langle 0|M|0\rangle is the top-left entry:

M=(1234)⟨0∣M∣0⟩=(10)⏟⟨0∣(1234)(10)⏟∣0⟩=1M=\begin{pmatrix}\textcolor{#0369a1}{1}&2\\3&4\end{pmatrix}\qquad\langle 0|M|0\rangle=\underbrace{\begin{pmatrix}1&0\end{pmatrix}}_{\langle 0|}\begin{pmatrix}1&2\\3&4\end{pmatrix}\underbrace{\begin{pmatrix}1\\0\end{pmatrix}}_{|0\rangle}=\textcolor{#0369a1}{1}

Equivalent entry rule

Equivalently, the tensor product is the matrix whose compound entry is found by multiplying the matching entry from MM with the matching entry from NN:

⟨ac∣M⊗N∣bd⟩=⟨a∣M∣b⟩ ⟨c∣N∣d⟩for all a,b∈Σ and c,d∈Γ\langle ac|M\otimes N|bd\rangle=\langle a|M|b\rangle\,\langle c|N|d\rangle\qquad\text{for all }a,b\in\Sigma\text{ and }c,d\in\Gamma
  • Read ⟨ac∣M⊗N∣bd⟩\langle ac|M\otimes N|bd\rangle from right to left: ∣bd⟩|bd\rangle is the starting column, M⊗NM\otimes N is the matrix you apply, and ⟨ac∣\langle ac| is the output row you read. For example, let's calculate ⟨10∣M⊗N∣00⟩\langle 10|M\otimes N|00\rangle. Start with the input column ∣00⟩|00\rangle, apply M⊗NM\otimes N, then read the value in the output row ⟨10∣\langle 10| — the result is 11.
    columns in ∣bd⟩|bd\rangle
    ∣00⟩|00\rangle
    ∣01⟩|01\rangle
    ∣10⟩|10\rangle
    ∣11⟩|11\rangle
    ↓\downarrow
    M⊗N=M\otimes N=
    ⟨00∣\langle 00|
    ⟨01∣\langle 01|
    ⟨10∣\langle 10|
    ⟨11∣\langle 11|
    0
    0
    1
    0
    0
    0
    0
    1
    1
    0
    0
    0
    0
    1
    0
    0
    ←\leftarrowrows out ⟨ac∣=⟨10∣\langle ac|=\langle 10|
    ⟨10∣M⊗N∣00⟩=1\langle 10|M\otimes N|00\rangle=1
  • The right side asks the same question one system at a time: ⟨a∣M∣b⟩\langle a|M|b\rangle asks how much MM sends ∣b⟩|b\rangle to ∣a⟩|a\rangle, and ⟨c∣N∣d⟩\langle c|N|d\rangle asks how much NN sends ∣d⟩|d\rangle to ∣c⟩|c\rangle. For the same example, look up each factor in its own matrix — ⟨1∣M∣0⟩\langle 1|M|0\rangle is the entry of MM in row ⟨1∣\langle 1|, column ∣0⟩|0\rangle, and ⟨0∣N∣0⟩\langle 0|N|0\rangle likewise for NN — then multiply them:
    M=M=
    ∣0⟩|0\rangle
    ∣1⟩|1\rangle
    ⟨0∣\langle 0|
    ⟨1∣\langle 1|
    0
    1
    1
    0
    N=N=
    ∣0⟩|0\rangle
    ∣1⟩|1\rangle
    ⟨0∣\langle 0|
    ⟨1∣\langle 1|
    1
    0
    0
    1
    ⟨1∣M∣0⟩⋅⟨0∣N∣0⟩=1⋅1=1\langle 1|M|0\rangle\cdot\langle 0|N|0\rangle=1\cdot 1=1
  • In this example, MM flips the first bit and NN leaves the second unchanged, so M⊗NM\otimes N sends ∣00⟩|00\rangle straight to ∣10⟩|10\rangle. The resulting 11 is the amplitude of that transition: it says all of the input lands on ∣10⟩|10\rangle and none on any other basis state — the mapping is exact and deterministic (a coefficient of 00 would mean ∣00⟩|00\rangle never reaches ∣10⟩|10\rangle).

Why the entry rule holds

We want the entry of M⊗NM\otimes N at row ⟨ac∣\langle ac| and column ∣bd⟩|bd\rangle. The rule says: take the matching entry from MM, take the matching entry from NN, then multiply them.

⟨ac∣M⊗N∣bd⟩=⟨a∣M∣b⟩ ⟨c∣N∣d⟩\langle ac|M\otimes N|bd\rangle=\langle a|M|b\rangle\,\langle c|N|d\rangle

Derivation

⟨ac∣M⊗N∣bd⟩\langle ac|M\otimes N|bd\rangle

The compound entry we want.

=⟨ac∣(∑i,j∈Σαij∣i⟩⟨j∣)⊗(∑k,l∈Γβkl∣k⟩⟨l∣)∣bd⟩=\langle ac|\left(\sum_{i,j\in\Sigma}\alpha_{ij}|i\rangle\langle j|\right)\otimes\left(\sum_{k,l\in\Gamma}\beta_{kl}|k\rangle\langle l|\right)|bd\rangle

Write MM and NN as sums of weighted outer products.

=∑i,j∈Σ∑k,l∈Γαijβkl ⟨ac∣ik⟩ ⟨jl∣bd⟩=\sum_{i,j\in\Sigma}\sum_{k,l\in\Gamma}\alpha_{ij}\beta_{kl}\,\langle ac|ik\rangle\,\langle jl|bd\rangle

Linearity pulls the coefficients out front, and ∣i⟩⟨j∣⊗∣k⟩⟨l∣|i\rangle\langle j|\otimes|k\rangle\langle l| becomes ∣ik⟩⟨jl∣|ik\rangle\langle jl|.

=∑i,j∈Σ∑k,l∈Γαijβkl ⟨a∣i⟩ ⟨c∣k⟩ ⟨j∣b⟩ ⟨l∣d⟩=\sum_{i,j\in\Sigma}\sum_{k,l\in\Gamma}\alpha_{ij}\beta_{kl}\,\langle a|i\rangle\,\langle c|k\rangle\,\langle j|b\rangle\,\langle l|d\rangle

Each compound overlap breaks into one bracket per system.

=∑i,j∈Σ∑k,l∈Γαijβkl δaiδckδjbδld=\sum_{i,j\in\Sigma}\sum_{k,l\in\Gamma}\alpha_{ij}\beta_{kl}\,\delta_{ai}\delta_{ck}\delta_{jb}\delta_{ld}

Those brackets vanish unless i=ai=a, j=bj=b, k=ck=c, l=dl=d.

=αabβcd=\alpha_{ab}\beta_{cd}

Only that single term survives the double sum.

=⟨a∣M∣b⟩ ⟨c∣N∣d⟩=\langle a|M|b\rangle\,\langle c|N|d\rangle

Exactly the per-system entry rule we set out to prove.

Selector picture

Focus on the αij\alpha_{ij} part first. The two brackets are two tests on the same cell:

  • ⟨a∣i⟩\langle a|i\rangle asks: is this cell in row i=ai=a?
  • ⟨j∣b⟩\langle j|b\rangle asks: is this cell in column j=bj=b?

A cell must pass both tests. If it fails either one, it gets multiplied by 00.

αij⟨a∣i⟩⟨j∣b⟩={αab,i=a and j=b0,otherwise\alpha_{ij}\langle a|i\rangle\langle j|b\rangle= \begin{cases} \alpha_{ab},&i=a\text{ and }j=b\\ 0,&\text{otherwise} \end{cases}

The βkl\beta_{kl} part does the same thing with k=ck=c and l=dl=d, so it leaves βcd\beta_{cd}.

Example: a=1a=1, b=1b=1

j=0j=0
j=1j=1
j=2j=2
i=0i=0
α00\alpha_{00}
α01\alpha_{01}
α02\alpha_{02}
i=1i=1
α10\alpha_{10}
α11\alpha_{11}
α12\alpha_{12}
i=2i=2
α20\alpha_{20}
α21\alpha_{21}
α22\alpha_{22}

Pale cells pass one test but fail the other. The highlighted cell passes both, so it is the only one left.

Equivalent action on product states

Equivalently, M⊗NM\otimes N is the unique matrix that satisfies the equation for all vectors ∣φ⟩|\varphi\rangle and ∣ψ⟩|\psi\rangle:

(M⊗N) (∣φ⟩⊗∣ψ⟩)=(M∣φ⟩)⊗(N∣ψ⟩)(M\otimes N)\,\bigl(|\varphi\rangle\otimes|\psi\rangle\bigr)=\bigl(M|\varphi\rangle\bigr)\otimes\bigl(N|\psi\rangle\bigr)

The tensor product is defined by one rule: apply MM to the first subsystem and NN to the second. Because every state can be built from basis states by adding them together, and linear maps preserve addition, this rule completely determines the full matrix.

Explicit formula

Collecting the entry rule into a single matrix gives the explicit form of the tensor product. Each entry is the product αab βcd=⟨ac∣M⊗N∣bd⟩\alpha_{ab}\,\beta_{cd}=\langle ac|M\otimes N|bd\rangle, so for MM with entries up to αmm\alpha_{mm} and NN with entries up to βnn\beta_{nn}:

M⊗N=(α00⋯α0m⋮⋱⋮αm0⋯αmm)⊗(β00⋯β0n⋮⋱⋮βn0⋯βnn)=(α00β00⋯α00β0n⋯α0mβ00⋯α0mβ0n⋮⋱⋮⋮⋱⋮α00βn0⋯α00βnn⋯α0mβn0⋯α0mβnn⋮⋮⋱⋮⋮αm0β00⋯αm0β0n⋯αmmβ00⋯αmmβ0n⋮⋱⋮⋮⋱⋮αm0βn0⋯αm0βnn⋯αmmβn0⋯αmmβnn)\begin{aligned}M\otimes N&=\begin{pmatrix}\alpha_{00}&\cdots&\alpha_{0m}\\\vdots&\ddots&\vdots\\\alpha_{m0}&\cdots&\alpha_{mm}\end{pmatrix}\otimes\begin{pmatrix}\beta_{00}&\cdots&\beta_{0n}\\\vdots&\ddots&\vdots\\\beta_{n0}&\cdots&\beta_{nn}\end{pmatrix}\\[8pt]&=\begin{pmatrix}\alpha_{00}\beta_{00}&\cdots&\alpha_{00}\beta_{0n}&\cdots&\alpha_{0m}\beta_{00}&\cdots&\alpha_{0m}\beta_{0n}\\[2pt]\vdots&\ddots&\vdots&&\vdots&\ddots&\vdots\\[2pt]\alpha_{00}\beta_{n0}&\cdots&\alpha_{00}\beta_{nn}&\cdots&\alpha_{0m}\beta_{n0}&\cdots&\alpha_{0m}\beta_{nn}\\[2pt]\vdots&&\vdots&\ddots&\vdots&&\vdots\\[2pt]\alpha_{m0}\beta_{00}&\cdots&\alpha_{m0}\beta_{0n}&\cdots&\alpha_{mm}\beta_{00}&\cdots&\alpha_{mm}\beta_{0n}\\[2pt]\vdots&\ddots&\vdots&&\vdots&\ddots&\vdots\\[2pt]\alpha_{m0}\beta_{n0}&\cdots&\alpha_{m0}\beta_{nn}&\cdots&\alpha_{mm}\beta_{n0}&\cdots&\alpha_{mm}\beta_{nn}\end{pmatrix}\end{aligned}

Three or more matrices

Nothing about the entry rule was special to two systems. For nn factors M1⊗⋯⊗MnM_1\otimes\cdots\otimes M_n, a compound entry still splits into one bracket per system — each MkM_k is asked the same question about its own bits in isolation, and the answers multiply:

⟨a1⋯an∣ M1⊗⋯⊗Mn ∣b1⋯bn⟩=⟨a1∣M1∣b1⟩  ⟨a2∣M2∣b2⟩⋯⟨an∣Mn∣bn⟩\langle a_1\cdots a_n|\,M_1\otimes\cdots\otimes M_n\,|b_1\cdots b_n\rangle=\langle a_1|M_1|b_1\rangle\;\langle a_2|M_2|b_2\rangle\cdots\langle a_n|M_n|b_n\rangle

The product also stays compatible with ordinary matrix multiplication. Running one stack of operators after another is the same as multiplying the matrices system by system (the mixed‑product property):

(M1⊗⋯⊗Mn)(N1⊗⋯⊗Nn)=(M1N1)⊗⋯⊗(MnNn)\bigl(M_1\otimes\cdots\otimes M_n\bigr)\bigl(N_1\otimes\cdots\otimes N_n\bigr)=\bigl(M_1N_1\bigr)\otimes\cdots\otimes\bigl(M_nN_n\bigr)

Multiple systems: quantum

Quantum states

A quantum state of several systems is represented by a column vector whose indices correspond to the Cartesian product of the individual systems' classical state sets — exactly the same index set as the compound classical system, now carrying complex amplitudes instead of probabilities.

For two systems XX with state set Σ\Sigma and YY with state set Γ\Gamma, the entries are indexed by Σ×Γ\Sigma\times\Gamma. If both are bits, the four indices are:

{0,1}×{0,1}={00, 01, 10, 11}\{0,1\}\times\{0,1\}=\{00,\,01,\,10,\,11\}

So a quantum state of the two-bit system XYXY is a four-entry column vector. Written as a combination of the standard basis states ∣00⟩,∣01⟩,∣10⟩,∣11⟩|00\rangle,|01\rangle,|10\rangle,|11\rangle:

∣ψ⟩=α00 ∣00⟩+α01 ∣01⟩+α10 ∣10⟩+α11 ∣11⟩=(α00α01α10α11)|\psi\rangle=\alpha_{00}\,|00\rangle+\alpha_{01}\,|01\rangle+\alpha_{10}\,|10\rangle+\alpha_{11}\,|11\rangle=\begin{pmatrix}\alpha_{00}\\\alpha_{01}\\\alpha_{10}\\\alpha_{11}\end{pmatrix}

As with a single system, the amplitudes are complex numbers and the vector is a unit vector: the squared absolute values sum to one, and ∣αab∣2|\alpha_{ab}|^2 is the probability of measuring the pair (a,b)(a,b).

∑(a,b)∈Σ×Γ∣αab∣2=1\sum_{(a,b)\in\Sigma\times\Gamma}|\alpha_{ab}|^2=1

Definite (basis) state

∣10⟩=(0010)|10\rangle=\begin{pmatrix}0\\0\\1\\0\end{pmatrix}

XX is certainly ∣1⟩|1\rangle and YY is certainly ∣0⟩|0\rangle.

Equal superposition

∣ψ⟩=12(∣00⟩+∣01⟩+∣10⟩+∣11⟩)=12(1111)|\psi\rangle=\tfrac{1}{2}\bigl(|00\rangle+|01\rangle+|10\rangle+|11\rangle\bigr)=\tfrac{1}{2}\begin{pmatrix}1\\1\\1\\1\end{pmatrix}

All four outcomes are equally likely, each with probability (12)2=14\bigl(\tfrac{1}{2}\bigr)^2=\tfrac{1}{4}.

Biased superposition

∣ψ⟩=12 ∣00⟩+12 ∣01⟩+12 ∣11⟩=(1212012)|\psi\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{2}\,|01\rangle+\tfrac{1}{2}\,|11\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{2}\\[4pt]0\\[4pt]\tfrac{1}{2}\end{pmatrix}

Still a unit vector: 12+14+14=1\tfrac{1}{2}+\tfrac{1}{4}+\tfrac{1}{4}=1.

Entangled (Bell) state

∣ϕ+⟩=12(∣00⟩+∣11⟩)=12(1001)|\phi^{+}\rangle=\tfrac{1}{\sqrt{2}}\bigl(|00\rangle+|11\rangle\bigr)=\tfrac{1}{\sqrt{2}}\begin{pmatrix}1\\0\\0\\1\end{pmatrix}

Measuring gives ∣00⟩|00\rangle or ∣11⟩|11\rangle with equal probability and never ∣01⟩|01\rangle or ∣10⟩|10\rangle, so the two bits always come out the same. Unlike the states above, it cannot be factored into a separate state for each bit — no ∣ψX⟩⊗∣ψY⟩|\psi_X\rangle\otimes|\psi_Y\rangle equals ∣ϕ+⟩|\phi^{+}\rangle. That inseparability is what entanglement means.

In general, combining nn systems with state sets Σ1,…,Σn\Sigma_1,\ldots,\Sigma_n gives a state vector indexed by Σ1×⋯×Σn\Sigma_1\times\cdots\times\Sigma_n, so its dimension is the product ∣Σ1∣⋯∣Σn∣|\Sigma_1|\cdots|\Sigma_n| of the individual sizes — for nn qubits, 2n2^n amplitudes.

Tensor products of states

The previous states described one compound system directly. We can also build a compound state by combining states of the parts: the tensor product of two quantum state vectors is again a quantum state vector.

Let ∣ϕ⟩|\phi\rangle be a state of system XX and ∣ψ⟩|\psi\rangle a state of system YY. Their tensor product is a state of the joint system (X,Y)(X,Y):

∣ϕ⟩⊗∣ψ⟩|\phi\rangle\otimes|\psi\rangle

States of this form are called product states. They describe the two systems acting independently — each part has its own well-defined state, with no correlation between them. (Entangled states like ∣ϕ+⟩|\phi^{+}\rangle cannot be written this way.)

More generally, if ∣ψ1⟩,…,∣ψn⟩|\psi_1\rangle,\ldots,|\psi_n\rangle are states of systems X1,…,XnX_1,\ldots,X_n, then their tensor product is a product state of the whole compound system (X1,…,Xn)(X_1,\ldots,X_n):

∣ψ1⟩⊗⋯⊗∣ψn⟩|\psi_1\rangle\otimes\cdots\otimes|\psi_n\rangle

Example: a state that is not a product state

∣ψ⟩=12 ∣00⟩+12 ∣11⟩|\psi\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{\sqrt{2}}\,|11\rangle

It is a valid quantum state — a unit vector, since the squared amplitudes sum to one:

∣12∣2+∣12∣2=12+12=1\left|\tfrac{1}{\sqrt{2}}\right|^2+\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}+\tfrac{1}{2}=1

But it cannot be written as a tensor product ∣ϕ⟩⊗∣ψ⟩|\phi\rangle\otimes|\psi\rangle. Any product (a∣0⟩+b∣1⟩)⊗(c∣0⟩+d∣1⟩)(a|0\rangle+b|1\rangle)\otimes(c|0\rangle+d|1\rangle) expands to ac ∣00⟩+ad ∣01⟩+bc ∣10⟩+bd ∣11⟩ac\,|00\rangle+ad\,|01\rangle+bc\,|10\rangle+bd\,|11\rangle. Matching our state forces ad=bc=0ad=bc=0 (no ∣01⟩|01\rangle or ∣10⟩|10\rangle terms) while acac and bdbd are both nonzero — which is impossible. So the two bits are entangled, not independent.

The Bell basis

The state from the previous example, ∣ϕ+⟩|\phi^{+}\rangle, is one of four Bell states. Each differs only by which basis pairs are combined and by a sign, and together they form the Bell basis — an orthonormal basis of the two-qubit space made entirely of maximally entangled states.

∣ϕ+⟩=12 ∣00⟩+12 ∣11⟩|\phi^{+}\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{\sqrt{2}}\,|11\rangle
∣ϕ−⟩=12 ∣00⟩−12 ∣11⟩|\phi^{-}\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle-\tfrac{1}{\sqrt{2}}\,|11\rangle
∣ψ+⟩=12 ∣01⟩+12 ∣10⟩|\psi^{+}\rangle=\tfrac{1}{\sqrt{2}}\,|01\rangle+\tfrac{1}{\sqrt{2}}\,|10\rangle
∣ψ−⟩=12 ∣01⟩−12 ∣10⟩|\psi^{-}\rangle=\tfrac{1}{\sqrt{2}}\,|01\rangle-\tfrac{1}{\sqrt{2}}\,|10\rangle

None of the four can be written as a tensor product ∣ϕ⟩⊗∣ψ⟩|\phi\rangle\otimes|\psi\rangle, so each is entangled. Because they are orthonormal, any two-qubit state can be expressed as a combination of these four.

Three-qubit states

Entanglement is not limited to pairs. Two famous three-qubit states show different ways three systems can be correlated — both are unit vectors and neither is a product state.

GHZ state

∣GHZ⟩=12 ∣000⟩+12 ∣111⟩|\mathrm{GHZ}\rangle=\tfrac{1}{\sqrt{2}}\,|000\rangle+\tfrac{1}{\sqrt{2}}\,|111\rangle

All three qubits are either ∣0⟩|0\rangle or all ∣1⟩|1\rangle. Measuring any one qubit instantly fixes the other two — the three-qubit analogue of the Bell state ∣ϕ+⟩|\phi^{+}\rangle.

W state

∣W⟩=13 ∣001⟩+13 ∣010⟩+13 ∣100⟩|\mathrm{W}\rangle=\tfrac{1}{\sqrt{3}}\,|001\rangle+\tfrac{1}{\sqrt{3}}\,|010\rangle+\tfrac{1}{\sqrt{3}}\,|100\rangle

Exactly one qubit is ∣1⟩|1\rangle and its position is in superposition. Each squared amplitude is (13)2=13\bigl(\tfrac{1}{\sqrt{3}}\bigr)^2=\tfrac{1}{3}, and the three sum to one.

Measurements

Measuring a compound quantum system works exactly like measuring a single one — provided every system is measured. A standard basis measurement of the whole system returns one combined classical outcome, drawn from the squared amplitudes just as in the single-system case.

If ∣ψ⟩|\psi\rangle is a quantum state of a system (X1,…,Xn)(X_1,\ldots,X_n) and all systems are measured, then each nn-tuple

(a1,…,an)∈Σ1×⋯×Σn(a_1,\ldots,a_n)\in\Sigma_1\times\cdots\times\Sigma_n

(or string a1⋯ana_1\cdots a_n) is obtained with probability equal to the squared absolute value of its amplitude:

Pr⁡(outcome=a1⋯an)=∣⟨a1⋯an∣ψ⟩∣2\Pr(\text{outcome}=a_1\cdots a_n)=\bigl|\langle a_1\cdots a_n|\psi\rangle\bigr|^2

The inner product ⟨a1⋯an∣ψ⟩\langle a_1\cdots a_n|\psi\rangle simply picks out the amplitude sitting in front of the basis state ∣a1⋯an⟩|a_1\cdots a_n\rangle — the entry of the state vector indexed by that outcome.

Example 1

Measuring both qubits of the Bell state ∣ϕ+⟩=12 ∣00⟩+12 ∣11⟩|\phi^{+}\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{\sqrt{2}}\,|11\rangle gives 0000 or 1111 only — the two qubits always agree:

Pr⁡(00)=∣12∣2=12Pr⁡(11)=∣12∣2=12Pr⁡(01)=Pr⁡(10)=0\Pr(00)=\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}\qquad\Pr(11)=\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}\qquad\Pr(01)=\Pr(10)=0

Example 2

Subsystems need not be qubits and amplitudes may be complex. For the pair (X,Y)(X,Y) in the state 35 ∣0⟩∣♡⟩−4i5 ∣1⟩∣♠⟩\tfrac{3}{5}\,|0\rangle|{\textcolor{#dc2626}{\heartsuit}}\rangle-\tfrac{4i}{5}\,|1\rangle|{\spadesuit}\rangle:

Pr⁡(0,♡)=∣35∣2=925Pr⁡(1,♠)=∣−4i5∣2=1625\Pr(0,\textcolor{#dc2626}{\heartsuit})=\left|\tfrac{3}{5}\right|^2=\tfrac{9}{25}\qquad\Pr(1,\spadesuit)=\left|-\tfrac{4i}{5}\right|^2=\tfrac{16}{25}

The two add to 11, and the ii drops out under ∣⋅∣2|\cdot|^2 — only amplitude magnitudes matter.

Measuring some systems

What if two systems (X,Y)(X,Y) share a quantum state but we measure only XX and leave YY alone? Write the joint state in the usual form:

∣ψ⟩=∑(a,b)∈Σ×Γαab ∣ab⟩|\psi\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}\alpha_{ab}\,|ab\rangle

If both were measured, outcome (a,b)(a,b) would appear with probability ∣⟨ab∣ψ⟩∣2=∣αab∣2|\langle ab|\psi\rangle|^2=|\alpha_{ab}|^2. Measuring only XX must give the same probability for X=aX=a as summing over every YY outcome:

Pr⁡(X=a)=∑b∈Γ∣⟨ab∣ψ⟩∣2=∑b∈Γ∣αab∣2\Pr(X=a)=\sum_{b\in\Gamma}|\langle ab|\psi\rangle|^2=\sum_{b\in\Gamma}|\alpha_{ab}|^2

Just as in the probabilistic setting, the state of YY changes as a result. The branch with X=aX=a survives and must be renormalized back to a unit vector — but because these are amplitudes, not probabilities, we divide by the square root of Pr⁡(X=a)\Pr(X=a):

∣ψY⟩=∑b∈Γαab ∣b⟩Pr⁡(X=a)=∑b∈Γαab ∣b⟩∑c∈Γ∣αac∣2|\psi_Y\rangle=\frac{\sum_{b\in\Gamma}\alpha_{ab}\,|b\rangle}{\sqrt{\Pr(X=a)}}=\frac{\sum_{b\in\Gamma}\alpha_{ab}\,|b\rangle}{\sqrt{\sum_{c\in\Gamma}|\alpha_{ac}|^2}}

That square-root normalization is the one real difference from the classical case — otherwise partial measurement collapses a quantum compound system exactly the way it collapses a probabilistic one.

Example 3

Suppose (X,Y)(X,Y) is in the state below and we measure only XX:

∣ψ⟩=12 ∣00⟩+12 ∣01⟩+i22 ∣10⟩−122 ∣11⟩|\psi\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{2}\,|01\rangle+\tfrac{i}{2\sqrt{2}}\,|10\rangle-\tfrac{1}{2\sqrt{2}}\,|11\rangle

We begin by writing it grouped by XX, factoring each term into ∣a⟩⊗∣b⟩|a\rangle\otimes|b\rangle:

∣ψ⟩=∣0⟩⊗(12 ∣0⟩+12 ∣1⟩)+∣1⟩⊗(i22 ∣0⟩−122 ∣1⟩)|\psi\rangle=|0\rangle\otimes\Bigl(\tfrac{1}{\sqrt{2}}\,|0\rangle+\tfrac{1}{2}\,|1\rangle\Bigr)+|1\rangle\otimes\Bigl(\tfrac{i}{2\sqrt{2}}\,|0\rangle-\tfrac{1}{2\sqrt{2}}\,|1\rangle\Bigr)

The probability of each outcome is the squared norm of its branch. Recall the norm ∥⋅∥\|\cdot\| is a vector's length, so the squared norm is just the sum of the squared amplitude magnitudes:

∥∑bγb ∣b⟩∥2=∑b∣γb∣2\Bigl\|\sum_b \gamma_b\,|b\rangle\Bigr\|^2=\sum_b |\gamma_b|^2

Then keep the surviving branch and divide by Pr⁡(X=a)\sqrt{\Pr(X=a)} so it is a unit vector again.

Outcome X=0X=0

Pr⁡(X=0)=∥12∣0⟩+12∣1⟩∥2=12+14=34\Pr(X=0)=\Bigl\|\tfrac{1}{\sqrt{2}}|0\rangle+\tfrac{1}{2}|1\rangle\Bigr\|^2=\tfrac{1}{2}+\tfrac{1}{4}=\tfrac{3}{4}

Divide the X=0X=0 branch by 34=32\sqrt{\tfrac{3}{4}}=\tfrac{\sqrt{3}}{2}:

∣0⟩⊗12∣0⟩+12∣1⟩3/4=∣0⟩⊗(23 ∣0⟩+13 ∣1⟩)|0\rangle\otimes\frac{\tfrac{1}{\sqrt{2}}|0\rangle+\tfrac{1}{2}|1\rangle}{\sqrt{3/4}}=|0\rangle\otimes\Bigl(\sqrt{\tfrac{2}{3}}\,|0\rangle+\tfrac{1}{\sqrt{3}}\,|1\rangle\Bigr)

Outcome X=1X=1

Pr⁡(X=1)=∥i22∣0⟩−122∣1⟩∥2=18+18=14\Pr(X=1)=\Bigl\|\tfrac{i}{2\sqrt{2}}|0\rangle-\tfrac{1}{2\sqrt{2}}|1\rangle\Bigr\|^2=\tfrac{1}{8}+\tfrac{1}{8}=\tfrac{1}{4}

Divide the X=1X=1 branch by 14=12\sqrt{\tfrac{1}{4}}=\tfrac{1}{2}:

∣1⟩⊗i22∣0⟩−122∣1⟩1/4=∣1⟩⊗(i2 ∣0⟩−12 ∣1⟩)|1\rangle\otimes\frac{\tfrac{i}{2\sqrt{2}}|0\rangle-\tfrac{1}{2\sqrt{2}}|1\rangle}{\sqrt{1/4}}=|1\rangle\otimes\Bigl(\tfrac{i}{\sqrt{2}}\,|0\rangle-\tfrac{1}{\sqrt{2}}\,|1\rangle\Bigr)

Each squared amplitude in the post-measurement state now sums back to 11: for X=0X=0, 23+13=1\tfrac{2}{3}+\tfrac{1}{3}=1; for X=1X=1, 12+12=1\tfrac{1}{2}+\tfrac{1}{2}=1 — measuring XX has left YY in a valid quantum state.

Measuring some: Y

The same works the other way round. Take the same state but measure only YY, grouping each term by YY instead:

∣ψ⟩=(12 ∣0⟩+i22 ∣1⟩)⊗∣0⟩+(12 ∣0⟩−122 ∣1⟩)⊗∣1⟩|\psi\rangle=\Bigl(\tfrac{1}{\sqrt{2}}\,|0\rangle+\tfrac{i}{2\sqrt{2}}\,|1\rangle\Bigr)\otimes|0\rangle+\Bigl(\tfrac{1}{2}\,|0\rangle-\tfrac{1}{2\sqrt{2}}\,|1\rangle\Bigr)\otimes|1\rangle

Now each branch carries the amplitudes of XX. The same squared-norm rule gives the outcome probabilities, and the surviving branch is divided by Pr⁡(Y=b)\sqrt{\Pr(Y=b)}:

Outcome Y=0Y=0

Pr⁡(Y=0)=∥12∣0⟩+i22∣1⟩∥2=12+18=58\Pr(Y=0)=\Bigl\|\tfrac{1}{\sqrt{2}}|0\rangle+\tfrac{i}{2\sqrt{2}}|1\rangle\Bigr\|^2=\tfrac{1}{2}+\tfrac{1}{8}=\tfrac{5}{8}

Divide the Y=0Y=0 branch by 58\sqrt{\tfrac{5}{8}}:

12∣0⟩+i22∣1⟩5/8⊗∣0⟩=(45 ∣0⟩+i5 ∣1⟩)⊗∣0⟩\frac{\tfrac{1}{\sqrt{2}}|0\rangle+\tfrac{i}{2\sqrt{2}}|1\rangle}{\sqrt{5/8}}\otimes|0\rangle=\Bigl(\sqrt{\tfrac{4}{5}}\,|0\rangle+\tfrac{i}{\sqrt{5}}\,|1\rangle\Bigr)\otimes|0\rangle

Outcome Y=1Y=1

Pr⁡(Y=1)=∥12∣0⟩−122∣1⟩∥2=14+18=38\Pr(Y=1)=\Bigl\|\tfrac{1}{2}|0\rangle-\tfrac{1}{2\sqrt{2}}|1\rangle\Bigr\|^2=\tfrac{1}{4}+\tfrac{1}{8}=\tfrac{3}{8}

Divide the Y=1Y=1 branch by 38\sqrt{\tfrac{3}{8}}:

12∣0⟩−122∣1⟩3/8⊗∣1⟩=(23 ∣0⟩−13 ∣1⟩)⊗∣1⟩\frac{\tfrac{1}{2}|0\rangle-\tfrac{1}{2\sqrt{2}}|1\rangle}{\sqrt{3/8}}\otimes|1\rangle=\Bigl(\sqrt{\tfrac{2}{3}}\,|0\rangle-\tfrac{1}{\sqrt{3}}\,|1\rangle\Bigr)\otimes|1\rangle

As before the two probabilities add to 58+38=1\tfrac{5}{8}+\tfrac{3}{8}=1, and each collapsed XX state is a unit vector again.

Example 4 — three qubits

Nothing changes with more systems. Take the W state of (X,Y,Z)(X,Y,Z) and measure only the first qubit XX, grouping the rest as ∣x⟩⊗∣yz⟩|x\rangle\otimes|yz\rangle:

∣W⟩=13(∣001⟩+∣010⟩+∣100⟩)=∣0⟩⊗13(∣01⟩+∣10⟩)+∣1⟩⊗13∣00⟩|\mathrm{W}\rangle=\tfrac{1}{\sqrt{3}}\bigl(|001\rangle+|010\rangle+|100\rangle\bigr)=|0\rangle\otimes\tfrac{1}{\sqrt{3}}\bigl(|01\rangle+|10\rangle\bigr)+|1\rangle\otimes\tfrac{1}{\sqrt{3}}|00\rangle

Outcome X=0X=0

Pr⁡(X=0)=∥13∣01⟩+13∣10⟩∥2=13+13=23\Pr(X=0)=\Bigl\|\tfrac{1}{\sqrt{3}}|01\rangle+\tfrac{1}{\sqrt{3}}|10\rangle\Bigr\|^2=\tfrac{1}{3}+\tfrac{1}{3}=\tfrac{2}{3}

Divide the branch by 23\sqrt{\tfrac{2}{3}} — and (Y,Z)(Y,Z) is left in the Bell state ∣ψ+⟩|\psi^{+}\rangle:

13∣01⟩+13∣10⟩2/3=12(∣01⟩+∣10⟩)\frac{\tfrac{1}{\sqrt{3}}|01\rangle+\tfrac{1}{\sqrt{3}}|10\rangle}{\sqrt{2/3}}=\tfrac{1}{\sqrt{2}}\bigl(|01\rangle+|10\rangle\bigr)

Outcome X=1X=1

Pr⁡(X=1)=∥13∣00⟩∥2=13\Pr(X=1)=\Bigl\|\tfrac{1}{\sqrt{3}}|00\rangle\Bigr\|^2=\tfrac{1}{3}

Divide the branch by 13\sqrt{\tfrac{1}{3}} — and (Y,Z)(Y,Z) is left in the definite state ∣00⟩|00\rangle:

13∣00⟩1/3=∣00⟩\frac{\tfrac{1}{\sqrt{3}}|00\rangle}{\sqrt{1/3}}=|00\rangle

This shows the robustness of the W state mentioned earlier: with probability 23\tfrac{2}{3} losing one qubit still leaves the other two entangled — whereas a GHZ state would collapse to a fully definite ∣00⟩|00\rangle or ∣11⟩|11\rangle on either outcome.

Key idea: no matter how many systems there are, we can always regroup the state into the part being measured and the part left unmeasured — one ⊗\otimes branch per outcome of the measured part. Each outcome's probability is the squared norm of its branch, and the unmeasured part is left in that branch, renormalized by Pr⁡(outcome)\sqrt{\Pr(\text{outcome})}. The split into “measured” vs. “unmeasured” is all that matters — the same recipe handles any subset of any number of systems.

Unitary operations

Just like for a single system, a quantum operation on a compound system is represented by a unitary matrix — but now its rows and columns are indexed by the Cartesian product of the individual classical state sets, the same index set as the compound state vector it acts on.

For instance, if XX has state set {1,2,3}\{1,2,3\} and YY has state set {0,1}\{0,1\}, the compound system has 3×2=63\times 2=6 classical states, so an operation on (X,Y)(X,Y) is a 6×66\times 6 unitary matrix:

U=(121212001212i2−1200−i212−121200−120001212012−i2−1200i2000−12120)U=\begin{pmatrix} \tfrac{1}{2} & \tfrac{1}{2} & \tfrac{1}{2} & 0 & 0 & \tfrac{1}{2}\\[2pt] \tfrac{1}{2} & \tfrac{i}{2} & -\tfrac{1}{2} & 0 & 0 & -\tfrac{i}{2}\\[2pt] \tfrac{1}{2} & -\tfrac{1}{2} & \tfrac{1}{2} & 0 & 0 & -\tfrac{1}{2}\\[2pt] 0 & 0 & 0 & \tfrac{1}{\sqrt{2}} & \tfrac{1}{\sqrt{2}} & 0\\[2pt] \tfrac{1}{2} & -\tfrac{i}{2} & -\tfrac{1}{2} & 0 & 0 & \tfrac{i}{2}\\[2pt] 0 & 0 & 0 & -\tfrac{1}{\sqrt{2}} & \tfrac{1}{\sqrt{2}} & 0 \end{pmatrix}

Independent operations: tensor product

The general matrix above can entangle the systems it acts on. But often each system is acted on independently — a separate gate on each, with no interaction between them. Just as the tensor product combined separate states into one compound vector, it combines these separate operations into one compound unitary. If X1,…,XnX_1,\ldots,X_n carry the operations U1,…,UnU_1,\ldots,U_n, the combined action on (X1,…,Xn)(X_1,\ldots,X_n) is their tensor product:

U1⊗⋯⊗UnU_1\otimes\cdots\otimes U_n

Read it slot by slot: the matrix in each position is the operation that system experiences on its own. A tensor product of unitaries is again unitary, so independent operations always assemble into a single valid operation on the whole compound system.

The most common case is acting on just one part and leaving the rest alone — and doing nothing is itself a unitary, the identity matrix II. For instance, applying the Hadamard gate HH to the first qubit while leaving the second untouched is H⊗IH\otimes I, which passes the second qubit through unchanged:

H⊗I=(121212−12)⊗(1001)=(120120012012120−1200120−12)H\otimes I=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\otimes\begin{pmatrix}1&0\\0&1\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&0&\tfrac{1}{\sqrt{2}}&0\\[4pt]0&\tfrac{1}{\sqrt{2}}&0&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&0&-\tfrac{1}{\sqrt{2}}&0\\[4pt]0&\tfrac{1}{\sqrt{2}}&0&-\tfrac{1}{\sqrt{2}}\end{pmatrix}

Swapping the order applies HH to the second qubit instead — now HH appears as two identical blocks along the diagonal:

I⊗H=(1001)⊗(121212−12)=(12120012−12000012120012−12)I\otimes H=\begin{pmatrix}1&0\\0&1\end{pmatrix}\otimes\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\\[4pt]0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}

Operations that aren't tensor products: SWAP

A tensor product describes systems acted on independently, but not every unitary on a compound system factors that way. The standard example is the SWAP operation, which exchanges the contents of two systems XX and YY sharing the same classical state set Σ\Sigma:

SWAP(∣φ⟩⊗∣ψ⟩)=∣ψ⟩⊗∣φ⟩\mathrm{SWAP}\bigl(|\varphi\rangle\otimes|\psi\rangle\bigr)=|\psi\rangle\otimes|\varphi\rangle

To build that operation, consider one possible input basis state ∣a,b⟩|a,b\rangle. The bra ⟨a,b∣\langle a,b| recognizes that input, while the ket ∣b,a⟩|b,a\rangle supplies its swapped output. Their outer product therefore handles one input, and summing over every possible pair handles them all:

SWAP=∑a,b∈Σ∣b,a⟩⟨a,b∣=∑a,b∈Σ∣b⟩⟨a∣⊗∣a⟩⟨b∣\mathrm{SWAP}=\sum_{a,b\in\Sigma}|b,a\rangle\langle a,b|=\sum_{a,b\in\Sigma}|b\rangle\langle a|\otimes|a\rangle\langle b|

Each tensor-product term handles one specific pair: ∣b⟩⟨a∣|b\rangle\langle a| changes the first system from aa to bb, while ∣a⟩⟨b∣|a\rangle\langle b| changes the second from bb to aa. Together they map ∣a,b⟩|a,b\rangle to ∣b,a⟩|b,a\rangle. So every such tensor-product term swaps one specific pair.

For example, let's consider a simple system of two qubits. Each qubit has basis-state set Σ={0,1}\Sigma=\{0,1\}, so both aa and bb can be either 0 or 1. Writing out all four possible pairs in the sum gives:

SWAP=∣0⟩⟨0∣⊗∣0⟩⟨0∣⏟(a,b)=(0,0)+∣1⟩⟨0∣⊗∣0⟩⟨1∣⏟(a,b)=(0,1)+∣0⟩⟨1∣⊗∣1⟩⟨0∣⏟(a,b)=(1,0)+∣1⟩⟨1∣⊗∣1⟩⟨1∣⏟(a,b)=(1,1)\mathrm{SWAP}=\underbrace{|0\rangle\langle0|\otimes|0\rangle\langle0|}_{(a,b)=(0,0)}+\underbrace{|1\rangle\langle0|\otimes|0\rangle\langle1|}_{(a,b)=(0,1)}+\underbrace{|0\rangle\langle1|\otimes|1\rangle\langle0|}_{(a,b)=(1,0)}+\underbrace{|1\rangle\langle1|\otimes|1\rangle\langle1|}_{(a,b)=(1,1)}
=[(10)(10)]⊗[(10)(10)]+[(01)(10)]⊗[(10)(01)]+[(10)(01)]⊗[(01)(10)]+[(01)(01)]⊗[(01)(01)]\begin{aligned}&=\left[\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}\right]\otimes\left[\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}\right]+\left[\begin{pmatrix}0\\1\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}\right]\otimes\left[\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix}\right]\\[6pt]&\quad+\left[\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix}\right]\otimes\left[\begin{pmatrix}0\\1\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}\right]+\left[\begin{pmatrix}0\\1\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix}\right]\otimes\left[\begin{pmatrix}0\\1\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix}\right]\end{aligned}
=(1000)⊗(1000)+(0010)⊗(0100)+(0100)⊗(0010)+(0001)⊗(0001)=\begin{pmatrix}1&0\\0&0\end{pmatrix}\otimes\begin{pmatrix}1&0\\0&0\end{pmatrix}+\begin{pmatrix}0&0\\1&0\end{pmatrix}\otimes\begin{pmatrix}0&1\\0&0\end{pmatrix}+\begin{pmatrix}0&1\\0&0\end{pmatrix}\otimes\begin{pmatrix}0&0\\1&0\end{pmatrix}+\begin{pmatrix}0&0\\0&1\end{pmatrix}\otimes\begin{pmatrix}0&0\\0&1\end{pmatrix}
=(1000000000000000)+(0000000001000000)+(0000001000000000)+(0000000000000001)=\begin{pmatrix}1&0&0&0\\0&0&0&0\\0&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&0\\0&1&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&1&0\\0&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&0\\0&0&0&0\\0&0&0&1\end{pmatrix}
=(1000001001000001)=\begin{pmatrix}1&0&0&0\\0&0&1&0\\0&1&0&0\\0&0&0&1\end{pmatrix}

To apply SWAP, multiply its matrix by the compound state vector. For example, the basis state ∣01⟩|01\rangle is the second standard basis vector, so the multiplication moves its 1 into the position for ∣10⟩|10\rangle:

SWAP∣01⟩=(1000001001000001)(0100)=(0010)=∣10⟩\mathrm{SWAP}|01\rangle=\begin{pmatrix}1&0&0&0\\0&0&1&0\\0&1&0&0\\0&0&0&1\end{pmatrix}\begin{pmatrix}0\\1\\0\\0\end{pmatrix}=\begin{pmatrix}0\\0\\1\\0\end{pmatrix}=|10\rangle

Controlled operations

Suppose that XX is a qubit and YY is an arbitrary quantum system. A controlled-UU operation uses XX as a switch for a unitary operation UU on YY. When the control is ∣0⟩|0\rangle, the target system is left unchanged; when the control is ∣1⟩|1\rangle, the operation UU is applied to it.

The projectors ∣0⟩⟨0∣|0\rangle\langle0| and ∣1⟩⟨1∣|1\rangle\langle1| select those two branches, giving the following operation on the pair (X,Y)(X,Y):

controlled⁡-U=∣0⟩⟨0∣⊗IY+∣1⟩⟨1∣⊗U=(IY00U)\operatorname{controlled}\text{-}U=|0\rangle\langle0|\otimes I_Y+|1\rangle\langle1|\otimes U=\begin{pmatrix}I_Y&0\\0&U\end{pmatrix}

The matrix (IY00U)\begin{pmatrix}I_Y&0\\0&U\end{pmatrix} is a block matrix, not an ordinary 2 × 2 matrix. If YY has dd basis states, then IYI_Y, UU, and each zero are all d×dd\times d blocks. The complete controlled-UU matrix is therefore 2d×2d2d\times2d. In particular, if YY is a qubit, the blocks are 2×22\times2 and the full matrix is 4×44\times4.

Example: controlled-NOT

Let the target YY also be a qubit and choose U=σxU=\sigma_x. On one qubit, σx\sigma_xflips the two basis states:

σx=(0110),σx∣0⟩=∣1⟩,σx∣1⟩=∣0⟩\sigma_x=\begin{pmatrix}0&1\\1&0\end{pmatrix},\qquad \sigma_x|0\rangle=|1\rangle,\qquad \sigma_x|1\rangle=|0\rangle

Now substitute U=σxU=\sigma_x and IY=I2I_Y=I_2 into the controlled-UU formula. The following equalities turn the general controlled operation into the controlled-NOT, or CNOT, gate:

CNOT=∣0⟩⟨0∣⊗I2+∣1⟩⟨1∣⊗σx\mathrm{CNOT}=|0\rangle\langle0|\otimes I_2+|1\rangle\langle1|\otimes\sigma_x
=(1000)⊗(1001)+(0001)⊗(0110)=\begin{pmatrix}1&0\\0&0\end{pmatrix}\otimes\begin{pmatrix}1&0\\0&1\end{pmatrix}+\begin{pmatrix}0&0\\0&1\end{pmatrix}\otimes\begin{pmatrix}0&1\\1&0\end{pmatrix}
=(I200σx)=\begin{pmatrix}I_2&0\\0&\sigma_x\end{pmatrix}
=(1000010000010010)=\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}

In this example, the first digit is the control and the second is the target. Therefore ∣10⟩|10\rangle means that the control is ∣1⟩|1\rangle and the target is ∣0⟩|0\rangle. A control value of 1 turns the operation on. The control stays at 1, while σx\sigma_x flips the target from 0 to 1:

CNOT∣10⟩=∣1⟩⊗σx∣0⟩=∣1⟩⊗∣1⟩=∣11⟩\mathrm{CNOT}|10\rangle=|1\rangle\otimes\sigma_x|0\rangle=|1\rangle\otimes|1\rangle=|11\rangle

In the standard basis order ∣00⟩,∣01⟩,∣10⟩,∣11⟩|00\rangle,|01\rangle,|10\rangle,|11\rangle, the state ∣10⟩|10\rangle is the third basis vector. The full matrix multiplication shows the same change from the third basis vector to the fourth:

CNOT∣10⟩=(1000010000010010)(0010)⏟∣10⟩=(0001)⏟∣11⟩\mathrm{CNOT}|10\rangle=\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}\underbrace{\begin{pmatrix}0\\0\\1\\0\end{pmatrix}}_{|10\rangle}=\underbrace{\begin{pmatrix}0\\0\\0\\1\end{pmatrix}}_{|11\rangle}

The control can be any qubit

The first qubit is the control above only because we chose it that way. A controlled operation can use any qubit as its control. If the second qubit controls an operation UU on the first, the projectors move to the second tensor factor:

CU(2→1)=I2⊗∣0⟩⟨0∣+U⊗∣1⟩⟨1∣C_U^{(2\to1)}=I_2\otimes|0\rangle\langle0|+U\otimes|1\rangle\langle1|

With that ordering, the second digit of ∣a,b⟩|a,b\rangle is the control, and the first digit is the target.

Example: Fredkin operation

A controlled-SWAP on three qubits swaps the last two qubits only when the first qubit is 1. This is called the Fredkin operation, or Fredkin gate:

Fredkin=∣0⟩⟨0∣⊗I4+∣1⟩⟨1∣⊗SWAP=(1000000001000000001000000001000000001000000000100000010000000001)\mathrm{Fredkin}=|0\rangle\langle0|\otimes I_4+|1\rangle\langle1|\otimes\mathrm{SWAP}=\begin{pmatrix}1&0&0&0&0&0&0&0\\0&1&0&0&0&0&0&0\\0&0&1&0&0&0&0&0\\0&0&0&1&0&0&0&0\\0&0&0&0&1&0&0&0\\0&0&0&0&0&0&1&0\\0&0&0&0&0&1&0&0\\0&0&0&0&0&0&0&1\end{pmatrix}

Example: Toffoli operation

A controlled-controlled-NOT on three qubits flips the third qubit only when both the first and second qubits are 1. This is called the Toffoli operation, or Toffoli gate:

Toffoli=∣0⟩⟨0∣⊗I2⊗I2+∣1⟩⟨1∣⊗(∣0⟩⟨0∣⊗I2+∣1⟩⟨1∣⊗σx)=(1000000001000000001000000001000000001000000001000000000100000010)\mathrm{Toffoli}=|0\rangle\langle0|\otimes I_2\otimes I_2+|1\rangle\langle1|\otimes\left(|0\rangle\langle0|\otimes I_2+|1\rangle\langle1|\otimes\sigma_x\right)=\begin{pmatrix}1&0&0&0&0&0&0&0\\0&1&0&0&0&0&0&0\\0&0&1&0&0&0&0&0\\0&0&0&1&0&0&0&0\\0&0&0&0&1&0&0&0\\0&0&0&0&0&1&0&0\\0&0&0&0&0&0&0&1\\0&0&0&0&0&0&1&0\end{pmatrix}

Quantum circuits

Circuits as models of computation

A circuit is a graphical model of computation. It describes how information moves through a computation and which operations are applied along the way.

Wires represent the paths along which values travel from one part of the computation to another. Gates receive one or more input values, perform an operation on them, and pass the resulting values along their outgoing wires.

Interactive Boolean circuit editor

Click X or Y to toggle it. Drag nodes to move them, drag square wire handles to reroute, and drag circular output ports onto input ports to connect. Drag an occupied input port to reconnect its wire. Right-click anywhere on the board for context-specific actions.

INPUTSOUTPUT0000101X0Y0FANOUT · 0NOT · 1AND · 0OR · 11Z

The quantum circuit model

In the quantum circuit model, wires represent qubits and gates represent both unitary operations and measurements.

The terminology comes from physical electrical circuits, where wires carry current and components act on it. Here, wires and gates describe information flow, not electricity. The circuits considered here are acyclic: information flows in one direction, conventionally from left to right. A wire never loops back to an earlier gate, so the gates define an unambiguous order of computation.

A one-qubit circuit

From left to right, qubit XX passes through H, S, H, TH,\ S,\ H,\ T. Operators act on a ket from the left, so the rightmost matrix in a product acts first. The output is therefore ∣ψout⟩=THSH∣ψin⟩|\psi_{\mathrm{out}}\rangle=THSH|\psi_{\mathrm{in}}\rangle: the combined operation is written THSHTHSH, the reverse of the visual HSHTHSHT order. Hover over, tap or focus a gate to see its full operation.

X

The three gate matrices are:

H=12(111−1),S=(100i),T=(100eiπ/4)=(1001+i2)H=\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix},\qquad S=\begin{pmatrix}1&0\\0&i\end{pmatrix},\qquad T=\begin{pmatrix}1&0\\0&e^{i\pi/4}\end{pmatrix}=\begin{pmatrix}1&0\\0&\tfrac{1+i}{\sqrt{2}}\end{pmatrix}

Multiply one step at a time, starting at the right:

SH=(100i)12(111−1)=12(11i−i)SH=\begin{pmatrix}1&0\\0&i\end{pmatrix}\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix}=\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\i&-i\end{pmatrix}
HSH=H(SH)=12(111−1)12(11i−i)=12(1+i1−i1−i1+i)HSH=H(SH)=\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix}\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\i&-i\end{pmatrix}=\frac{1}{2}\begin{pmatrix}1+i&1-i\\1-i&1+i\end{pmatrix}
THSH=T(HSH)=(1001+i2)12(1+i1−i1−i1+i)=(1+i21−i212i2)THSH=T(HSH)=\begin{pmatrix}1&0\\0&\tfrac{1+i}{\sqrt{2}}\end{pmatrix}\frac{1}{2}\begin{pmatrix}1+i&1-i\\1-i&1+i\end{pmatrix}=\begin{pmatrix}\tfrac{1+i}{2}&\tfrac{1-i}{2}\\\tfrac{1}{\sqrt{2}}&\tfrac{i}{\sqrt{2}}\end{pmatrix}

A controlled-NOT circuit

The Hadamard gate first acts on the upper qubit YY. The filled dot then makes YY the CNOT control, while the circled plus marks XX as its target. If Y=0Y=0, the target is unchanged; if Y=1Y=1, the target is flipped.

YX

For the matrix calculation, use the register order (X,Y)(X,Y): ∣00⟩XY,∣01⟩XY,∣10⟩XY,∣11⟩XY|00\rangle_{XY},|01\rangle_{XY},|10\rangle_{XY},|11\rangle_{XY}. The Hadamard acts on YY, the second tensor factor, so its two-qubit matrix is I2⊗HI_2\otimes H:

I2⊗H=(1001)⊗12(111−1)=(12120012−12000012120012−12)I_2\otimes H=\begin{pmatrix}1&0\\0&1\end{pmatrix}\otimes\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\\0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}

With YY as the control, CNOT exchanges ∣01⟩|01\rangle and ∣11⟩|11\rangle while leaving the other two basis states unchanged:

CNOTY→X=(1000000100100100)\mathrm{CNOT}_{Y\to X}=\begin{pmatrix}1&0&0&0\\0&0&0&1\\0&0&1&0\\0&1&0&0\end{pmatrix}

CNOT acts after the Hadamard, so it appears on the left in the product:

U=CNOTY→X(I2⊗H)=(1000000100100100)(12120012−12000012120012−12)=(1212000012−1200121212−1200)U=\mathrm{CNOT}_{Y\to X}(I_2\otimes H)=\begin{pmatrix}1&0&0&0\\0&0&0&1\\0&0&1&0\\0&1&0&0\end{pmatrix}\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\\0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\\0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\end{pmatrix}

Applying that matrix to ∣00⟩XY|00\rangle_{XY} gives the Bell state:

U∣00⟩XY=(1212000012−1200121212−1200)(1000)=(120012)=∣00⟩+∣11⟩2U|00\rangle_{XY}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\\0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\end{pmatrix}\begin{pmatrix}1\\0\\0\\0\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\0\\0\\\tfrac{1}{\sqrt{2}}\end{pmatrix}=\frac{|00\rangle+|11\rangle}{\sqrt{2}}

For all four computational-basis inputs, the circuit produces the four Bell states below. The final minus sign is a global phase and does not affect measurement probabilities.

U∣00⟩XY=∣00⟩+∣11⟩2=∣ϕ+⟩,U∣01⟩XY=∣00⟩−∣11⟩2=∣ϕ−⟩,U∣10⟩XY=∣01⟩+∣10⟩2=∣ψ+⟩,U∣11⟩XY=−∣01⟩+∣10⟩2=−∣ψ−⟩.\begin{aligned}U|00\rangle_{XY}&=\frac{|00\rangle+|11\rangle}{\sqrt{2}}=|\phi^{+}\rangle,\\[6pt]U|01\rangle_{XY}&=\frac{|00\rangle-|11\rangle}{\sqrt{2}}=|\phi^{-}\rangle,\\[6pt]U|10\rangle_{XY}&=\frac{|01\rangle+|10\rangle}{\sqrt{2}}=|\psi^{+}\rangle,\\[6pt]U|11\rangle_{XY}&=\frac{-|01\rangle+|10\rangle}{\sqrt{2}}=-|\psi^{-}\rangle.\end{aligned}

Follow the state through the circuit

To find the output for one input, we do not need to multiply the full gate matrices. Instead, follow the state from left to right and update it at each gate; the slices below are labeled ∣π0⟩,∣π1⟩,∣π2⟩|\pi_0\rangle,|\pi_1\rangle,|\pi_2\rangle.

∣0⟩|0\rangle
∣0⟩|0\rangle
∣π0⟩|\pi_0\rangle
∣π1⟩|\pi_1\rangle
∣π2⟩|\pi_2\rangle
∣π0⟩=∣0⟩X∣0⟩Y=∣00⟩XY,∣π1⟩=∣0⟩X∣+⟩Y=∣00⟩+∣01⟩2,∣π2⟩=CNOTY→X∣π1⟩=∣00⟩+∣11⟩2=∣ϕ+⟩.\begin{aligned}|\pi_0\rangle&=|0\rangle_X|0\rangle_Y=|00\rangle_{XY},\\[6pt]|\pi_1\rangle&=|0\rangle_X|+\rangle_Y=\frac{|00\rangle+|01\rangle}{\sqrt{2}},\\[6pt]|\pi_2\rangle&=\mathrm{CNOT}_{Y\to X}|\pi_1\rangle=\frac{|00\rangle+|11\rangle}{\sqrt{2}}=|\phi^+\rangle.\end{aligned}

Classical states in a quantum circuit

A measurement connects quantum and classical information: it collapses a qubit to ∣0⟩|0\rangle or ∣1⟩|1\rangle and writes the corresponding bit onto a double classical wire. Here the measurements store Y in B and X in A. Because the qubits are in ∣ϕ+⟩|\phi^+\rangle, the classical result is (A,B)=(0,0)(A,B)=(0,0) or (1,1)(1,1), each with probability one half; those bits can then control later operations or be read as output.

YXBA

Quantum-circuits symbols

Single-qubit gates

Controlled-NOT (CNOT)

SWAP

Toffoli (CCNOT)

Fredkin (controlled-SWAP)

Arbitrary unitary U

Controlled-U

Quantum states and measurements

Inner products

Suppose that we have two kets ∣ψ⟩\lvert\psi\rangle and ∣ϕ⟩\lvert\phi\rangle with complex amplitudes:

∣ψ⟩=(α1⋮αn)∣ϕ⟩=(β1⋮βn)\lvert\psi\rangle=\begin{pmatrix}\alpha_1\\\vdots\\\alpha_n\end{pmatrix}\qquad\lvert\phi\rangle=\begin{pmatrix}\beta_1\\\vdots\\\beta_n\end{pmatrix}

To calculate the inner product of ∣ψ⟩\lvert\psi\rangle and ∣ϕ⟩\lvert\phi\rangle, first turn ∣ψ⟩\lvert\psi\rangle into the bra ⟨ψ∣=∣ψ⟩†\langle\psi\rvert=\lvert\psi\rangle^\dagger — its conjugate transpose, whose entries are the complex conjugates αi‾\overline{\alpha_i}— then multiply it by ∣ϕ⟩\lvert\phi\rangle to get a single number:

⟨ψ∣ϕ⟩=(α1‾⋯αn‾)(β1⋮βn)=α1‾β1+⋯+αn‾βn\langle\psi\vert\phi\rangle=\begin{pmatrix}\overline{\alpha_1}&\cdots&\overline{\alpha_n}\end{pmatrix}\begin{pmatrix}\beta_1\\\vdots\\\beta_n\end{pmatrix}=\overline{\alpha_1}\beta_1+\cdots+\overline{\alpha_n}\beta_n

For example, take the two qubit states:

∣ψ⟩=(12i2)∣ϕ⟩=(1212)\lvert\psi\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{i}{\sqrt{2}}\end{pmatrix}\qquad\lvert\phi\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}\end{pmatrix}

Conjugating the amplitudes of ∣ψ⟩\lvert\psi\rangle flips the ii to −i-i:

⟨ψ∣ϕ⟩=(12−i2)(1212)=12⋅12−i2⋅12=1−i2\langle\psi\vert\phi\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&-\tfrac{i}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[2pt]\tfrac{1}{\sqrt{2}}\end{pmatrix}=\tfrac{1}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}-\tfrac{i}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}=\frac{1-i}{2}

Alternatively, suppose the same vectors are written as sums over a basis Σ\Sigma:

∣ψ⟩=∑a∈Σαa∣a⟩∣ϕ⟩=∑b∈Σβb∣b⟩\lvert\psi\rangle=\sum_{a\in\Sigma}\alpha_a\lvert a\rangle\qquad\lvert\phi\rangle=\sum_{b\in\Sigma}\beta_b\lvert b\rangle

This gives the same formula as before, just in a different format. Expanding the product term by term, each ⟨a∣b⟩\langle a\vert b\rangle is 11 when a=ba=b and 00 otherwise, so only the matching terms survive:

⟨ψ∣ϕ⟩=(∑a∈Σαa‾⟨a∣)(∑b∈Σβb∣b⟩)=∑a∈Σ∑b∈Σαa‾βb⟨a∣b⟩=∑a∈Σαa‾βa\begin{aligned}\langle\psi\vert\phi\rangle&=\left(\sum_{a\in\Sigma}\overline{\alpha_a}\langle a\rvert\right)\left(\sum_{b\in\Sigma}\beta_b\lvert b\rangle\right)\\[2pt]&=\sum_{a\in\Sigma}\sum_{b\in\Sigma}\overline{\alpha_a}\beta_b\langle a\vert b\rangle\\[2pt]&=\sum_{a\in\Sigma}\overline{\alpha_a}\beta_a\end{aligned}

Our example states in this form, over the basis Σ={0,1}\Sigma=\{0,1\}:

∣ψ⟩=12∣0⟩+i2∣1⟩∣ϕ⟩=12∣0⟩+12∣1⟩\lvert\psi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{i}{\sqrt{2}}\lvert1\rangle\qquad\lvert\phi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{\sqrt{2}}\lvert1\rangle

Substituting our values into the formula, we get (the ⟨0∣1⟩\langle 0\vert 1\rangle and ⟨1∣0⟩\langle 1\vert 0\rangle terms are 00, so they cancel):

⟨ψ∣ϕ⟩=(12⟨0∣−i2⟨1∣)(12∣0⟩+12∣1⟩)=12⟨0∣0⟩+12⟨0∣1⟩−i2⟨1∣0⟩−i2⟨1∣1⟩=12−i2=1−i2\begin{aligned}\langle\psi\vert\phi\rangle&=\left(\tfrac{1}{\sqrt{2}}\langle 0\rvert-\tfrac{i}{\sqrt{2}}\langle 1\rvert\right)\left(\tfrac{1}{\sqrt{2}}\lvert 0\rangle+\tfrac{1}{\sqrt{2}}\lvert 1\rangle\right)\\[4pt]&=\tfrac{1}{2}\langle 0\vert 0\rangle+\cancel{\tfrac{1}{2}\langle 0\vert 1\rangle}-\cancel{\tfrac{i}{2}\langle 1\vert 0\rangle}-\tfrac{i}{2}\langle 1\vert 1\rangle\\[4pt]&=\tfrac{1}{2}-\tfrac{i}{2}=\frac{1-i}{2}\end{aligned}

Inner product as an angle

When the amplitudes are all real, the conjugation does nothing (α‾=α\overline{\alpha}=\alpha), and the inner product has a clean geometric meaning: for two unit vectors it equals the cosine of the angle between them.

⟨ψ∣ϕ⟩=α0β0+α1β1=cos⁡θ\langle\psi\vert\phi\rangle=\alpha_0\beta_0+\alpha_1\beta_1=\cos\theta
|0⟩−|0⟩|1⟩−|1⟩105°ψϕ
|ψ⟩ = 1/√2|0⟩ − 1/√2|1⟩
|ϕ⟩ = 1/2|0⟩ + √3/2|1⟩
θ = 105°
⟨ψ|ϕ⟩ = cos θ = -0.26
Drag either tip around the circle to see how cos changes value

This geometric interpretation only works when all amplitudes are real. With complex amplitudes, the inner product becomes a complex number, so it no longer represents the cosine of an ordinary angle. The quantity that remains physically meaningful is its magnitude, ∣⟨ψ∣ϕ⟩∣\lvert\langle\psi\vert\phi\rangle\rvert, which determines measurement probabilities.

Inner product as an angle in multidimensional space

The idea of the inner product as an angle for real amplitudes generalizes to any number of dimensions: for real unit vectors the inner product is still the cosine of the angle between them — it just lives on a higher-dimensional sphere. Below is the three-dimensional case (a qutrit with real amplitudes, basis ∣0⟩,∣1⟩,∣2⟩\lvert0\rangle,\lvert1\rangle,\lvert2\rangle):

⟨ψ∣ϕ⟩=α0β0+α1β1+α2β2=cos⁡θ\langle\psi\vert\phi\rangle=\alpha_0\beta_0+\alpha_1\beta_1+\alpha_2\beta_2=\cos\theta
|0⟩|1⟩|2⟩
|ψ⟩ = 1/√2|0⟩ + 0|1⟩ + 1/√2|2⟩
|ϕ⟩ = 0|0⟩ + 1/√2|1⟩ + 1/√2|2⟩
θ = 60°
⟨ψ|ϕ⟩ = cos θ = 0.50
Drag a tip to move it on the sphere; drag the background to rotate the view.

Properties of the inner product

Relationship to the Euclidean norm

Take the inner product of a vector ∣ψ⟩=∑a∈Σαa∣a⟩\lvert\psi\rangle=\sum_{a\in\Sigma}\alpha_a\lvert a\rangle with itself. Each conjugate pair collapses to a squared magnitude, αa‾αa=∣αa∣2\overline{\alpha_a}\alpha_a=\lvert\alpha_a\rvert^2, so the result is the sum of squared amplitudes — exactly the squared Euclidean length of the vector:

⟨ψ∣ψ⟩=∑a∈Σαa‾αa=∑a∈Σ∣αa∣2=∥ ∣ψ⟩ ∥2\langle\psi\vert\psi\rangle=\sum_{a\in\Sigma}\overline{\alpha_a}\alpha_a=\sum_{a\in\Sigma}\lvert\alpha_a\rvert^2=\bigl\lVert\,\lvert\psi\rangle\,\bigr\rVert^2

When a state is normalized (that is, it represents a valid quantum state), its inner product with itself equals 11. Equivalently, its Euclidean norm equals 11. For example, for the vector ∣ψ⟩=12∣0⟩+i2∣1⟩\lvert\psi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{i}{\sqrt{2}}\lvert1\rangle:

⟨ψ∣ψ⟩=∣12∣2+∣i2∣2=12+12=1\langle\psi\vert\psi\rangle=\left\lvert\tfrac{1}{\sqrt{2}}\right\rvert^2+\left\lvert\tfrac{i}{\sqrt{2}}\right\rvert^2=\tfrac{1}{2}+\tfrac{1}{2}=1

More generally, the Euclidean norm of any vector is the square root of its inner product with itself:

∥ ∣ψ⟩ ∥=⟨ψ∣ψ⟩\bigl\lVert\,\lvert\psi\rangle\,\bigr\rVert=\sqrt{\langle\psi\vert\psi\rangle}

Conjugate symmetry

Swapping the order of the two vectors conjugates a different set of amplitudes. For ∣ψ⟩=∑a∈Σαa∣a⟩\lvert\psi\rangle=\sum_{a\in\Sigma}\alpha_a\lvert a\rangle and ∣ϕ⟩=∑a∈Σβa∣a⟩\lvert\phi\rangle=\sum_{a\in\Sigma}\beta_a\lvert a\rangle, the two orderings are

⟨ψ∣ϕ⟩=∑a∈Σαa‾βaand⟨ϕ∣ψ⟩=∑a∈Σβa‾αa,\langle\psi\vert\phi\rangle=\sum_{a\in\Sigma}\overline{\alpha_a}\beta_a\qquad\text{and}\qquad\langle\phi\vert\psi\rangle=\sum_{a\in\Sigma}\overline{\beta_a}\alpha_a,

which differ only in which factor carries the bar. In fact the two are complex conjugates of each other:

⟨ψ∣ϕ⟩‾=⟨ϕ∣ψ⟩.\overline{\langle\psi\vert\phi\rangle}=\langle\phi\vert\psi\rangle.

Why. Conjugate ⟨ψ∣ϕ⟩\langle\psi\vert\phi\rangle, pulling the bar inside the sum and onto each factor:

⟨ψ∣ϕ⟩‾=∑a∈Σαa‾βa‾=∑a∈Σαa‾βa‾=∑a∈Σαaβa‾.\overline{\langle\psi\vert\phi\rangle}=\overline{\sum_{a\in\Sigma}\overline{\alpha_a}\beta_a}=\sum_{a\in\Sigma}\overline{\overline{\alpha_a}\beta_a}=\sum_{a\in\Sigma}\alpha_a\overline{\beta_a}.

Because complex multiplication is commutative, αaβa‾=βa‾αa\alpha_a\overline{\beta_a}=\overline{\beta_a}\alpha_a, the final summand is exactly that of ⟨ϕ∣ψ⟩\langle\phi\vert\psi\rangle, which proves the identity.

Example. Take the two states

∣ψ⟩=12∣0⟩+i2∣1⟩∣ϕ⟩=12∣0⟩+12∣1⟩\lvert\psi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{i}{\sqrt{2}}\lvert1\rangle\qquad\lvert\phi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{\sqrt{2}}\lvert1\rangle

Computing both orderings yields a conjugate pair:

⟨ψ∣ϕ⟩=12⋅12−i2⋅12=1−i2\langle\psi\vert\phi\rangle=\tfrac{1}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}-\tfrac{i}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}=\frac{1-i}{2}
⟨ϕ∣ψ⟩=12⋅12+12⋅i2=1+i2=(1−i2)‾.\langle\phi\vert\psi\rangle=\tfrac{1}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}+\tfrac{1}{\sqrt{2}}\cdot\tfrac{i}{\sqrt{2}}=\frac{1+i}{2}=\overline{\left(\tfrac{1-i}{2}\right)}.

Linearity in the second argument

Suppose that ∣ψ⟩\lvert\psi\rangle, ∣ϕ1⟩\lvert\phi_1\rangle, and ∣ϕ2⟩\lvert\phi_2\rangle are vectors and α1\alpha_1 and α2\alpha_2 are complex numbers. If we define a new vector

∣ϕ⟩=α1∣ϕ1⟩+α2∣ϕ2⟩,\lvert\phi\rangle=\alpha_1\lvert\phi_1\rangle+\alpha_2\lvert\phi_2\rangle,

then the inner product distributes over the combination, with each coefficient pulled out in front:

⟨ψ∣ϕ⟩=⟨ψ∣(α1∣ϕ1⟩+α2∣ϕ2⟩)=α1⟨ψ∣ϕ1⟩+α2⟨ψ∣ϕ2⟩.\langle\psi\vert\phi\rangle=\langle\psi\vert\bigl(\alpha_1\lvert\phi_1\rangle+\alpha_2\lvert\phi_2\rangle\bigr)=\alpha_1\langle\psi\vert\phi_1\rangle+\alpha_2\langle\psi\vert\phi_2\rangle.

Conjugate linearity in the first argument

Suppose that ∣ψ1⟩\lvert\psi_1\rangle, ∣ψ2⟩\lvert\psi_2\rangle, and ∣ϕ⟩\lvert\phi\rangle are vectors and β1\beta_1 and β2\beta_2 are complex numbers. If we define a new vector

∣ψ⟩=β1∣ψ1⟩+β2∣ψ2⟩,\lvert\psi\rangle=\beta_1\lvert\psi_1\rangle+\beta_2\lvert\psi_2\rangle,

then the inner product is again linear in each term — but forming the bra ⟨ψ∣\langle\psi\vert conjugates every coefficient as it comes out in front:

⟨ψ∣ϕ⟩=(β1‾⟨ψ1∣+β2‾⟨ψ2∣)∣ϕ⟩=β1‾⟨ψ1∣ϕ⟩+β2‾⟨ψ2∣ϕ⟩.\langle\psi\vert\phi\rangle=\bigl(\overline{\beta_1}\langle\psi_1\vert+\overline{\beta_2}\langle\psi_2\vert\bigr)\lvert\phi\rangle=\overline{\beta_1}\langle\psi_1\vert\phi\rangle+\overline{\beta_2}\langle\psi_2\vert\phi\rangle.

The Cauchy–Schwarz inequality

For every choice of vectors ∣ψ⟩\lvert\psi\rangle and ∣ϕ⟩\lvert\phi\rangle, the magnitude of the inner product can never be larger than the product of the vectors' lengths:

∣⟨ψ∣ϕ⟩∣≤∥ ∣ψ⟩ ∥  ∥ ∣ϕ⟩ ∥.\bigl\lvert\langle\psi\vert\phi\rangle\bigr\rvert\le\bigl\lVert\,\lvert\psi\rangle\,\bigr\rVert\;\bigl\lVert\,\lvert\phi\rangle\,\bigr\rVert.

In other words, the overlap between two vectors cannot exceed what their lengths allow.

Equality, ∣⟨ψ∣ϕ⟩∣=∥ ∣ψ⟩ ∥  ∥ ∣ϕ⟩ ∥\bigl\lvert\langle\psi\vert\phi\rangle\bigr\rvert=\bigl\lVert\,\lvert\psi\rangle\,\bigr\rVert\;\bigl\lVert\,\lvert\phi\rangle\,\bigr\rVert, holds only when the two vectors are linearly dependent — that is, one is simply a scalar multiple of the other.

Orthogonality and orthonormality

Two vectors ∣ψ⟩\lvert\psi\rangle and ∣ϕ⟩\lvert\phi\rangle are orthogonal if their inner product is zero:

⟨ψ∣ϕ⟩=0.\langle\psi\vert\phi\rangle=0.

An orthogonal set {∣ψ1⟩,…,∣ψm⟩}\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\} is one in which every pair is orthogonal:

⟨ψj∣ψk⟩=0(for all j≠k).\langle\psi_j\vert\psi_k\rangle=0\qquad(\text{for all }j\neq k).

An orthonormal set {∣ψ1⟩,…,∣ψm⟩}\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\} is an orthogonal set of unit vectors — each pair is orthogonal and each vector has length one:

⟨ψj∣ψk⟩={1j=k0j≠k(for all j,k).\langle\psi_j\vert\psi_k\rangle=\begin{cases}1 & j=k\\[2pt]0 & j\neq k\end{cases}\qquad(\text{for all }j,k).

An orthonormal basis {∣ψ1⟩,…,∣ψm⟩}\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\} is a set of orthonormal vectors that spans the entire vector space. This means every vector in the space can be written as a linear combination of the basis vectors.

Constructing basis sets

The key idea is that any orthonormal set can always be extended into an orthonormal basis. Suppose that {∣ψ1⟩,…,∣ψm⟩}\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\} is an orthonormal set of vectors in an nn-dimensional space. Because orthonormal sets are always linearly independent, these vectors span a subspace of dimension m≤nm\leq n.

If m<nm<n, then there must exist additional vectors ∣ψm+1⟩,…,∣ψn⟩\lvert\psi_{m+1}\rangle,\dots,\lvert\psi_n\rangle so that {∣ψ1⟩,…,∣ψn⟩}\{\lvert\psi_1\rangle,\dots,\lvert\psi_n\rangle\} forms an orthonormal basis.

The Gram–Schmidt orthogonalization process can be used to construct these extra vectors. It repeatedly subtracts a vector's projections onto the existing basis vectors, leaving an orthogonal remainder that is then normalized to get the missing basis vectors.

Two dependent vectors — the span is only a line.

Independent vectors span the whole plane.

Orthonormal basis — vectors perpendicular and have unit length.

Orthonormal bases and unitary matrices

A unitary matrix is just a matrix whose columns are an orthonormal basis. More precisely, a square matrix UU is unitary if and only if its columns form an orthonormal basis — and equivalently, if and only if its rows do. These conditions are equivalent:

  1. U†U=I=UU†U^\dagger U=I=UU^\dagger (UU is unitary).
  2. The columns of UU form an orthonormal basis.
  3. The rows of UU form an orthonormal basis.

To see why, write the columns of UU as vectors and take the conjugate transpose, which turns each column ket into the corresponding row bra. Multiplying U†UU^\dagger U then gives, in row jj and column kk, the inner product ⟨ψj∣ψk⟩\langle\psi_j\vert\psi_k\rangle:

U=[∣∣∣∣ψ1⟩∣ψ2⟩⋯∣ψn⟩∣∣∣]U†=[⟨ψ1∣⟨ψ2∣⋮⟨ψn∣](U†U)jk=⟨ψj∣ψk⟩U=\left[\begin{array}{cccc}\textcolor{#0284c7}{\rvert}&\textcolor{#7c3aed}{\rvert}& &\textcolor{#db2777}{\rvert}\\[1pt]\textcolor{#0284c7}{\lvert\psi_1\rangle}&\textcolor{#7c3aed}{\lvert\psi_2\rangle}&\cdots&\textcolor{#db2777}{\lvert\psi_n\rangle}\\[1pt]\textcolor{#0284c7}{\rvert}&\textcolor{#7c3aed}{\rvert}& &\textcolor{#db2777}{\rvert}\end{array}\right]\qquad U^\dagger=\left[\begin{array}{ccc}\textcolor{#0284c7}{\rule[0.45ex]{1.3em}{0.5pt}}&\textcolor{#0284c7}{\langle\psi_1\rvert}&\textcolor{#0284c7}{\rule[0.45ex]{1.3em}{0.5pt}}\\[4pt]\textcolor{#7c3aed}{\rule[0.45ex]{1.3em}{0.5pt}}&\textcolor{#7c3aed}{\langle\psi_2\rvert}&\textcolor{#7c3aed}{\rule[0.45ex]{1.3em}{0.5pt}}\\[4pt]&\vdots&\\[4pt]\textcolor{#db2777}{\rule[0.45ex]{1.3em}{0.5pt}}&\textcolor{#db2777}{\langle\psi_n\rvert}&\textcolor{#db2777}{\rule[0.45ex]{1.3em}{0.5pt}}\end{array}\right]\qquad (U^\dagger U)_{jk}=\langle\psi_j\vert\psi_k\rangle

For example, suppose two columns of the unitary matrix are the vectors ∣ψj⟩\lvert\psi_j\rangle and ∣ψk⟩\lvert\psi_k\rangle below. Taking the conjugate transpose of the first turns it into its bra (row) form ⟨ψj∣\langle\psi_j\rvert, and the inner product ⟨ψj∣ψk⟩\langle\psi_j\vert\psi_k\rangle is exactly the dot product of the two columns, with complex conjugation applied to the first vector:

∣ψj⟩=(a1a2⋮an),∣ψk⟩=(b1b2⋮bn)⟨ψj∣=(a1‾,…,an‾)⟨ψj∣ψk⟩=a1‾ b1+⋯+an‾ bn\lvert\psi_j\rangle=\begin{pmatrix}a_1\\ a_2\\ \vdots\\ a_n\end{pmatrix},\quad\lvert\psi_k\rangle=\begin{pmatrix}b_1\\ b_2\\ \vdots\\ b_n\end{pmatrix}\qquad\langle\psi_j\rvert=(\overline{a_1},\dots,\overline{a_n})\qquad\langle\psi_j\vert\psi_k\rangle=\overline{a_1}\,b_1+\cdots+\overline{a_n}\,b_n

There are two important cases:

  • If j=kj=k,
    ⟨ψj∣ψj⟩=∣a1∣2+∣a2∣2+⋯+∣an∣2=∥ψj∥2.\langle\psi_j\vert\psi_j\rangle=|a_1|^2+|a_2|^2+\cdots+|a_n|^2=\lVert\psi_j\rVert^2.
    So the diagonal entries of U†UU^\dagger U are the squared lengths of the columns.
  • If j≠kj\neq k, then ⟨ψj∣ψk⟩\langle\psi_j\vert\psi_k\rangle measures how closely the two vectors point in the same direction.

    For real vectors, ⟨ψj∣ψk⟩=∥ψj∥ ∥ψk∥cos⁡θ\langle\psi_j\vert\psi_k\rangle=\lVert\psi_j\rVert\,\lVert\psi_k\rVert\cos\theta, where θ\theta is the angle between the vectors. Thus, the inner product measures their overlap:

    • large magnitude means they point in similar directions,
    • zero means they are perpendicular (orthogonal).

Therefore U†UU^\dagger U is the matrix of all pairwise inner products between the columns of UU — each diagonal entry a squared length, each off-diagonal entry the overlap between two different columns:

U†U=(⟨ψ1∣ψ1⟩⟨ψ1∣ψ2⟩⋯⟨ψ1∣ψn⟩⟨ψ2∣ψ1⟩⟨ψ2∣ψ2⟩⋯⟨ψ2∣ψn⟩⋮⋮⋱⋮⟨ψn∣ψ1⟩⟨ψn∣ψ2⟩⋯⟨ψn∣ψn⟩)U^\dagger U=\begin{pmatrix}\textcolor{#0d9488}{\langle\psi_1\vert\psi_1\rangle}&\textcolor{#d97706}{\langle\psi_1\vert\psi_2\rangle}&\cdots&\textcolor{#d97706}{\langle\psi_1\vert\psi_n\rangle}\\[2pt]\textcolor{#d97706}{\langle\psi_2\vert\psi_1\rangle}&\textcolor{#0d9488}{\langle\psi_2\vert\psi_2\rangle}&\cdots&\textcolor{#d97706}{\langle\psi_2\vert\psi_n\rangle}\\[2pt]\vdots&\vdots&\ddots&\vdots\\[2pt]\textcolor{#d97706}{\langle\psi_n\vert\psi_1\rangle}&\textcolor{#d97706}{\langle\psi_n\vert\psi_2\rangle}&\cdots&\textcolor{#0d9488}{\langle\psi_n\vert\psi_n\rangle}\end{pmatrix}

If UU is unitary, then U†U=IU^\dagger U=I, and matching the entries of these two matrices gives

  • every diagonal entry equals 11: ⟨ψi∣ψi⟩=1\langle\psi_i\vert\psi_i\rangle=1, so every column has unit length;
  • every off-diagonal entry equals 00: ⟨ψi∣ψj⟩=0 (i≠j)\langle\psi_i\vert\psi_j\rangle=0\ (i\neq j), so every pair of distinct columns is orthogonal.

So, for a unitary matrix, its columns form an orthonormal set, and since there are exactly nn columns in an nn-dimensional space, an orthonormal set of nn vectors is automatically an orthonormal basis.

Applying the same argument to UU†=IUU^\dagger=I shows that the rows are also an orthonormal basis.

Projections

A square matrix Π\Pi is called a projection if it satisfies two properties:

  1. Π=Π†\Pi=\Pi^\dagger (it is Hermitian);
  2. Π2=Π\Pi^2=\Pi (it is idempotent).

Example: a single unit vector

For example, if ∣ψ⟩\lvert\psi\rangle is a unit vector, then its outer product with itself is a projection:

Π=∣ψ⟩⟨ψ∣.\Pi=\lvert\psi\rangle\langle\psi\rvert.

To confirm this, we check the two defining properties in turn.

  1. Hermitian — taking the conjugate transpose reverses the order of a product and daggers each factor, (AB)†=B†A†(AB)^\dagger=B^\dagger A^\dagger. Since (∣ψ⟩)†=⟨ψ∣(\lvert\psi\rangle)^\dagger=\langle\psi\rvert and (⟨ψ∣)†=∣ψ⟩(\langle\psi\rvert)^\dagger=\lvert\psi\rangle, the two factors swap back into their original places:
    Π†=(∣ψ⟩⟨ψ∣)†=(⟨ψ∣)†(∣ψ⟩)†=∣ψ⟩⟨ψ∣=Π.\Pi^\dagger=(\lvert\psi\rangle\langle\psi\rvert)^\dagger=(\langle\psi\rvert)^\dagger(\lvert\psi\rangle)^\dagger=\lvert\psi\rangle\langle\psi\rvert=\Pi.
  2. Idempotent — applying Π\Pi twice leaves the inner product ⟨ψ∣ψ⟩\langle\psi\vert\psi\rangle in the middle, and because ∣ψ⟩\lvert\psi\rangle is a unit vector that factor equals ⟨ψ∣ψ⟩=1\langle\psi\vert\psi\rangle=1, so the product collapses back to a single copy:
    Π2=(∣ψ⟩⟨ψ∣)2=∣ψ⟩⟨ψ∣ψ⟩⟨ψ∣=∣ψ⟩⟨ψ∣=Π.\Pi^2=(\lvert\psi\rangle\langle\psi\rvert)^2=\lvert\psi\rangle\langle\psi\vert\psi\rangle\langle\psi\rvert=\lvert\psi\rangle\langle\psi\rvert=\Pi.

Both properties hold, so Π=∣ψ⟩⟨ψ∣\Pi=\lvert\psi\rangle\langle\psi\rvert is indeed a projection.

Visual demo: projecting onto the line

With a single unit vector ∣ψ⟩\lvert\psi\rangle, the projection operator Π=∣ψ⟩⟨ψ∣\Pi=\lvert\psi\rangle\langle\psi\rvert takes any vector ∣v⟩\lvert v\rangle and projects it onto the line spanned by ∣ψ⟩\lvert\psi\rangle. The projection is Π∣v⟩=∣ψ⟩⟨ψ∣v⟩\Pi\lvert v\rangle=\lvert\psi\rangle\langle\psi\vert v\rangle, where ⟨ψ∣v⟩\langle\psi\vert v\rangle is the scalar giving the component of ∣v⟩\lvert v\rangle along ∣ψ⟩\lvert\psi\rangle. Thus Π∣v⟩\Pi\lvert v\rangle always lies on the line, while the remaining vector ∣v⟩−Π∣v⟩\lvert v\rangle-\Pi\lvert v\rangle is perpendicular to it. Applying the projection a second time changes nothing, because Π∣v⟩\Pi\lvert v\rangle is already on the line, so Π2=Π\Pi^2=\Pi.

vψΠv
ψ = (0.94, 0.34)
v = (1.55, 1.35)
⟨ψ|v⟩ = 1.92
Πv = ⟨ψ|v⟩·ψ = (1.80, 0.66)
Drag v (violet) or rotate ψ (blue). Πv is the foot of the perpendicular — the closest point on the line.

Generalization: an orthonormal set

This generalizes from a single vector to any orthonormal set. If {∣ψ1⟩,…,∣ψm⟩}\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\} is orthonormal, then the sum of their outer products is again a projection:

Π=∑k=1m∣ψk⟩⟨ψk∣.\Pi=\sum_{k=1}^{m}\lvert\psi_k\rangle\langle\psi_k\rvert.

The same two checks go through, now using orthonormality ⟨ψj∣ψk⟩=δjk\langle\psi_j\vert\psi_k\rangle=\delta_{jk} (equal to 11 when j=kj=k and 00 otherwise).

  1. Hermitian — the dagger passes through the sum and daggers each term, and every term ∣ψk⟩⟨ψk∣\lvert\psi_k\rangle\langle\psi_k\rvert is Hermitian by the argument above:
    Π†=(∑k=1m∣ψk⟩⟨ψk∣)†=∑k=1m(∣ψk⟩⟨ψk∣)†=∑k=1m∣ψk⟩⟨ψk∣=Π.\Pi^\dagger=\Bigl(\sum_{k=1}^{m}\lvert\psi_k\rangle\langle\psi_k\rvert\Bigr)^\dagger=\sum_{k=1}^{m}(\lvert\psi_k\rangle\langle\psi_k\rvert)^\dagger=\sum_{k=1}^{m}\lvert\psi_k\rangle\langle\psi_k\rvert=\Pi.
  2. Idempotent — multiplying the two sums gives a double sum whose inner factor is ⟨ψj∣ψk⟩\langle\psi_j\vert\psi_k\rangle. Orthonormality removes every cross term (j≠kj\neq k) and leaves 11 on the diagonal (j=kj=k), collapsing the double sum back to a single one:
    Π2=∑j=1m∑k=1m∣ψj⟩⟨ψj∣ψk⟩⟨ψk∣=∑k=1m∣ψk⟩⟨ψk∣=Π.\Pi^2=\sum_{j=1}^{m}\sum_{k=1}^{m}\lvert\psi_j\rangle\langle\psi_j\vert\psi_k\rangle\langle\psi_k\rvert=\sum_{k=1}^{m}\lvert\psi_k\rangle\langle\psi_k\rvert=\Pi.

So any sum of outer products over an orthonormal set is a projection — geometrically, it projects onto the subspace those vectors span.

Projective measurements

A collection of projections {Π1,…,Πm}\{\Pi_1,\dots,\Pi_m\} that satisfies Π1+⋯+Πm=I\Pi_1+\dots+\Pi_m=I describes a projective measurement.

When such a measurement is performed on a system in the state ∣ψ⟩\lvert\psi\rangle, two things happen:

  1. The outcome k∈{1,…,m}k\in\{1,\dots,m\} of the measurement is chosen randomly:
    Pr⁡(outcome is k)=∥Πk∣ψ⟩∥2=⟨Πkψ∣Πkψ⟩=⟨ψ∣Πk†Πk∣ψ⟩=⟨ψ∣ΠkΠk∣ψ⟩=⟨ψ∣Πk∣ψ⟩.\begin{aligned}\Pr(\text{outcome is }k)&=\bigl\lVert\Pi_k\lvert\psi\rangle\bigr\rVert^2\\&=\langle\Pi_k\psi\vert\Pi_k\psi\rangle\\&=\langle\psi\rvert\Pi_k^\dagger\Pi_k\lvert\psi\rangle\\&=\langle\psi\rvert\Pi_k\Pi_k\lvert\psi\rangle\\&=\langle\psi\rvert\Pi_k\lvert\psi\rangle.\end{aligned}

    The two forms ∥Πk∣ψ⟩∥2\lVert\Pi_k\lvert\psi\rangle\rVert^2 and ⟨ψ∣Πk∣ψ⟩\langle\psi\rvert\Pi_k\lvert\psi\rangle are mathematically identical. The first makes the geometry — the squared projection length — much more obvious, while the second connects naturally to the general framework of expectation values.

  2. The state of the system becomes
    Πk∣ψ⟩∥Πk∣ψ⟩∥.\dfrac{\Pi_k\lvert\psi\rangle}{\lVert\Pi_k\lvert\psi\rangle\rVert}.

The outcomes need not be labelled 1,…,m1,\dots,m. We are free to name them however is convenient — letters a,b,c,…a,b,c,\dots, or any index set Γ\Gamma, so that a family {Πa:a∈Γ}\{\Pi_a:a\in\Gamma\} with ∑a∈ΓΠa=I\sum_{a\in\Gamma}\Pi_a=I describes a projective measurement with outcomes in Γ\Gamma. The rules are exactly the same.

Visual demo: measuring a qutrit in the standard basis

Take the three rank-one projections Πk=∣k⟩⟨k∣\Pi_k=\lvert k\rangle\langle k\rvert onto the basis axes ∣0⟩,∣1⟩,∣2⟩\lvert0\rangle,\lvert1\rangle,\lvert2\rangle. They add up to the identity, Π0+Π1+Π2=I\Pi_0+\Pi_1+\Pi_2=I, so they form a projective measurement:

Π0+Π1+Π2=(100000000)+(000010000)+(000000001)=(100010001)=I.\Pi_0+\Pi_1+\Pi_2=\begin{pmatrix}1&0&0\\0&0&0\\0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0\\0&1&0\\0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0\\0&0&0\\0&0&1\end{pmatrix}=\begin{pmatrix}1&0&0\\0&1&0\\0&0&1\end{pmatrix}=I.

If outcome kk occurs, the state ∣ψ⟩\lvert\psi\rangle is projected onto the corresponding subspace, and the probability pk=⟨ψ∣Πk∣ψ⟩=∣αk∣2p_k=\langle\psi\rvert\Pi_k\lvert\psi\rangle=\lvert\alpha_k\rvert^2 is the squared length of the projected vector — so the squared projection lengths (the probabilities) always sum to 1.

|0⟩|1⟩|2⟩Π₀|ψ⟩Π₁|ψ⟩Π₂|ψ⟩|ψ⟩
‖Π₀|ψ⟩‖
0.59² = 0.35
‖Π₁|ψ⟩‖
0.46² = 0.21
‖Π₂|ψ⟩‖
0.66² = 0.43
‖Π₀|ψ⟩‖² + ‖Π₁|ψ⟩‖² + ‖Π₂|ψ⟩‖² = 1.00
|ψ⟩ = 0.59|0⟩ + 0.46|1⟩ + 0.66|2⟩
Press Measure to sample an outcome. Drag the dark handle to move the state; drag the background to rotate the view.

Measuring one subsystem

Suppose we have two systems, XX and YY, but we only measure XX in the standard basis. We don't need a new mathematical procedure for measuring composite systems — just use the projectors.

{∣a⟩⟨a∣⊗IY:a∈Σ}\{\textcolor{#2563eb}{\lvert a\rangle\langle a\rvert\otimes I_Y}:a\in\Sigma\}

Each projector acts on the pair (X,Y)(X,Y) in two independent steps:

  • ∣a⟩⟨a∣\textcolor{#2563eb}{\lvert a\rangle\langle a\rvert} checks whether X=aX=a — it keeps the part of the state carrying that outcome and discards the rest.
  • IY\textcolor{#2563eb}{I_Y} does nothing to YY — the second system is left exactly as it was.

Since these projectors sum to the identity, they form a valid projective measurement:

∑a∈Σ(∣a⟩⟨a∣⊗IY)=(∑a∈Σ∣a⟩⟨a∣)⊗IY=IX⊗IY=I.\sum_{a\in\Sigma}\textcolor{#2563eb}{\bigl(\lvert a\rangle\langle a\rvert\otimes I_Y\bigr)}=\Bigl(\sum_{a\in\Sigma}\lvert a\rangle\langle a\rvert\Bigr)\otimes I_Y=I_X\otimes I_Y=I.

The projective measurement rules apply as before:

Pr⁡(outcome is a)=∥Πa∣ψ⟩∥2=∥(∣a⟩⟨a∣⊗IY)∣ψ⟩∥2,\Pr(\text{outcome is }a)=\bigl\lVert\textcolor{#2563eb}{\Pi_a}\lvert\psi\rangle\bigr\rVert^2=\bigl\lVert\textcolor{#2563eb}{(\lvert a\rangle\langle a\rvert\otimes I_Y)}\lvert\psi\rangle\bigr\rVert^2,

and after observing outcome aa, the state of (X,Y)(X,Y) becomes

∣ψ⟩  ⟶  Πa∣ψ⟩∥Πa∣ψ⟩∥=(∣a⟩⟨a∣⊗IY)∣ψ⟩∥(∣a⟩⟨a∣⊗IY)∣ψ⟩∥.\lvert\psi\rangle\;\longrightarrow\;\dfrac{\textcolor{#2563eb}{\Pi_a}\lvert\psi\rangle}{\bigl\lVert\textcolor{#2563eb}{\Pi_a}\lvert\psi\rangle\bigr\rVert}=\dfrac{\textcolor{#2563eb}{(\lvert a\rangle\langle a\rvert\otimes I_Y)}\lvert\psi\rangle}{\bigl\lVert\textcolor{#2563eb}{(\lvert a\rangle\langle a\rvert\otimes I_Y)}\lvert\psi\rangle\bigr\rVert}.

Example — subsystem measurement by projection

For example, suppose (X,Y)(X,Y) is a pair of qubits in the state below, written so that each term is grouped by XX:

∣ψ⟩=∣0⟩⊗(12 ∣0⟩+12 ∣1⟩)+∣1⟩⊗(i22 ∣0⟩−122 ∣1⟩)\lvert\psi\rangle=\textcolor{#AB6108}{\lvert0\rangle}\otimes\textcolor{#0A74A9}{\Bigl(\tfrac{1}{\sqrt{2}}\,\lvert0\rangle+\tfrac{1}{2}\,\lvert1\rangle\Bigr)}+\textcolor{#AB6108}{\lvert1\rangle}\otimes\textcolor{#0A74A9}{\Bigl(\tfrac{i}{2\sqrt{2}}\,\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\,\lvert1\rangle\Bigr)}

Outcome a=0a=0

The projector for this outcome is

Π0=∣0⟩⟨0∣⊗IY.\textcolor{#2563eb}{\Pi_0}=\textcolor{#2563eb}{\lvert0\rangle\langle0\rvert\otimes I_Y}.

Apply it to the full state and expand. Each factor is handled separately, so the X\textcolor{#AB6108}{X} kets meet ⟨0∣\textcolor{#2563eb}{\langle0\rvert} while the identity IY\textcolor{#2563eb}{I_Y} leaves the Y\textcolor{#0A74A9}{Y} part unchanged:

Π0∣ψ⟩=(∣0⟩⟨0∣⊗IY)[∣0⟩⊗(12∣0⟩+12∣1⟩)+∣1⟩⊗(i22∣0⟩−122∣1⟩)]=(∣0⟩⟨0∣0⟩⏟= 1)⊗IY⏟no effect(12∣0⟩+12∣1⟩)+(∣0⟩⟨0∣1⟩⏟= 0)⊗IY⏟no effect(i22∣0⟩−122∣1⟩)=∣0⟩⊗(12∣0⟩+12∣1⟩).\begin{aligned}\textcolor{#2563eb}{\Pi_0}\lvert\psi\rangle&=\textcolor{#2563eb}{\bigl(\lvert0\rangle\langle0\rvert\otimes I_Y\bigr)}\Bigl[\textcolor{#AB6108}{\lvert0\rangle}\otimes\textcolor{#0A74A9}{\bigl(\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{2}\lvert1\rangle\bigr)}+\textcolor{#AB6108}{\lvert1\rangle}\otimes\textcolor{#0A74A9}{\bigl(\tfrac{i}{2\sqrt{2}}\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\lvert1\rangle\bigr)}\Bigr]\\[6pt]&=\Bigl(\textcolor{#2563eb}{\lvert0\rangle}\underbrace{\textcolor{#2563eb}{\langle0\vert}\textcolor{#AB6108}{0\rangle}}_{\textcolor{#16a34a}{=\,1}}\Bigr)\otimes\underbrace{\textcolor{#2563eb}{I_Y}}_{\text{no effect}}\textcolor{#0A74A9}{\bigl(\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{2}\lvert1\rangle\bigr)}+\Bigl(\textcolor{#2563eb}{\lvert0\rangle}\underbrace{\textcolor{#2563eb}{\langle0\vert}\textcolor{#AB6108}{1\rangle}}_{\textcolor{#dc2626}{=\,0}}\Bigr)\otimes\underbrace{\textcolor{#2563eb}{I_Y}}_{\text{no effect}}\textcolor{#0A74A9}{\bigl(\tfrac{i}{2\sqrt{2}}\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\lvert1\rangle\bigr)}\\[6pt]&=\textcolor{#2563eb}{\lvert0\rangle}\otimes\textcolor{#0A74A9}{\bigl(\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{2}\lvert1\rangle\bigr)}.\end{aligned}

The probability is the squared norm of what survived — the sum of its squared amplitude magnitudes:

∥Π0∣ψ⟩∥2=∣12∣2+∣12∣2=12+14=34\bigl\lVert\textcolor{#2563eb}{\Pi_0}\lvert\psi\rangle\bigr\rVert^2=\left|\tfrac{1}{\sqrt{2}}\right|^2+\left|\tfrac{1}{2}\right|^2=\tfrac{1}{2}+\tfrac{1}{4}=\tfrac{3}{4}

The survivor is not a unit vector, so divide it by its norm 34=32\sqrt{\tfrac{3}{4}}=\tfrac{\sqrt{3}}{2}:

∣0⟩⊗(12∣0⟩+12∣1⟩)3/4=∣0⟩⊗(23 ∣0⟩+13 ∣1⟩)\frac{\lvert0\rangle\otimes\Bigl(\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{2}\lvert1\rangle\Bigr)}{\sqrt{3/4}}=\lvert0\rangle\otimes\Bigl(\sqrt{\tfrac{2}{3}}\,\lvert0\rangle+\tfrac{1}{\sqrt{3}}\,\lvert1\rangle\Bigr)

Outcome a=1a=1

The other projector does the mirror image: now ⟨1∣\textcolor{#2563eb}{\langle1\rvert} keeps the ∣1⟩\lvert1\rangle branch and discards the ∣0⟩\lvert0\rangle one:

(∣1⟩⟨1∣⊗IY)∣ψ⟩=∣1⟩⊗(i22∣0⟩−122∣1⟩)\textcolor{#2563eb}{(\lvert1\rangle\langle1\rvert\otimes I_Y)}\lvert\psi\rangle=\lvert1\rangle\otimes\Bigl(\tfrac{i}{2\sqrt{2}}\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\lvert1\rangle\Bigr)

Its squared norm is the remaining probability. The ii drops out under ∣⋅∣2\lvert\cdot\rvert^2 — only magnitudes matter:

∥(∣1⟩⟨1∣⊗IY)∣ψ⟩∥2=∣i22∣2+∣−122∣2=18+18=14\bigl\lVert\textcolor{#2563eb}{(\lvert1\rangle\langle1\rvert\otimes I_Y)}\lvert\psi\rangle\bigr\rVert^2=\left|\tfrac{i}{2\sqrt{2}}\right|^2+\left|-\tfrac{1}{2\sqrt{2}}\right|^2=\tfrac{1}{8}+\tfrac{1}{8}=\tfrac{1}{4}

Divide by the norm 14=12\sqrt{\tfrac{1}{4}}=\tfrac{1}{2} as before:

∣1⟩⊗(i22∣0⟩−122∣1⟩)1/4=∣1⟩⊗(i2 ∣0⟩−12 ∣1⟩)\frac{\lvert1\rangle\otimes\Bigl(\tfrac{i}{2\sqrt{2}}\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\lvert1\rangle\Bigr)}{\sqrt{1/4}}=\lvert1\rangle\otimes\Bigl(\tfrac{i}{\sqrt{2}}\,\lvert0\rangle-\tfrac{1}{\sqrt{2}}\,\lvert1\rangle\Bigr)

Implementing projective measurements

So far the projectors have been pure mathematics. Hardware offers only two things: unitary operations and standard basis measurements. That is already enough — any projective measurement can be assembled out of those two.

The trick is to bring in one extra system alongside the one being measured, with a classical state for each possible outcome. A single unitary then files each outcome into its own branch: the extra system carries the outcome's label, and travelling alongside it is the matching projected state of the measured system. Reading the extra system in the standard basis picks one branch — with exactly the probability the projection rule demands — and leaves the measured system in that branch's projected, renormalised state. No projector is ever built as hardware; the projections emerge from the branch structure.

Below, that idea runs on a pair of systems XX and YY. The bottom wire is the extra qubit — the only thing ever read.

XY∣0⟩\lvert0\rangle

The first Hadamard splits the extra qubit into two branches. The controlled-SWAP exchanges XX and YY on one branch only, so the two branches now hold the state seen from two different angles. The second Hadamard makes them interfere, which sorts the state into a part unchanged by the exchange and a part that the exchange reverses — and those two parts are precisely the projections. Measuring the extra qubit says which one you landed in, and turns it into a classical bit.

Three gates and a single qubit read out: an abstract pair of projections has become something a machine actually does.

Limitations of quantum measurements

Irrelevance of global phases

Quantum amplitudes are complex numbers, so each one carries both a magnitude and a phase. Phases show up in two different ways:

  • A global phase multiplies every amplitude in a state by the same unit-modulus factor eiθe^{i\theta}.
  • A relative phase changes the phase difference between the amplitudes within a superposition.

The distinction is fundamental. A global phase rotates the entire state vector by the same amount, leaving the relationships between its amplitudes unchanged. A relative phase changes those relationships and therefore affects interference.

A scalar factor in a global phase must preserve the norm, so ∣α∣=1\lvert\alpha\rvert=1. That condition is met exactly by the complex numbers of the form α=eiθ\alpha=e^{i\theta} for some real θ\theta: every such number sits on the unit circle of the complex plane, and every unit-modulus number can be written this way. Two state vectors therefore differ by a global phase when ∣ϕ⟩=eiθ∣ψ⟩\lvert\phi\rangle=e^{i\theta}\lvert\psi\rangle — every amplitude is multiplied by the same unit-modulus factor eiθe^{i\theta}.

For example, the two amplitudes of ∣ψ⟩=α0∣0⟩+α1∣1⟩\lvert\psi\rangle=\alpha_0\lvert0\rangle+\alpha_1\lvert1\rangle are drawn as arrows on the complex plane. Their lengths are the magnitudes, their angles the phases.

Drag global phase — every bar holds still, so the state is physically unchanged. Drag relative phase — the standard basis still won't move, but the ± measurement swings: the relative phase is observable. The amplitude split resizes the arrows and shifts the standard-basis odds.

ReImα₀α₁
0°
90°
0.71 · 0.71

Probability of each outcome if we measure ∣ψ⟩\lvert\psi\rangle in that basis:

Standard basis{∣0⟩,∣1⟩}\{\lvert0\rangle,\lvert1\rangle\}

P(0)=∣⟨0∣ψ⟩∣2P(0)=\lvert\langle0\vert\psi\rangle\rvert^250%
P(1)=∣⟨1∣ψ⟩∣2P(1)=\lvert\langle1\vert\psi\rangle\rvert^250%

Hadamard basis{∣+⟩,∣−⟩}\{\lvert+\rangle,\lvert-\rangle\}

P(+)=∣⟨+∣ψ⟩∣2P(+)=\lvert\langle+\vert\psi\rangle\rvert^250%
P(−)=∣⟨−∣ψ⟩∣2P(-)=\lvert\langle-\vert\psi\rangle\rvert^250%

Mathematically, two state vectors differ by a global phase if ∣ϕ⟩=eiθ∣ψ⟩\lvert\phi\rangle=e^{i\theta}\lvert\psi\rangle.

Standard-basis measurement. The probability of outcome aa is the squared magnitude of the amplitude ⟨a∣ϕ⟩\langle a\vert\phi\rangle. Substituting ∣ϕ⟩=eiθ∣ψ⟩\lvert\phi\rangle=e^{i\theta}\lvert\psi\rangle, the phase factors out and its magnitude ∣eiθ∣2=1\lvert e^{i\theta}\rvert^2=1 disappears:

∣⟨a∣ϕ⟩∣2=∣eiθ ⟨a∣ψ⟩∣2=∣eiθ∣2 ∣⟨a∣ψ⟩∣2=∣⟨a∣ψ⟩∣2.\bigl\lvert\langle a\vert\phi\rangle\bigr\rvert^2=\bigl\lvert e^{i\theta}\,\langle a\vert\psi\rangle\bigr\rvert^2=\lvert e^{i\theta}\rvert^2\,\bigl\lvert\langle a\vert\psi\rangle\bigr\rvert^2=\bigl\lvert\langle a\vert\psi\rangle\bigr\rvert^2.

Projective measurement. The same cancellation holds for any projective measurement {Π1,…,Πm}\{\Pi_1,\dots,\Pi_m\}. Each outcome probability is the squared norm ∥Πk∣ϕ⟩∥2\lVert\Pi_k\lvert\phi\rangle\rVert^2, and pulling the scalar eiθe^{i\theta} out of the norm leaves the factor ∣eiθ∣2=1\lvert e^{i\theta}\rvert^2=1 again:

∥Πk∣ϕ⟩∥2=∥eiθ Πk∣ψ⟩∥2=∣eiθ∣2 ∥Πk∣ψ⟩∥2=∥Πk∣ψ⟩∥2.\bigl\lVert\Pi_k\lvert\phi\rangle\bigr\rVert^2=\bigl\lVert e^{i\theta}\,\Pi_k\lvert\psi\rangle\bigr\rVert^2=\lvert e^{i\theta}\rvert^2\,\bigl\lVert\Pi_k\lvert\psi\rangle\bigr\rVert^2=\bigl\lVert\Pi_k\lvert\psi\rangle\bigr\rVert^2.

Therefore all measurements produce exactly the same statistics for ∣ψ⟩\lvert\psi\rangle and eiθ∣ψ⟩e^{i\theta}\lvert\psi\rangle. States that differ by a global phase are considered equivalent — they represent the same physical state. A relative phase, by contrast, changes the interference between amplitudes and is observable, as the demo above shows.

A note on representation. This global phase is a degeneracy of the state-vector picture: the same physical state maps to a whole circle of vectors {eiθ∣ψ⟩}\{e^{i\theta}\lvert\psi\rangle\} that are all indistinguishable. It is an artifact of describing states at this simplified level of generality (it's called “simplified,” though personally I cried at the word “simple”).

The more general formalism uses density matrices: a state vector ∣ψ⟩\lvert\psi\rangle is replaced by the operator ρ=∣ψ⟩⟨ψ∣\rho=\lvert\psi\rangle\langle\psi\rvert. Here the global phase cancels automatically, since (eiθ∣ψ⟩)(eiθ∣ψ⟩)†=eiθe−iθ ∣ψ⟩⟨ψ∣=ρ\bigl(e^{i\theta}\lvert\psi\rangle\bigr)\bigl(e^{i\theta}\lvert\psi\rangle\bigr)^{\dagger}=e^{i\theta}e^{-i\theta}\,\lvert\psi\rangle\langle\psi\rvert=\rho, so equivalent states share exactly one density matrix and the redundancy disappears. Density matrices also describe mixed states (classical uncertainty over several vectors), which no single state vector can capture.

No-cloning theorem

Copying classical information is trivial — you read a bit and write it down twice. It is natural to ask whether a quantum computer can do the same for an unknown state ∣ψ⟩\lvert\psi\rangle, producing two identical copies of it. The no-cloning theorem says this is impossible: no single unitary can duplicate an arbitrary unknown quantum state, which is exactly why quantum information cannot simply be copied and why quantum key distribution is secure.

Formally, let XX and YY both have the classical state set {0,…,d−1}\{0,\dots,d-1\} with d≥2d\ge 2. A cloner would be a unitary UU on the pair (X,Y)(X, Y) that takes the state in XX with a blank ∣0⟩\lvert0\rangle in YY and writes a copy into YY. The theorem states no such UU exists:

∀ ∣ψ⟩:U(∣ψ⟩⊗∣0⟩)=∣ψ⟩⊗∣ψ⟩.\forall\,\lvert\psi\rangle:\quad U\bigl(\lvert\psi\rangle\otimes\lvert0\rangle\bigr)=\lvert\psi\rangle\otimes\lvert\psi\rangle.

Drawn as a circuit, this cloner would feed ∣ψ⟩\lvert\psi\rangle and a blank register ∣0⋯0⟩\lvert0\cdots0\rangle into UU and read out two copies — the box below that cannot exist for every input:

∣0⋯0⟩\lvert0\cdots0\rangle∣ψ⟩\lvert\psi\rangle∣ψ⟩\lvert\psi\rangle∣ψ⟩\lvert\psi\rangle
U
No such U exists

Why no such U can exist

Suppose there existed a unitary operator UU that could perfectly clone any quantum state. For every state ∣ψ⟩\textcolor{#2563eb}{\lvert\psi\rangle}, it would satisfy U(∣ψ⟩⊗∣0⟩)=∣ψ⟩⊗∣ψ⟩U\bigl(\textcolor{#2563eb}{\lvert\psi\rangle}\otimes\lvert0\rangle\bigr)=\textcolor{#2563eb}{\lvert\psi\rangle}\otimes\textcolor{#2563eb}{\lvert\psi\rangle}, where ∣0⟩\lvert0\rangle is a blank qubit that receives the copy.

Applying UU to basis states results in: U(∣0⟩∣0⟩)=∣0⟩∣0⟩,U(∣1⟩∣0⟩)=∣1⟩∣1⟩.U\bigl(\textcolor{#2563eb}{\lvert0\rangle}\lvert0\rangle\bigr)=\textcolor{#2563eb}{\lvert0\rangle}\textcolor{#2563eb}{\lvert0\rangle},\quad U\bigl(\textcolor{#2563eb}{\lvert1\rangle}\lvert0\rangle\bigr)=\textcolor{#2563eb}{\lvert1\rangle}\textcolor{#2563eb}{\lvert1\rangle}.

But now consider a plus (superposition) state ∣+⟩=∣0⟩+∣1⟩2\lvert+\rangle=\dfrac{\lvert0\rangle+\lvert1\rangle}{\sqrt2}. Because every quantum gate is linear, the cloning operation must satisfy:

U(∣+⟩∣0⟩)=U ⁣(∣0⟩+∣1⟩2⊗∣0⟩)=12 U(∣0⟩∣0⟩)+12 U(∣1⟩∣0⟩)=12 ∣0⟩∣0⟩+12 ∣1⟩∣1⟩=∣00⟩+∣11⟩2.\begin{aligned} U\bigl(\lvert+\rangle\lvert0\rangle\bigr) &=U\!\left(\dfrac{\lvert0\rangle+\lvert1\rangle}{\sqrt2}\otimes\lvert0\rangle\right)\\[4pt] &=\dfrac{1}{\sqrt2}\,U\bigl(\textcolor{#2563eb}{\lvert0\rangle}\lvert0\rangle\bigr)+\dfrac{1}{\sqrt2}\,U\bigl(\textcolor{#2563eb}{\lvert1\rangle}\lvert0\rangle\bigr)\\[4pt] &=\dfrac{1}{\sqrt2}\,\textcolor{#2563eb}{\lvert0\rangle}\textcolor{#2563eb}{\lvert0\rangle}+\dfrac{1}{\sqrt2}\,\textcolor{#2563eb}{\lvert1\rangle}\textcolor{#2563eb}{\lvert1\rangle}\\[4pt] &=\dfrac{\lvert00\rangle+\lvert11\rangle}{\sqrt2}. \end{aligned}

But if UU were truly a cloning machine, the output should instead be

U(∣+⟩∣0⟩)=∣+⟩∣+⟩.U\bigl(\textcolor{#2563eb}{\lvert+\rangle}\lvert0\rangle\bigr)=\textcolor{#2563eb}{\lvert+\rangle}\textcolor{#2563eb}{\lvert+\rangle}.

This expands to

∣+⟩∣+⟩=(∣0⟩+∣1⟩2)⊗(∣0⟩+∣1⟩2)=∣00⟩+∣01⟩+∣10⟩+∣11⟩2.\lvert+\rangle\lvert+\rangle=\left(\dfrac{\lvert0\rangle+\lvert1\rangle}{\sqrt2}\right)\otimes\left(\dfrac{\lvert0\rangle+\lvert1\rangle}{\sqrt2}\right)=\dfrac{\lvert00\rangle+\lvert01\rangle+\lvert10\rangle+\lvert11\rangle}{2}.

These two states are different:

∣00⟩+∣11⟩2  ≠  ∣00⟩+∣01⟩+∣10⟩+∣11⟩2.\dfrac{\lvert00\rangle+\lvert11\rangle}{\sqrt2}\;\neq\;\dfrac{\lvert00\rangle+\lvert01\rangle+\lvert10\rangle+\lvert11\rangle}{2}.

The contradiction arises because linearity forces one output, while perfect cloning requires another. Therefore, no unitary operation can perfectly clone an arbitrary unknown quantum state.

Remarks

  • Approximate forms of the cloning theorem are known.
  • Copying a standard basis state is possible — the no-cloning theorem does not contradict this.

    For example, a CNOT\mathrm{CNOT} controlled by a basis value ∣a⟩\lvert a\rangle copies it into a blank ∣0⟩\lvert0\rangle, giving ∣a⟩∣a⟩\lvert a\rangle\lvert a\rangle:

    ∣0⟩\lvert0\rangle∣a⟩\lvert a\rangle∣a⟩\lvert a\rangle∣a⟩\lvert a\rangle
  • Cloning a probabilistic state (classically) is also impossible.
  • Perfect clones are possible if they are all encrypted. A 2026 protocol, encrypted cloning, deterministically produces any number of perfect copies of an unknown state — as long as the copies are simultaneously locked with a single-use quantum decryption key. Decrypting one clone consumes the key and renders every other clone indecipherable, so no two usable copies ever coexist and the theorem still holds. The real constraint is not the copying, but that the decryption mechanism must be single-use. The payoff is redundancy, like keeping backups: you hold many encrypted copies and recover the original from any one that survives. It has been demonstrated on IBM Heron-R2 hardware with up to 154 qubits (Yamaguchi et al., 2026).

Discriminating non-orthogonal states

It is not possible to perfectly discriminate two non-orthogonal quantum states. Equivalently, if we can discriminate two quantum states perfectly, then they must be orthogonal.

Two states ∣ψ⟩\lvert\psi\rangle and ∣ϕ⟩\lvert\phi\rangle can be discriminated perfectly if there is a unitary operation UU that works like this:

∣π0⟩\lvert\pi_0\rangle0
U(∣0⋯0⟩∣ψ⟩)=∣π0⟩∣0⟩U\bigl(\lvert0\cdots0\rangle\lvert\psi\rangle\bigr)=\lvert\pi_0\rangle\lvert0\rangle
∣π1⟩\lvert\pi_1\rangle1
U(∣0⋯0⟩∣ϕ⟩)=∣π1⟩∣1⟩U\bigl(\lvert0\cdots0\rangle\lvert\phi\rangle\bigr)=\lvert\pi_1\rangle\lvert1\rangle

Suppose a unitary operator UU perfectly distinguishes two states ∣ψ⟩\lvert\psi\rangle and ∣ϕ⟩\lvert\phi\rangle. By definition,

U(∣0⋯0⟩∣ψ⟩)=∣π0⟩∣0⟩,U(∣0⋯0⟩∣ϕ⟩)=∣π1⟩∣1⟩,\begin{aligned} U\bigl(\lvert0\cdots0\rangle\lvert\psi\rangle\bigr)&=\lvert\pi_0\rangle\lvert0\rangle,\\[4pt] U\bigl(\lvert0\cdots0\rangle\lvert\phi\rangle\bigr)&=\lvert\pi_1\rangle\lvert1\rangle, \end{aligned}

where the final qubit stores the measurement result and the remaining qubits ∣π0⟩\lvert\pi_0\rangle and ∣π1⟩\lvert\pi_1\rangle represent arbitrary ancilla states.

The overlap of two states is their inner product ⟨a∣b⟩\langle a\vert b\rangle— a number that measures how similar they are. It is 11 for identical states and 00 for orthogonal (perfectly distinguishable) ones.

Now take two states ∣a⟩\lvert a\rangle and ∣b⟩\lvert b\rangle and apply UU to each. To form the overlap of the outputs, the first ket U∣a⟩U\lvert a\rangle becomes a bra by taking its conjugate transpose, which flips UU into U†U^\dagger:

⟨Ua∣Ub⟩=⟨a∣ U†U ∣b⟩=⟨a∣ I ∣b⟩=⟨a∣b⟩,\langle Ua\vert Ub\rangle=\langle a\rvert\,U^\dagger U\,\lvert b\rangle=\langle a\rvert\,I\,\lvert b\rangle=\langle a\vert b\rangle,

where the middle U†UU^\dagger U collapses to the identity II because UU is unitary. So passing both states through the same UU leaves their overlap unchanged — the input overlap equals the output overlap.

For the discriminating unitary, the two input vectors are ∣0⋯0⟩∣ψ⟩\lvert0\cdots0\rangle\lvert\psi\rangle and ∣0⋯0⟩∣ϕ⟩\lvert0\cdots0\rangle\lvert\phi\rangle. Their overlap factorizes across the ancilla and state registers, and because the ancilla register is the same in both inputs ⟨0⋯0∣0⋯0⟩=1\langle0\cdots0\vert0\cdots0\rangle=1:

(⟨0⋯0∣⟨ψ∣)(∣0⋯0⟩∣ϕ⟩)=⟨0⋯0∣0⋯0⟩⋅⟨ψ∣ϕ⟩=⟨ψ∣ϕ⟩.\bigl(\langle0\cdots0\rvert\langle\psi\rvert\bigr)\bigl(\lvert0\cdots0\rangle\lvert\phi\rangle\bigr)=\langle0\cdots0\vert0\cdots0\rangle\cdot\langle\psi\vert\phi\rangle=\langle\psi\vert\phi\rangle.

The output overlap factorizes the same way. The measurement outcomes are different, so ⟨0∣1⟩=0\langle0\vert1\rangle=0, and the whole overlap collapses to zero:

(⟨π0∣⟨0∣)(∣π1⟩∣1⟩)=⟨π0∣π1⟩ ⟨0∣1⟩=⟨π0∣π1⟩⋅0=0.\bigl(\langle\pi_0\rvert\langle0\rvert\bigr)\bigl(\lvert\pi_1\rangle\lvert1\rangle\bigr)=\langle\pi_0\vert\pi_1\rangle\,\langle0\vert1\rangle=\langle\pi_0\vert\pi_1\rangle\cdot0=0.

Equating the input and output overlaps, and recalling the output overlap is zero, gives

  ⟨ψ∣ϕ⟩=0.  \boxed{\;\langle\psi\vert\phi\rangle=0.\;}

In other words, perfect discrimination is possible only for orthogonal quantum states.

The demo illustrates this result visually. Drag either state to change their overlap ⟨ψ∣ϕ⟩\langle\psi\vert\phi\rangle. As the overlap decreases, the states become easier to distinguish. When ⟨ψ∣ϕ⟩=0\langle\psi\vert\phi\rangle=0, they are orthogonal and can be distinguished perfectly.

60°ψϕ
θ = 60°
⟨ψ|ϕ⟩ = cos θ = 0.50
best distinguishing probability0.93
0.5 · coin flip1.0 · perfect
Non-orthogonal — the overlap is nonzero, so no measurement can tell the two apart with certainty. Drag the arrows 90° apart.

Entanglement

Two qubits are entangled when their joint state cannot be written as ∣a⟩⊗∣b⟩|a\rangle\otimes|b\rangle. Measure both systems several times and compare the results: the separable state behaves like two independent coin flips, while the entangled Bell state always produces matching outcomes.

Separable∣+⟩⊗∣+⟩=12∣00⟩+12∣01⟩+12∣10⟩+12∣11⟩|{+}\rangle\otimes|{+}\rangle=\tfrac12|00\rangle+\tfrac12|01\rangle+\tfrac12|10\rangle+\tfrac12|11\rangle
A–B–✗ Different
|00⟩
0.00
|01⟩
0.00
|10⟩
0.00
|11⟩
0.00
CorrelationIndependent • 0%
Entangled∣ϕ+⟩=12∣00⟩+12∣11⟩|\phi^{+}\rangle=\tfrac{1}{\sqrt2}|00\rangle+\tfrac{1}{\sqrt2}|11\rangle
A–B–✗ Different
|00⟩
0.00
|01⟩
0.00
|10⟩
0.00
|11⟩
0.00
CorrelationPerfectly correlated • 0%
0 measurements

These correlations are stronger than anything independent classical systems can share, so entanglement is treated as a resource. One maximally entangled pair ∣ϕ+⟩|\phi^{+}\rangle is one unit of it — an e-bit.

Quantum teleportation

Quantum teleportation is a protocol that allows a sender to send quantum information to a receiver using only entanglement and classical communication to accomplish that transmission.

Setup

  • Alice holds a qubit QQ in an unknown state ∣ψ⟩|\psi\rangle that she wants to transfer to Bob.
  • Alice and Bob share an entangled pair (an e-bit) in the state ∣ϕ+⟩|\phi^{+}\rangle. Alice holds qubit AA, and Bob holds qubit BB. How or when they established this shared entanglement—for example, during an earlier meeting—is irrelevant to the protocol.
  • Alice can communicate with Bob only by sending classical bits.
  • An unknown quantum state cannot be completely described by classical bits, so classical communication alone is insufficient.
  • Because of the no-cloning theorem, once the protocol is complete and Bob's qubit is in state ∣ψ⟩|\psi\rangle, Alice no longer has a copy of that quantum state.
∣ψ⟩|\psi\rangleQABAliceBob
∣ψ⟩|\psi\rangle

Protocol

  1. 1Alice performs a controlled-NOT operation, where QQ is the control and AA is the target.
  2. 2Alice performs a Hadamard operation on QQ.
  3. 3Alice measures AA and QQ, obtaining binary outcomes mAm_A and mQm_Q, respectively.
  4. 4Alice sends mAm_A and mQm_Q to Bob.
  5. 5Bob performs these two steps on qubit BB:
    • If mA=1m_A = 1, Bob applies an XX operation.
    • If mQ=1m_Q = 1, Bob applies a ZZ operation.

Superdense coding

Superdense coding is a protocol that allows a sender to transmit two classical bits to a receiver by sending only a single qubit, using one shared e-bit of entanglement to accomplish that transmission.

Scenario

  • Alice has two classical bits (a,b)(a,b) that she wishes to transmit to Bob.
  • Alice is able to send only a single qubitsingle\ qubit to Bob.
  • Alice and Bob already share an entangled pair (an e-bit) in the state ∣ϕ+⟩|\phi^{+}\rangle.
  • Without the e-bit the task would be impossible: by Holevo's theorem, two classical bits cannot be reliably transmitted by a single qubit alone.
baAliceBob
ba

Protocol

  1. 1Alice applies XaZbX^{a}Z^{b} to her qubit — an XX when a=1a=1 and a ZZ when b=1b=1.
  2. 2Alice sends her qubit to Bob.
  3. 3Bob applies a controlled-NOT, with the qubit received from Alice as the control and his own qubit as the target.
  4. 4Bob applies a Hadamard to the qubit received from Alice.
  5. 5Bob measures both qubits, reading off aa and bb.

CHSH game

A nonlocal game is a mathematical and physical framework modeling two or more cooperating players who try to win a game against a referee. The defining rule is that once the game begins, players cannot communicate.

Set-up

  • The players Alice and Bob cooperate as a team against a referee.
  • The referee runs the game: it sends each player a question and checks their answers against a fixed rule.
  • Alice and Bob may prepare a strategy together beforehand — including sharing entanglement.
  • But once the game starts they are forbidden from communicating: neither learns the other's question or answer.
Alice
Bob
Referee
xxaabbyy

One round

The referee asks. The referee picks two questions using randomness and sends one to each player — xx to Alice and yy to Bob. Each player sees only their own question.

The CHSH referee

The CHSH game is one example of a nonlocal game, in which the referee follows these rules. Questions and answers are all bits x,y,a,b∈{0,1}x,y,a,b \in \{0,1\}, the questions xx and yy are chosen uniformly at random, and the team wins exactly when a⊕b=x∧ya \oplus b = x \wedge y.

(x,y)(x,y)x∧yx \wedge yWinning condition
(0,0)0a=ba = b
(0,1)0a=ba = b
(1,0)0a=ba = b
(1,1)1a≠ba \neq b

CHSH — Deterministic strategy

Program Alice and Bob before the game begins. For each possible question, choose the answer they will always give. Once the game starts, they cannot communicate or change their strategy. Can you find a strategy that wins all four rounds?

Strategy

Alice

If asked x=0x=0, answer
If asked x=1x=1, answer

Bob

If asked y=0y=0, answer
If asked y=1y=1, answer

CHSH game results

(0,0)x∧yx \wedge y= 0needa=ba = ba=0, b=0→a=ba = b
✓ Win
(0,1)x∧yx \wedge y= 0needa=ba = ba=0, b=0→a=ba = b
✓ Win
(1,0)x∧yx \wedge y= 0needa=ba = ba=0, b=0→a=ba = b
✓ Win
(1,1)x∧yx \wedge y= 1needa≠ba \neq ba=0, b=0→a=ba = b
✗ Lose

3 / 4 Wins

75%

🏆 This is the best possible deterministic strategy.

CHSH — Probabilistic strategy

Now let Alice and Bob answer at random. For each question, set how often they reply with a 1. The game is unchanged — only the strategy is now a coin flip. Can randomness push them past 75%?

Strategy

Alice

P(a=1∣x=0)P(a{=}1 \mid x{=}0)50%
P(a=1∣x=1)P(a{=}1 \mid x{=}1)50%

Bob

P(b=1∣y=0)P(b{=}1 \mid y{=}0)50%
P(b=1∣y=1)P(b{=}1 \mid y{=}1)50%

CHSH game results

0 wins2,000 losses

Only a few thousand rounds, so the result is noisy — a lucky sample can drift above or below the true rate. Re-run it a few times to see it bounce around the expected value.

0.0%

win rate over 2,000 rounds

expected 50.0% · best possible 75%

CHSH — Quantum strategy

Now Alice and Bob share an entangled pair prepared before the game. Each question picks a measurement angle instead of a fixed answer. Same game, same rule — but can entanglement beat the classical 75% ceiling?

Measurements are angles

Measuring a qubit is not simply reading a fixed bit — the player first chooses a direction to measure along. That choice is an angle θ\theta: outcome 0 corresponds to the direction ∣ψθ⟩|\psi_\theta\rangle, and outcome 1 to the perpendicular direction ∣ψθ+π/2⟩|\psi_{\theta+\pi/2}\rangle. The measurement asks: which of these two is the qubit closer to?

|0⟩-|0⟩|1⟩-|1⟩|ψθ+π/2⟩|ψθ⟩θcos θsin θ

Measurement basis

∣ψθ⟩=cos⁡(θ) ∣0⟩+sin⁡(θ) ∣1⟩|\psi_\theta\rangle = \cos(\theta)\,|0\rangle + \sin(\theta)\,|1\rangle
∠ θ=67.5∘\angle\,\theta = 67.5^\circ
cos⁡θ=0.383\cos\theta = 0.383sin⁡θ=0.924\sin\theta = 0.924
Common angles and their exact trigonometric values
θdegcos θsin θ
000°1100
π8\tfrac{\pi}{8}22.5°2+22\tfrac{\sqrt{2+\sqrt{2}}}{2}2−22\tfrac{\sqrt{2-\sqrt{2}}}{2}
π4\tfrac{\pi}{4}45°12\tfrac{1}{\sqrt{2}}12\tfrac{1}{\sqrt{2}}
3π8\tfrac{3\pi}{8}67.5°2−22\tfrac{\sqrt{2-\sqrt{2}}}{2}2+22\tfrac{\sqrt{2+\sqrt{2}}}{2}
π2\tfrac{\pi}{2}90°0011

So how does the angle determine the probability?

The qubit is also represented by a direction on the same circle. Suppose it points at angle φ\varphi, so its state is ∣ψφ⟩|\psi_\varphi\rangle.

A measurement compares the qubit's direction with the measurement direction θ\theta. The closer they are, the more likely the measurement returns outcome 0. If they point in exactly the same direction, the result is always 0. If they are perpendicular, outcome 0 is impossible.

Quantum mechanics quantifies this “closeness” using the inner product (also called the overlap). For a qubit pointing at angle φ\varphi and a measurement at angle θ\theta, the overlap between the qubit state ∣ψφ⟩|\psi_\varphi\rangle and the measurement direction ∣ψθ⟩|\psi_\theta\rangle is

⟨ψθ∣ψφ⟩=cos⁡(θ−φ).\langle \psi_\theta | \psi_\varphi \rangle = \cos(\theta - \varphi).

The probability of obtaining a measurement outcome is the square of the overlap with the corresponding measurement direction. Since outcome 0 corresponds to ∣ψθ⟩|\psi_\theta\rangle,

Pr⁡[outcome 0]=∣⟨ψθ∣ψφ⟩∣2=cos⁡2(θ−φ).\begin{aligned}\Pr[\text{outcome }0] &= \big|\langle \psi_\theta | \psi_\varphi \rangle\big|^2 \\ &= \cos^2(\theta - \varphi).\end{aligned}

Likewise, outcome 1 corresponds to the perpendicular direction ∣ψθ+π/2⟩|\psi_{\theta+\pi/2}\rangle, so

Pr⁡[outcome 1]=sin⁡2(θ−φ).\Pr[\text{outcome }1] = \sin^2(\theta - \varphi).

The simplest choice is θ=0∘\theta = 0^\circ, where the measurement direction lines up with the horizontal axis. In this case,

∣ψ0⟩=∣0⟩,∣ψπ/2⟩=∣1⟩.|\psi_0\rangle = |0\rangle, \qquad |\psi_{\pi/2}\rangle = |1\rangle.

So the two measurement outcomes are simply the familiar states ∣0⟩|0\rangle and ∣1⟩|1\rangle. This is called the standard (or computational, ZZ) basis. Measuring in this basis is the familiar question: “Is the qubit 0 or 1?”

Of course, nothing requires us to measure in this basis. We can rotate the measurement direction to any angle, creating a different pair of measurement states. A few of these angles are used so often that they have their own names. Each basis is simply a different measurement angle — choose one below, or drag the arrows to rotate the measurement basis and watch the outcome probabilities change while the qubit itself stays fixed.

|0⟩|1⟩01|ψ⟩
θ=0∘\theta = 0^\circstandard basis
outcome 0 · cos²(θ−φ)33%
outcome 1 · sin²(θ−φ)67%

From one qubit to an entangled pair

So far we've measured a single qubit. Now imagine Alice and Bob each receive one qubit from a shared entangled pair prepared before the game.

Just as before, each player independently chooses a measurement angle — Alice uses α\alpha, Bob uses β\beta. Each individual measurement still looks completely random: Alice sees 0 or 1 with equal probability, and so does Bob.

The surprise is that their outcomes are correlated. The probability that they obtain the same result depends only on the angle between their measurements:

Pr⁡[Alice=Bob]=cos⁡2(α−β).\Pr[\text{Alice} = \text{Bob}] = \cos^2(\alpha - \beta).

Just like for a single qubit, only the difference between the two angles matters — not their absolute positions.

  • If α=β\alpha = \beta, Alice and Bob always obtain the same result.
  • If the measurement directions are 90∘90^\circ apart, they always obtain opposite results.
  • Between these extremes, the probability changes smoothly as the angle changes.
|0⟩|1⟩α−ββα
α = 0°β = 30°|α−β| = 30.0°
same outcome · cos²(α−β)75%
different · sin²(α−β)25%

The CHSH game circuit

Alice and Bob begin with the shared Bell state ∣ϕ+⟩|\phi^{+}\rangle. Their CHSH questions don't determine the answers — they determine which measurement basis each player uses. In the circuit below, question xx selects Alice's rotation and question yy selects Bob's. After applying these rotations, both qubits are measured in the standard basis.

∣ϕ+⟩|\phi^{+}\rangleyyxxBobAlice
bbaa

Strategy

Alice

angle Alice measures at for each question x

question x=0x = 0→A0A_00°
question x=1x = 1→A1A_145°

Bob

angle Bob measures at for each question y

question y=0y = 0→B0B_022.5°
question y=1y = 1→B1B_1157.5°
|0⟩|1⟩A₀A₁B₀B₁

CHSH game results

0 wins2,000 losses

0.0%

win rate over 2,000 rounds

expected 85.4% · maximum 85.36% (Tsirelson bound)

⚛️ Optimal — this reaches the Tsirelson bound, the quantum maximum.