<!-- Source: https://web3-lab.annaburd.me/how-quantum-computing-works/quantum-information-basics/ -->

# Quantum information basics

From classical bits to qubits: state vectors and Dirac notation, measurement, compound systems,
quantum circuits, and entanglement, with teleportation, superdense coding and the CHSH game.

## Classical information

Classical states

Suppose we have a physical system $X$ that stores information. In the classical model, the system has a finite set of possible classical states. Call that set $\Sigma$.

At any given moment, $X$ is in exactly one of those states. If the current state is written as $x(t)$, then $x(t)\in\Sigma$.

$$
\Sigma=\{0,1\}
$$

The system is either 0 or 1, never both.

Dice

$$
\Sigma=\{1,2,3,4,5,6\}
$$

The upward face can be only one value at a time.

Probabilistic states

When we are uncertain about which state $X$ is in, we describe our knowledge using a probabilistic state. A probability $\Pr(x=a)\geq 0$ is assigned to each $a\in\Sigma$, and all probabilities together sum to 1.

$$
\Pr(x{=}0)=\tfrac{3}{4} \qquad \Pr(x{=}1)=\tfrac{1}{4}
$$

Probability vectors

The same distribution can be written as a probability vector - a column of numbers, one per state in $\Sigma$, in a fixed order:

$$
\begin{pmatrix}\tfrac{3}{4}\\[4pt]\tfrac{1}{4}\end{pmatrix}\begin{matrix}\leftarrow\;\Pr(x{=}0)\text{, the probability the system is in state }0\\[4pt]\leftarrow\;\Pr(x{=}1)\text{, the probability the system is in state }1\end{matrix}
$$

Dirac notation

The same probability vector can be written with Dirac notation. Suppose the elements of the state set are ordered as $\Sigma=(a_1,\ldots,a_{|\Sigma|})$. For any state $a\in\Sigma$, $\lvert a\rangle$ is the column vector having a 1 in the entry corresponding to $a$ in that ordering, with 0 for all other entries.

Standard basis vectors:

$$
\lvert 0\rangle=\begin{pmatrix}1\\0\end{pmatrix}\qquad \lvert 1\rangle=\begin{pmatrix}0\\1\end{pmatrix}
$$

Any vector is then a combination of standard basis vectors. In our example, $\tfrac{3}{4}$ is the probability of $\lvert 0\rangle$ and $\tfrac{1}{4}$ is the probability of $\lvert 1\rangle$:

$$
\begin{pmatrix}\tfrac{3}{4}\\[4pt]\tfrac{1}{4}\end{pmatrix}=\tfrac{3}{4}\lvert 0\rangle+\tfrac{1}{4}\lvert 1\rangle
$$

These notation versions all describe the same distribution. The probability statements, the column vector, and the Dirac notation are just different tools for writing the same information, each useful in a different setting.

Deterministic operations

A deterministic operation has no chance involved: each input state has exactly one output state. We can write it as a function from the state set to itself:

$$
f:\Sigma\to\Sigma,\qquad a\mapsto f(a)
$$

In Dirac notation, each state $a$ is represented by a basis vector $\lvert a\rangle$. The operation becomes a matrix $M_f$ that acts on those basis vectors:

$$
M_f\lvert a\rangle=\lvert f(a)\rangle\qquad\text{for every }a\in\Sigma
$$

- feed in the basis vector $\lvert a\rangle$ for state $a$
- get out the basis vector $\lvert f(a)\rangle$ for state $f(a)$

The matrix entries are defined by:

$$
(M_f)_{b,a}=\begin{cases}1,&b=f(a)\\0,&b\neq f(a)\end{cases}
$$

This says:

- column $a$ describes what happens to input state $a$
- the only $1$ in that column appears in the row for the output state $f(a)$
- all other entries are 0

For example, if $\Sigma=\{0,1,2\}$ and $f(0)=1,\ f(1)=1,\ f(2)=0$, then the operation always sends, with no uncertainty involved:

- $0\mapsto1$
- $1\mapsto1$
- $2\mapsto0$

This function can be represented by the matrix $M_f$. Its columns are the possible input states, and its rows are the possible output states. For each input $a$, column $a$ contains a single $1$ in row $f(a)$, with every other entry equal to $0$:

$$
M_f =
$$

$$
\begin{pmatrix}0&0&1\\1&1&0\\0&0&0\end{pmatrix}
$$

input 0

←

input 1

←

input 2

←

←row 0: output slot for state 0

←row 1: output slot for state 1

←row 2: output slot for state 2

Matrix-vector updates

Instead of tracking individual state transitions, we can apply the matrix to the entire probability distribution at once. If $v$ is the current probability vector, the updated distribution is:

$$
v'=M_fv
$$

Consider a probability distribution where input 0 has probability $\tfrac{1}{2}$, input 1 has probability $\tfrac{1}{3}$, and input 2 has probability $\tfrac{1}{6}$. The vector $v$ represents this distribution. When $M_f$ is applied, the probability associated with input 2 moves to output 0, the probabilities of inputs 0 and 1 are merged at output 1, and output 2 receives no probability mass:

$$
v=\begin{pmatrix}\tfrac{1}{2}\\[4pt]\tfrac{1}{3}\\[4pt]\tfrac{1}{6}\end{pmatrix}\qquad
M_fv=
\begin{pmatrix}
0&0&1\\
1&1&0\\
0&0&0
\end{pmatrix}
\begin{pmatrix}
\tfrac{1}{2}\\[4pt]
\tfrac{1}{3}\\[4pt]
\tfrac{1}{6}
\end{pmatrix}
=
\begin{pmatrix}
\tfrac{1}{6}\\[4pt]
\tfrac{1}{2}+\tfrac{1}{3}\\[4pt]
0
\end{pmatrix}
=
\begin{pmatrix}
\tfrac{1}{6}\\[4pt]
\tfrac{5}{6}\\[4pt]
0
\end{pmatrix}
$$

We could update the probabilities by following each input separately and moving its probability mass to the corresponding output. The matrix $M_f$ packages all of these transfers into a single matrix-vector multiplication, producing the updated probability distribution in one step.

One-bit matrices

For a single bit $\Sigma=\{0,1\}$, there are four possible deterministic operations:

Set to 0

$$
0,1\mapsto0
$$

$$
M_0=\begin{pmatrix}1&1\\0&0\end{pmatrix}
$$

Identity

$$
0\mapsto0,\quad1\mapsto1
$$

$$
I=\begin{pmatrix}1&0\\0&1\end{pmatrix}
$$

NOT / bit flip

$$
0\mapsto1,\quad1\mapsto0
$$

$$
X=\begin{pmatrix}0&1\\1&0\end{pmatrix}
$$

Set to 1

$$
0,1\mapsto1
$$

$$
M_1=\begin{pmatrix}0&0\\1&1\end{pmatrix}
$$

Bras and inner products

In Dirac notation, column vectors are called kets and are written as $\lvert 0\rangle$ and $\lvert 1\rangle$. Their row-vector counterparts are called bras and are written with the bracket facing the opposite direction:

$$
\langle\text{bra}\mid\text{ket}\rangle\qquad \langle 0\rvert=\begin{pmatrix}1&0\end{pmatrix}\qquad \langle 1\rvert=\begin{pmatrix}0&1\end{pmatrix}
$$

Multiplying a row vector by a column vector produces a single number. This operation is called the inner product, or dot product. It measures how much two vectors overlap:

$$
\begin{pmatrix}r_1&r_2&\cdots&r_n\end{pmatrix}\begin{pmatrix}c_1\\c_2\\\vdots\\c_n\end{pmatrix}=r_1c_1+r_2c_2+\cdots+r_nc_n=\sum_{i=1}^{n}r_ic_i
$$

For basis states, kets and bras contain a single 1 and zeros everywhere else. If we multiply matching states, the 1s line up. If we multiply different states, the 1s occur in different positions, so every term in the sum is zero:

$$
\langle 0\vert0\rangle=\begin{pmatrix}1&0\end{pmatrix}\begin{pmatrix}1\\0\end{pmatrix}=1\cdot1+0\cdot0=1\qquad
\langle 0\vert1\rangle=\begin{pmatrix}1&0\end{pmatrix}\begin{pmatrix}0\\1\end{pmatrix}=1\cdot0+0\cdot1=0
$$

So multiplying a bra by a ket, denoted as $\langle a\vert b\rangle$, acts like an equality test: it returns 1 when the states are the same and 0 when they are different:

$$
\langle a\vert b\rangle=\begin{cases}1,&a=b\\0,&a\neq b\end{cases}
$$

Ket-bra products

If row-by-column multiplication gives a single number, then column-by-row multiplication gives a matrix. This operation is called an outer product:

$$
\begin{pmatrix}c_1\\c_2\\\vdots\\c_m\end{pmatrix}\begin{pmatrix}r_1&r_2&\cdots&r_n\end{pmatrix}
=
\begin{pmatrix}
c_1r_1&c_1r_2&\cdots&c_1r_n\\
c_2r_1&c_2r_2&\cdots&c_2r_n\\
\vdots&\vdots&\ddots&\vdots\\
c_mr_1&c_mr_2&\cdots&c_mr_n
\end{pmatrix}
$$

For one-bit basis states, a ket-bra product creates a matrix with a single 1 and zeros everywhere else:

$$
\lvert0\rangle\langle0\rvert=
\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}
=
\begin{pmatrix}1&0\\0&0\end{pmatrix}
\qquad
\lvert0\rangle\langle1\rvert=
\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix}
=
\begin{pmatrix}0&1\\0&0\end{pmatrix}
$$

Suppose we are given a deterministic operation $f:\Sigma\to\Sigma$ where, for each state $b\in\Sigma$, the function specifies an output state $f(b)$, representing this individual rule $b\mapsto f(b)$ with the ket-bra operator $\lvert f(b)\rangle\langle b\rvert$, where the bra $\langle b\rvert$ identifies the input state $\lvert b\rangle$ while the ket $\lvert f(b)\rangle$ specifies the state that should be produced, so that adding one such term for every possible input state yields a matrix that implements the entire deterministic operation:

$$
M_f=\sum_{b\in\Sigma}\lvert f(b)\rangle\langle b\rvert
$$

For the function $f(0)=1,\ f(1)=1,\ f(2)=0$, the deterministic rules are $0\mapsto1,\ 1\mapsto1,\ 2\mapsto0$, and each rule contributes one ket-bra term: the rule $0\mapsto1$ becomes $\lvert1\rangle\langle0\rvert$, the rule $1\mapsto1$ becomes $\lvert1\rangle\langle1\rvert$, and the rule $2\mapsto0$ becomes $\lvert0\rangle\langle2\rvert$ — note that while we are used to reading input on the left and output on the right, in ket-bra notation the output comes first: $\lvert1\rangle\langle0\rvert$ means "produce $\lvert1\rangle$ when you see $\lvert0\rangle$."

Adding these terms gives:

$$
M_f=\lvert1\rangle\langle0\rvert+\lvert1\rangle\langle1\rvert+\lvert0\rangle\langle2\rvert
$$

Applying it to the basis state $\lvert2\rangle$, the first two inner products vanish and only the matching term survives:

$$
M_f\lvert2\rangle=\lvert1\rangle\underbrace{\langle0\vert2\rangle}_{0}+\lvert1\rangle\underbrace{\langle1\vert2\rangle}_{0}+\lvert0\rangle\underbrace{\langle2\vert2\rangle}_{1}=\lvert0\rangle
$$

Each of the three ket-bra terms in $M_f$ is an ordinary matrix:

$$
\lvert1\rangle\langle0\rvert=
\begin{pmatrix}0&0&0\\1&0&0\\0&0&0\end{pmatrix}
\qquad
\lvert1\rangle\langle1\rvert=
\begin{pmatrix}0&0&0\\0&1&0\\0&0&0\end{pmatrix}
\qquad
\lvert0\rangle\langle2\rvert=
\begin{pmatrix}0&0&1\\0&0&0\\0&0&0\end{pmatrix}
$$

Adding them together gives exactly the deterministic-operation matrix constructed earlier:

$$
M_f=
\begin{pmatrix}0&0&0\\1&0&0\\0&0&0\end{pmatrix}
+
\begin{pmatrix}0&0&0\\0&1&0\\0&0&0\end{pmatrix}
+
\begin{pmatrix}0&0&1\\0&0&0\\0&0&0\end{pmatrix}
=
\begin{pmatrix}0&0&1\\1&1&0\\0&0&0\end{pmatrix}
$$

The action of this matrix becomes clear when it is applied to a basis state $\lvert a\rangle$. In each term, the inner product $\langle b\vert a\rangle$ acts as an equality test: it equals 1 when $b=a$ and 0 otherwise. As a result, every term in the sum vanishes except the one corresponding to $b=a$, leaving:

$$
M_f\lvert a\rangle=
\left(\sum_{b\in\Sigma}\lvert f(b)\rangle\langle b\rvert\right)\lvert a\rangle
=\sum_{b\in\Sigma}\lvert f(b)\rangle\langle b\vert a\rangle
=\lvert f(a)\rangle
$$

Thus the matrix sends each basis state $\lvert a\rangle$ to the state specified by the function, $\lvert f(a)\rangle$.

Probabilistic operations

A deterministic operation maps each input state to exactly one output state. A probabilistic operation is more general: each input state produces a probability distribution over output states. The operation is described by a matrix $M$ where entry $M_{b,a}$ gives the probability that input state $a$ produces output state $b$.

For $M$ to represent a valid probabilistic operation, two conditions must hold. All entries must be nonnegative real numbers:

$$
M_{b,a}\geq 0\qquad\text{for all }a,b\in\Sigma
$$

And the entries in each column must sum to 1 — each column is a probability vector describing the output distribution for one input state:

$$
\sum_{b\in\Sigma}M_{b,a}=1\qquad\text{for all }a\in\Sigma
$$

A matrix satisfying both conditions is called a stochastic matrix. Every deterministic-operation matrix is a special case: its columns each contain a single 1 and zeros elsewhere, which is a valid probability vector.

Composing operations

When two operations are applied in sequence — first $M_1$, then $M_2$ — the result is $M_2(M_1 v)$. Because matrix multiplication is associative, this equals $(M_2 M_1)v$: the composition is itself a matrix, and the product of two stochastic matrices is stochastic.

$$
M_2(M_1 v)=(M_2 M_1)v
$$

Matrix multiplication is not commutative, however: $M_2 M_1$ and $M_1 M_2$ generally produce different results. The order in which operations are applied matters.

Using the one-bit matrices introduced earlier, we can see this directly. Applying $M_0$ (set to 0) and then $X$ (NOT) always produces 1 — equivalent to $M_1$. Doing the same two operations in the opposite order always produces 0 — equivalent to $M_0$ itself:

$$
X M_0=\begin{pmatrix}0&1\\1&0\end{pmatrix}\begin{pmatrix}1&1\\0&0\end{pmatrix}=\begin{pmatrix}0&0\\1&1\end{pmatrix}=M_1
$$

$$
M_0 X=\begin{pmatrix}1&1\\0&0\end{pmatrix}\begin{pmatrix}0&1\\1&0\end{pmatrix}=\begin{pmatrix}1&1\\0&0\end{pmatrix}=M_0
$$

## Quantum information

Two levels of description

Quantum information can be described at two levels of generality. The simplified picture — kets and unitary matrices — is enough for pure states and reversible operations. The general picture adds density matrices and a broader class of measurements and operations, covering mixed states, noise, and measurement.

|  | Simplified | General |
| --- | --- | --- |
| Quantum states | Kets — state vectors | Density matrices |
| Operations | Unitary matrices | More general class of measurements and operations |

Quantum states

A quantum state of a system is represented by a column vector whose indices are placed in correspondence with the classical states of that system:

$$
\lvert\psi\rangle=\begin{pmatrix}\alpha_1\\\vdots\\\alpha_n\end{pmatrix}
$$

A valid quantum state vector must satisfy two conditions:

- Entries are complex numbers — amplitudes, not probabilities: $\alpha_k\in\mathbb{C}$
- The sum of squared absolute values of the entries equals 1: $\sum_{k=1}^{n}|\alpha_k|^2=1$

The second condition uses the Euclidean norm — the generalisation of vector length to complex numbers:

$$
\|\psi\|=\sqrt{\sum_{k=1}^{n}|\alpha_k|^2}
$$

Real vector

$$
v=\begin{pmatrix}3\\4\end{pmatrix}
$$

$$
\|v\|=\sqrt{3^2+4^2}=\sqrt{25}=5
$$

$$
\begin{aligned}\|v\|&=\sqrt{3^2+4^2}\\&=\sqrt{25}=5\end{aligned}
$$

Complex vector

$$
v=\begin{pmatrix}\tfrac{3}{\sqrt{2}}\\[4pt]\tfrac{i}{\sqrt{2}}\end{pmatrix}
$$

$$
\|v\|=\sqrt{\left|\tfrac{3}{\sqrt{2}}\right|^2+\left|\tfrac{i}{\sqrt{2}}\right|^2}=\sqrt{\tfrac{9}{2}+\tfrac{1}{2}}=\sqrt{5}
$$

Quantum state vectors are unit vectors — $\|\psi\|=1$. The amplitudes encode probability amplitudes: the probability of observing the system in classical state $k$ is $|\alpha_k|^2$.

Qubit states

A qubit is a quantum system whose classical state set is $\Sigma=\{0,1\}$, so its state vector has two entries. The basis states $\lvert0\rangle$ and $\lvert1\rangle$ look like ordinary integer vectors, but their entries are complex numbers that happen to have zero imaginary part:

$$
\lvert0\rangle=\begin{pmatrix}1\\0\end{pmatrix}\qquad\lvert1\rangle=\begin{pmatrix}0\\1\end{pmatrix}
$$

A general qubit state is any normalised complex combination of these two basis states:

$$
\lvert\psi\rangle=\alpha\lvert0\rangle+\beta\lvert1\rangle,\qquad\alpha,\beta\in\mathbb{C},\qquad|\alpha|^2+|\beta|^2=1
$$

Two particularly important named states are the plus and minus states. They assign equal probability to the basis states $\lvert0\rangle$ and $\lvert1\rangle$ — 50% and 50% — but differ in the relative sign between their amplitudes:

$$
\lvert+\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{\sqrt{2}}\lvert1\rangle\qquad\lvert-\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle-\tfrac{1}{\sqrt{2}}\lvert1\rangle
$$

Most qubit states have no special name. Any normalised choice of $\alpha$ and $\beta$ is valid — for example:

$$
\lvert\phi\rangle=\tfrac{1+2i}{3}\lvert0\rangle-\tfrac{2}{3}\lvert1\rangle\qquad\left|\tfrac{1+2i}{3}\right|^2+\left|\tfrac{2}{3}\right|^2=\tfrac{5}{9}+\tfrac{4}{9}=1\checkmark
$$

Conjugate transpose

Every ket $\lvert\psi\rangle$ has a corresponding bra $\langle\psi\rvert$ obtained by the conjugate transpose — written with a $†$ (dagger):

$$
\langle\psi\rvert=\lvert\psi\rangle^\dagger
$$

Two steps: transpose the column vector into a row, then replace each entry with its complex conjugate — $a+bi\;\mapsto\;a-bi$. For the qubit state $\lvert\phi\rangle=\tfrac{1+2i}{3}\lvert0\rangle-\tfrac{2}{3}\lvert1\rangle$:

$$
\langle\phi\rvert=\lvert\phi\rangle^\dagger=\frac{1-2i}{3}\langle0\rvert-\frac{2}{3}\langle1\rvert
$$

The real amplitude $-\tfrac{2}{3}$ is unchanged — conjugating a real number leaves it the same. Only the complex entry $\tfrac{1+2i}{3}$ flips its imaginary part to give $\tfrac{1-2i}{3}$.

Why do we conjugate?

A complex number is not just a number—it can also be viewed as a point (or vector) in the complex plane.

$$
z = x + iy
$$

Its length is simply the Euclidean distance from the origin:

$$
\lvert z\rvert=\sqrt{x^2+y^2}
$$

The challenge is to compute this length using only complex arithmetic. Notice what happens if we multiply $z$ by its conjugate:

$$
\overline{z}\,z=(x-iy)(x+iy)=x^2\,\textcolor{#dc2626}{\cancel{+\,ixy}}\,\textcolor{#dc2626}{\cancel{-\,ixy}}\,+y^2=x^2+y^2
$$

The imaginary terms cancel, leaving exactly the square of the Euclidean length:

$$
\overline{z}\,z=\left(\sqrt{x^2+y^2}\right)^2=\lvert z\rvert^2
$$

z = 1.20 + 0.60i

z = 1.20 − 0.60i

zz= (1.20)² + (0.60)²= 1.44 + 0.36= 1.80 (real)

|z| = √1.80 = 1.34

Drag z around the plane. Its conjugate z is the mirror image across the real axis, and the product zz always lands on the real axis at |z|².

This is why the complex conjugate appears. It isn't an arbitrary rule—it is the operation that recovers the ordinary Euclidean length of a complex number.

The inner product extends this same idea to vectors by applying the conjugate transpose to every amplitude:

$$
\langle\psi\vert\psi\rangle=\sum_{k}\overline{\alpha_k}\,\alpha_k=\sum_{k}\lvert\alpha_k\rvert^2
$$

Measurements

Measuring a quantum state extracts classical information from it. In a standard basis measurement, the possible outcomes are the classical states — the same states that label the entries of the state vector. For a state $\lvert\psi\rangle=\sum_{a\in\Sigma}\alpha_a\lvert a\rangle$, the probability of obtaining outcome $a$ is the squared absolute value of the corresponding amplitude:

$$
\Pr(\text{outcome}=a)=|\alpha_a|^2
$$

For example, measuring the state $\lvert{+}\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{\sqrt{2}}\lvert1\rangle$ gives each outcome with equal probability:

$$
\Pr(\text{outcome}=0)=\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}\qquad\Pr(\text{outcome}=1)=\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}
$$

For the state with complex amplitudes:

$$
\lvert\psi\rangle=\frac{1+2i}{3}\lvert0\rangle-\frac{2}{3}\lvert1\rangle
$$

$$
\Pr(\text{outcome}=0)=\left|\frac{1+2i}{3}\right|^2=\frac{1^2+2^2}{9}=\frac{5}{9}\qquad\Pr(\text{outcome}=1)=\left|\frac{2}{3}\right|^2=\frac{4}{9}
$$

Measurement also changes the state. Once outcome $a$ is observed, the quantum state collapses to the corresponding basis state $\lvert a\rangle$. A second measurement on the collapsed state will always return the same outcome — this is called the collapse of the quantum state.

The same logic applies to ordinary probability. Flip a coin and let it land face-up on the table. Before you look, each side has probability $\tfrac{1}{2}$. The moment you look, the uncertainty is gone — the coin is showing heads, and re-checking it a second or third time still shows heads with certainty. What changes is not the coin but your state of knowledge about it. Quantum collapse works the same way from the outside: once you have a measurement result, subsequent measurements on that same state are no longer uncertain.

After the measurement, regardless of the pre-measurement state, the system is in a definite classical state $\lvert0\rangle$ or $\lvert1\rangle$. This places a fundamental limit on how much classical information can be extracted from a quantum state in a single measurement.

Unitary operations

Quantum operations are represented by unitary matrices — a different constraint from the stochastic matrices of classical probabilistic operations. A square matrix $U$ is unitary if its conjugate transpose is also its inverse:

$$
U^\dagger U = I = UU^\dagger
$$

This has two equivalent restatements:

- The inverse is simply the conjugate transpose — $U^{-1}=U^\dagger$— so inverting a unitary is cheap.
- Unitary matrices preserve the Euclidean norm: $\|U\lvert\psi\rangle\|=\|\lvert\psi\rangle\|$.

To see why the norm is preserved, first notice that for any column vector, multiplying it by its own conjugate transpose gives the squared norm — the conjugates pair with each entry to produce $\bar\alpha_k\alpha_k=|\alpha_k|^2$:

$$
v^\dagger v=\begin{pmatrix}\bar\alpha_1&\cdots&\bar\alpha_n\end{pmatrix}\begin{pmatrix}\alpha_1\\\vdots\\\alpha_n\end{pmatrix}=|\alpha_1|^2+\cdots+|\alpha_n|^2=\|v\|^2
$$

Or, in Dirac notation:

$$
\langle\psi\vert\psi\rangle=\bar\alpha_1\alpha_1+\cdots+\bar\alpha_n\alpha_n=|\alpha_1|^2+\cdots+|\alpha_n|^2=\|\lvert\psi\rangle\|^2
$$

Since $v^\dagger v=\|v\|^2$ holds for any vector, we can apply it to $v=U\lvert\psi\rangle$. The rule $(AB)^\dagger=B^\dagger A^\dagger$ says the conjugate transpose of a product reverses the order — so:

$$
(U\lvert\psi\rangle)^\dagger=\lvert\psi\rangle^\dagger U^\dagger=\langle\psi\rvert U^\dagger
$$

Substituting into $v^\dagger v$ and applying $U^\dagger U=I$:

$$
\|U\lvert\psi\rangle\|^2=\langle\psi\rvert U^\dagger U\lvert\psi\rangle=\langle\psi\rvert I\lvert\psi\rangle=\langle\psi\vert\psi\rangle=\|\lvert\psi\rangle\|^2
$$

Taking square roots gives $\|U\lvert\psi\rangle\|=\|\lvert\psi\rangle\|$. As a numeric check, applying $\sigma_x$ to the familiar state:

$$
\lvert\phi\rangle=\tfrac{1+2i}{3}\lvert0\rangle-\tfrac{2}{3}\lvert1\rangle
$$

It simply swaps the two entries:

$$
\sigma_x\lvert\phi\rangle=\begin{pmatrix}0&1\\1&0\end{pmatrix}\begin{pmatrix}\tfrac{1+2i}{3}\\[4pt]-\tfrac{2}{3}\end{pmatrix}=\begin{pmatrix}-\tfrac{2}{3}\\[4pt]\tfrac{1+2i}{3}\end{pmatrix}
$$

$$
\|\lvert\phi\rangle\|=\sqrt{\tfrac{5}{9}+\tfrac{4}{9}}=1\qquad\|\sigma_x\lvert\phi\rangle\|=\sqrt{\tfrac{4}{9}+\tfrac{5}{9}}=1
$$

Geometrically, a unitary transformation is the complex-number analogue of a rotation: it changes the direction of the vector, but never its length. Since quantum state vectors are unit vectors, a unitary operation always maps valid quantum states to valid quantum states — it can never take a state outside the unit sphere.

To check if a matrix is unitary, multiply it by its conjugate transpose and see if the result is the identity.

For $\sigma_x$, which is real and symmetric so $\sigma_x^\dagger=\sigma_x$:

$$
\sigma_x^\dagger\sigma_x=\begin{pmatrix}0&1\\1&0\end{pmatrix}\begin{pmatrix}0&1\\1&0\end{pmatrix}=\begin{pmatrix}1&0\\0&1\end{pmatrix}=I\checkmark
$$

$$
\begin{aligned}
\sigma_x^\dagger\sigma_x
&=\begin{pmatrix}0&1\\1&0\end{pmatrix}\begin{pmatrix}0&1\\1&0\end{pmatrix}\\
&=\begin{pmatrix}1&0\\0&1\end{pmatrix}=I\checkmark
\end{aligned}
$$

Qubit unitary operations

1. Pauli operations

The four Pauli matrices are the most common single-qubit unitary operations:

Identity

$$
I=\begin{pmatrix}1&0\\0&1\end{pmatrix}
$$

Bit flip

$$
\sigma_x=\begin{pmatrix}0&1\\1&0\end{pmatrix}
$$

Phase + bit flip

$$
\sigma_y=\begin{pmatrix}0&-i\\i&0\end{pmatrix}
$$

Phase flip

$$
\sigma_z=\begin{pmatrix}1&0\\0&-1\end{pmatrix}
$$

The Pauli matrices also happen to be Hermitian — a matrix is Hermitian if it equals its own conjugate transpose, $M^\dagger=M$. For a real symmetric matrix this just means symmetry across the diagonal. For a complex matrix, entries are mirrored across the diagonal and conjugated.

2. Hadamard

The Hadamard gate H, named after the French mathematician [Jacques Hadamard](https://en.wikipedia.org/wiki/Jacques_Hadamard), is one of the most important quantum gates. It acts as a bridge between two ways of describing a qubit:

- The computational basis: $\lvert0\rangle$ and $\lvert1\rangle$
- The superposition basis: $\lvert+\rangle$ and $\lvert-\rangle$

$$
H=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix}
$$

From 0, 1 to +, -:

$$
H\lvert0\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}1\\0\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}\end{pmatrix}=\lvert+\rangle
$$

$$
H\lvert1\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}0\\1\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[6pt]-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\lvert-\rangle
$$

From +, - back to 0, 1:

$$
H\lvert+\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}\end{pmatrix}=\begin{pmatrix}1\\0\end{pmatrix}=\lvert0\rangle
$$

$$
H\lvert-\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[6pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[6pt]-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\begin{pmatrix}0\\1\end{pmatrix}=\lvert1\rangle
$$

Applying a Hadamard gate transforms a computational basis state into an equal superposition of $\lvert0\rangle$ and $\lvert1\rangle$, meaning that a measurement would find each outcome with equal probability:

$$
H\lvert0\rangle=\lvert+\rangle=\frac{\lvert0\rangle+\lvert1\rangle}{\sqrt{2}},\qquad H\lvert1\rangle=\lvert-\rangle=\frac{\lvert0\rangle-\lvert1\rangle}{\sqrt{2}}
$$

You can think of the Hadamard gate as a quantum "basis changer" that lets us move between definite states and equal-probability superpositions, making it a fundamental building block of many quantum algorithms.

3. Phase gates

Phase gates leave $\lvert0\rangle$ unchanged and rotate $\lvert1\rangle$ by a complex phase. They do not change measurement probabilities on their own, but they shift the relative phase between amplitudes, which affects how states interfere in a larger circuit.

The general phase rotation gate $R_\phi$ applies a phase $e^{i\phi}$ to $\lvert1\rangle$ while leaving $\lvert0\rangle$ alone. $\phi$ is a real number, making $i\phi$ always purely imaginary. This means $e^{i\phi}$ always lies on the complex unit circle — $|e^{i\phi}|=1$ — so it is a pure phase that rotates without scaling:

$$
R_\phi=\begin{pmatrix}1&0\\0&e^{i\phi}\end{pmatrix}
$$

Two standard choices give the S gate and the T gate:

$$
S=R_{\pi/2}=\begin{pmatrix}1&0\\0&i\end{pmatrix}\qquad T=R_{\pi/4}=\begin{pmatrix}1&0\\0&e^{i\pi/4}\end{pmatrix}
$$

S applies a quarter-turn phase (90°) and satisfies $S^2=Z$, where Z is the Pauli phase-flip gate from section 1: $Z=\begin{pmatrix}1&0\\0&-1\end{pmatrix}$. T applies an eighth-turn phase (45°) and satisfies $T^2=S$ and $T^4=Z$. Together with H, T generates a gate set that can approximate any single-qubit unitary to arbitrary precision.

Applied to the basis states:

$$
S\lvert0\rangle=\lvert0\rangle,\quad S\lvert1\rangle=i\lvert1\rangle
$$

$$
T\lvert0\rangle=\lvert0\rangle,\quad T\lvert1\rangle=e^{i\pi/4}\lvert1\rangle
$$

Composing unitary operations

Gates compose by matrix multiplication. In the product the rightmost matrix acts first — the state travels right to left through the sequence:

$$
\lvert\psi_\text{out}\rangle=\overset{(3)}{U_3}\;\overset{(2)}{U_2}\;\overset{(1)}{U_1}\lvert\psi_\text{in}\rangle
$$

For example, in $\textcolor{#2563eb}{H}\textcolor{#d97706}{S}\textcolor{#7c3aed}{H}$ the rightmost gate (purple H) acts first, then S, then the leftmost H (blue). Expanding step by step:

$$
\textcolor{#2563eb}{H}\textcolor{#d97706}{S}\textcolor{#7c3aed}{H}=\textcolor{#2563eb}{\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}}\textcolor{#d97706}{\begin{pmatrix}1&0\\0&i\end{pmatrix}}\textcolor{#7c3aed}{\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}}=\textcolor{#2563eb}{\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}}\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{i}{\sqrt{2}}&-\tfrac{i}{\sqrt{2}}\end{pmatrix}=\frac{1}{2}\begin{pmatrix}1+i&1-i\\1-i&1+i\end{pmatrix}=\sqrt{X}
$$

This matrix is called the square root of NOT (written $\sqrt{X}$), because applying it twice gives the Pauli X (NOT) gate. To see why, insert $HH=I$ in the middle and use $S^2=Z$:

$$
(HSH)^2=HS\underbrace{HH}_{I}SH=HS^2H=HZH=X
$$

In ordinary arithmetic, squaring something makes it "more of the same" — you would never expect a number squared to flip a sign. But unitary matrices can have complex eigenvalues such as $i$, and $i^2=-1$ introduces the sign change that turns a partial rotation into a full logical inversion. This is one of the ways quantum gates behave fundamentally differently from classical boolean operations.

As another example, consider HTH. T is a rotation around the Z axis by $\pi/4$. Placing H on both sides redirects that same rotation onto the X axis:

$$
HTH=\frac{1}{2}\begin{pmatrix}1+e^{i\pi/4}&1-e^{i\pi/4}\\1-e^{i\pi/4}&1+e^{i\pi/4}\end{pmatrix}=e^{i\pi/8}\begin{pmatrix}\cos\tfrac{\pi}{8}&-i\sin\tfrac{\pi}{8}\\-i\sin\tfrac{\pi}{8}&\cos\tfrac{\pi}{8}\end{pmatrix}=e^{i\pi/8}R_x\!\left(\tfrac{\pi}{4}\right)
$$

So T and HTH are rotations around two non-parallel axes — Z and X — each by $\pi/4$. Combining rotations around any two non-parallel axes generates all rotations of the sphere, so any single-qubit unitary can be reached by some finite sequence of T and HTH to any desired precision.

## Multiple systems: classical

Classical states

Suppose we have two systems:

- $X$ with classical state set $\Sigma$.
- $Y$ with classical state set $\Gamma$.

Together they form a compound system, written $(X,Y)$ or simply $XY$. At any moment the compound system is in exactly one state - a pair $(a,b)$ where $a\in\Sigma$ is the state of $X$ and $b\in\Gamma$ is the state of $Y$.

The full set of possible states of $XY$ is the Cartesian product:

$$
\Sigma\times\Gamma=\{(a,b):a\in\Sigma,\,b\in\Gamma\}
$$

For example, if both $X$ and $Y$ are bits so $\Sigma=\Gamma=\{0,1\}$, the compound system has four possible states:

$$
\Sigma\times\Gamma=\{(0,0),(0,1),(1,0),(1,1)\}
$$

General formula: if we combine $n$ classical systems with state sets $\Sigma_1,\ldots,\Sigma_n$, their compound state set is:

$$
\Sigma_1\times\cdots\times\Sigma_n=\{(a_1,\ldots,a_n):a_i\in\Sigma_i\text{ for }i=1,\ldots,n\}
$$

Convention notes

- When we list a Cartesian product, we usually use lexicographic order: compare the first coordinate, then the second, and so on.
- This assumes each state set already has an order. For bits, we use $0<1$.
- Significance decreases from left to right: the leftmost coordinate is the most significant, and the rightmost coordinate changes fastest.

Example: for three bits, $\Sigma_1=\Sigma_2=\Sigma_3=\{0,1\}$. In lexicographic order:

$$
\Sigma_1\times\Sigma_2\times\Sigma_3=\{(0,0,0),(0,0,1),(0,1,0),(0,1,1),(1,0,0),(1,0,1),(1,1,0),(1,1,1)\}
$$

Probabilistic states

For a compound classical system, a probabilistic state assigns one probability to each state in the Cartesian product. If $x$ is the current state of $X$ and $y$ is the current state of $Y$, then:

$$
\Pr\bigl((x,y)=(a,b)\bigr)\geq 0\quad\text{for all }(a,b)\in\Sigma\times\Gamma,\qquad\sum_{(a,b)\in\Sigma\times\Gamma}\Pr\bigl((x,y)=(a,b)\bigr)=1
$$

For example, for the system of two bits, where $\Sigma=\Gamma=\{0,1\}$ and the possible compound states are $(0,0),(0,1),(1,0),(1,1)$, one of multiple possible probabilistic states can be:

$$
\begin{aligned}
\Pr\bigl((x,y)=(0,0)\bigr)&=\tfrac{1}{2}\\[4pt]
\Pr\bigl((x,y)=(0,1)\bigr)&=0\\[4pt]
\Pr\bigl((x,y)=(1,0)\bigr)&=0\\[4pt]
\Pr\bigl((x,y)=(1,1)\bigr)&=\tfrac{1}{2}
\end{aligned}
$$

The two-bit system is equally likely to be in state $(0,0)$ or state $(1,1)$, and has probability zero of being in the other two states. In vector form (using lexicographic order):

$$
u=\begin{pmatrix}
\tfrac{1}{2}\\[4pt]
0\\[4pt]
0\\[4pt]
\tfrac{1}{2}
\end{pmatrix}
\begin{matrix}
\leftarrow\;\text{probability associated with state }00\\[4pt]
\leftarrow\;\text{probability associated with state }01\\[4pt]
\leftarrow\;\text{probability associated with state }10\\[4pt]
\leftarrow\;\text{probability associated with state }11
\end{matrix}
$$

For another two-bit system, the probability state might look like:

$$
v=\begin{pmatrix}
\tfrac{1}{4}\\[4pt]
\tfrac{1}{4}\\[4pt]
\tfrac{1}{4}\\[4pt]
\tfrac{1}{4}
\end{pmatrix}
\begin{matrix}
\leftarrow\;\text{probability associated with state }00\\[4pt]
\leftarrow\;\text{probability associated with state }01\\[4pt]
\leftarrow\;\text{probability associated with state }10\\[4pt]
\leftarrow\;\text{probability associated with state }11
\end{matrix}
$$

For a given probabilistic state of $(X,Y)$, we say $X$ and $Y$ are independent if $\Pr\bigl((x,y)=(a,b)\bigr)=\Pr(x=a)\,\Pr(y=b)$ for all $a\in\Sigma$ and $b\in\Gamma$.

Let's check $u$ and $v$ against this rule — by looking for a contradiction.

u — suppose the rule holds:

$$
u=\begin{pmatrix}\tfrac{1}{2}\\[4pt]0\\[4pt]0\\[4pt]\tfrac{1}{2}\end{pmatrix}\begin{array}{l}\leftarrow\text{ ① }\Pr(x{=}0)>0\\[4pt]\leftarrow\text{ ③ }=0,\text{ so }\Pr(x{=}0){=}0\text{ or }\Pr(y{=}1){=}0\text{ — contradicts ①② ✗}\\[4pt]\phantom{\leftarrow}\\[4pt]\leftarrow\text{ ② }\Pr(y{=}1)>0\end{array}
$$

Contradiction — $X$ and $Y$ are not independent under $u$.

v — no contradiction arises:

$$
v=\begin{pmatrix}\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\end{pmatrix}\begin{matrix}\leftarrow\;\tfrac{1}{4}=\tfrac{1}{2}\cdot\tfrac{1}{2}=\Pr(x{=}0)\,\Pr(y{=}0)\;\checkmark\\[4pt]\leftarrow\;\tfrac{1}{4}=\tfrac{1}{2}\cdot\tfrac{1}{2}=\Pr(x{=}0)\,\Pr(y{=}1)\;\checkmark\\[4pt]\leftarrow\;\tfrac{1}{4}=\tfrac{1}{2}\cdot\tfrac{1}{2}=\Pr(x{=}1)\,\Pr(y{=}0)\;\checkmark\\[4pt]\leftarrow\;\tfrac{1}{4}=\tfrac{1}{2}\cdot\tfrac{1}{2}=\Pr(x{=}1)\,\Pr(y{=}1)\;\checkmark\end{matrix}
$$

$X$ and $Y$ are independent under $v$.

Correlation is, in a sense, a lack of independence: when two systems are not independent, the state of one carries information about the state of the other.

Dirac notation

The same rule for independence can be written in Dirac notation. Suppose that a probabilistic state of $(X,Y)$ is expressed as a vector:

$$
|\pi\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|ab\rangle
$$

The systems $X$ and $Y$ are independent if there exist probability vectors

$$
|\phi\rangle=\sum_{a\in\Sigma}q_a\,|a\rangle\quad\text{and}\quad|\psi\rangle=\sum_{b\in\Gamma}r_b\,|b\rangle
$$

such that $p_{ab}=q_a r_b$ for all $a\in\Sigma$ and $b\in\Gamma$.

- $|\pi\rangle$ — the probabilistic state of the compound system $(X,Y)$
- $p_{ab}$ — the probability assigned to outcome $(a,b)$
- $|ab\rangle$ — a basis vector labelling the outcome $X{=}a,\,Y{=}b$for $X,Y\in\{0,1\}$ these are the four standard basis vectors:

$$
|00\rangle=\begin{pmatrix}1\\0\\0\\0\end{pmatrix},\quad|01\rangle=\begin{pmatrix}0\\1\\0\\0\end{pmatrix},\quad|10\rangle=\begin{pmatrix}0\\0\\1\\0\end{pmatrix},\quad|11\rangle=\begin{pmatrix}0\\0\\0\\1\end{pmatrix}
$$

Returning to $u$ and $v$ from the earlier example. For $u$, its column vector alongside its $|\pi\rangle$ expansion:

$$
u=\begin{pmatrix}\textcolor{blue}{\tfrac{1}{2}}\\[4pt]0\\[4pt]0\\[4pt]\textcolor{orange}{\tfrac{1}{2}}\end{pmatrix}\qquad|\pi\rangle=\textcolor{blue}{\tfrac{1}{2}}\,|00\rangle+0\cdot|01\rangle+0\cdot|10\rangle+\textcolor{orange}{\tfrac{1}{2}}\,|11\rangle
$$

For $v$, its column vector alongside its $|\pi\rangle$ expansion:

$$
v=\begin{pmatrix}\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\\[4pt]\tfrac{1}{4}\end{pmatrix}\qquad|\pi\rangle=\tfrac{1}{4}\,|00\rangle+\tfrac{1}{4}\,|01\rangle+\tfrac{1}{4}\,|10\rangle+\tfrac{1}{4}\,|11\rangle
$$

Since $X$ and $Y$ are independent under $v$, there exist probability vectors $|\phi\rangle$ and $|\psi\rangle$ such that each coefficient in $|\pi\rangle$ is a product of one factor from each:

$$
|\phi\rangle=\textcolor{blue}{\tfrac{1}{2}}\,|0\rangle+\textcolor{teal}{\tfrac{1}{2}}\,|1\rangle\qquad|\psi\rangle=\textcolor{orange}{\tfrac{1}{2}}\,|0\rangle+\textcolor{violet}{\tfrac{1}{2}}\,|1\rangle
$$

$$
|\pi\rangle=\textcolor{blue}{\tfrac{1}{2}}\cdot\textcolor{orange}{\tfrac{1}{2}}\,|00\rangle+\textcolor{blue}{\tfrac{1}{2}}\cdot\textcolor{violet}{\tfrac{1}{2}}\,|01\rangle+\textcolor{teal}{\tfrac{1}{2}}\cdot\textcolor{orange}{\tfrac{1}{2}}\,|10\rangle+\textcolor{teal}{\tfrac{1}{2}}\cdot\textcolor{violet}{\tfrac{1}{2}}\,|11\rangle
$$

Tensor products of vectors

When working with multiple systems, we need a way to combine their vector spaces into a larger one. The mathematical operation that does this is the tensor product.

Given two vectors, $|\phi\rangle$ and $|\psi\rangle$, their tensor product, written $|\phi\rangle\otimes|\psi\rangle$, represents the combined state of both systems. Every basis state of the first vector is paired with every basis state of the second, and the corresponding coefficients are multiplied.

$$
|\phi\rangle=\sum_{a\in\Sigma}\alpha_a\,|a\rangle\quad\text{and}\quad|\psi\rangle=\sum_{b\in\Gamma}\beta_b\,|b\rangle
$$

$$
|\phi\rangle\otimes|\psi\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}\alpha_a\beta_b\,|ab\rangle
$$

- $|\phi\rangle$ — probability vector for system $X$, with coefficients $\alpha_a$
- $|\psi\rangle$ — probability vector for system $Y$, with coefficients $\beta_b$
- $\alpha_a\beta_b$ — coefficient of the compound basis state $|ab\rangle$ in $|\phi\rangle\otimes|\psi\rangle$
- $\Sigma\times\Gamma$ — the Cartesian product of the two state sets; the sum runs over all possible pairs $(a,b)$

Inner product form

Equivalently, the vector $|\pi\rangle=|\phi\rangle\otimes|\psi\rangle$ is defined by this condition:

$$
\langle ab|\pi\rangle=\langle a|\phi\rangle\langle b|\psi\rangle\qquad(\text{for all }a\in\Sigma\text{ and }b\in\Gamma)
$$

Here, $\langle a|\phi\rangle$ is a single number — the coefficient (amplitude) of basis state $|a\rangle$ in $|\phi\rangle$; and $\langle b|\psi\rangle$ is also a single number — the coefficient of basis state $|b\rangle$ in $|\psi\rangle$:

$$
\langle a|\phi\rangle=\alpha_a\qquad\langle b|\psi\rangle=\beta_b
$$

So the condition says: the coefficient of the combined basis state $|ab\rangle$ in $|\pi\rangle$ is the product of the coefficients of $|a\rangle$ and $|b\rangle$ separately. $\langle ab|\pi\rangle$ means: take the vector $|\pi\rangle$ and ask what its coefficient is along basis state $|ab\rangle$.

For example, suppose

$$
|\pi\rangle=0.3\,|00\rangle+\textcolor{teal}{0.4}\,|01\rangle+0.1\,|10\rangle+0.2\,|11\rangle
$$

Then $\langle \textcolor{teal}{01}|\pi\rangle=\textcolor{teal}{0.4}$. It is just extracting one coefficient.

Now suppose

$$
|\phi\rangle=\textcolor{blue}{\tfrac{1}{\sqrt{2}}}\,|0\rangle+\tfrac{1}{\sqrt{2}}\,|1\rangle\qquad\text{and}\qquad|\psi\rangle=\tfrac{3}{5}\,|0\rangle+\textcolor{orange}{\tfrac{4}{5}}\,|1\rangle
$$

Then:

$$
\langle \textcolor{blue}{0}|\phi\rangle=\textcolor{blue}{\tfrac{1}{\sqrt{2}}},\qquad\langle \textcolor{orange}{1}|\psi\rangle=\textcolor{orange}{\tfrac{4}{5}}
$$

Therefore:

$$
\langle \textcolor{blue}{0}\textcolor{orange}{1}|\pi\rangle=\langle \textcolor{blue}{0}|\phi\rangle\langle \textcolor{orange}{1}|\psi\rangle=\textcolor{blue}{\tfrac{1}{\sqrt{2}}}\cdot\textcolor{orange}{\tfrac{4}{5}}
$$

The amplitude of the combined outcome $01$ equals the product of the amplitudes of the individual outcomes $0$ and $1$.

Column vector form

The tensor product can be viewed as a "multiply every entry by every entry" operation. Each coefficient of the first vector is paired with every coefficient of the second, producing a larger vector that represents all possible combinations of the two systems. As a result, dimensions multiply: if $|\phi\rangle$ has $m$ entries and $|\psi\rangle$ has $k$ entries, then $|\phi\rangle\otimes|\psi\rangle$ has $m\cdot k$ entries.

$$
\begin{pmatrix}\alpha_1\\\vdots\\\alpha_m\end{pmatrix}\otimes\begin{pmatrix}\beta_1\\\vdots\\\beta_k\end{pmatrix}=\begin{pmatrix}\alpha_1\beta_1\\\vdots\\\alpha_1\beta_k\\\alpha_2\beta_1\\\vdots\\\alpha_2\beta_k\\\vdots\\\alpha_m\beta_1\\\vdots\\\alpha_m\beta_k\end{pmatrix}
$$

Tensor product of standard basis vectors

The tensor product of two standard basis vectors is often written by simply combining their labels into a single basis label. Thus, instead of writing $|a\rangle\otimes|b\rangle$, we commonly write $|ab\rangle$, which can be viewed as shorthand for the basis vector indexed by the pair $(a,b)$. More explicitly, one could write $|(a,b)\rangle$, but in practice the parentheses are usually omitted and the notation $|a,b\rangle$ is preferred. This follows a common mathematical convention of removing symbols that do not add information or eliminate ambiguity. From a mathematician's perspective, once the structure is understood, the parentheses are carrying no real content and can be safely discarded. Of course, for anyone still getting comfortable with Dirac notation, the more explicit form $|(a,b)\rangle$ can be a useful stepping stone — it makes the two-label structure impossible to miss, and once that structure feels natural, dropping the parentheses costs nothing.

For example, for two-bit systems where $\Sigma=\Gamma=\{0,1\}$, each of the four standard basis vectors of the compound system arises as a tensor product:

$$
|0\rangle\otimes|0\rangle=|00\rangle,\quad|0\rangle\otimes|1\rangle=|01\rangle,\quad|1\rangle\otimes|0\rangle=|10\rangle,\quad|1\rangle\otimes|1\rangle=|11\rangle
$$

To see this concretely, take $|0\rangle\otimes|1\rangle$ and apply the column-vector rule — multiply every entry of the first vector by every entry of the second:

$$
|0\rangle\otimes|1\rangle=\begin{pmatrix}1\\0\end{pmatrix}\otimes\begin{pmatrix}0\\1\end{pmatrix}=\begin{pmatrix}1\cdot0\\1\cdot1\\0\cdot0\\0\cdot1\end{pmatrix}=\begin{pmatrix}0\\1\\0\\0\end{pmatrix}=|01\rangle
$$

The result is exactly the standard basis vector $|01\rangle$ — confirming that the label shorthand and the column-vector computation agree.

The same shorthand is used for basis bras: $\langle a|\otimes\langle b|=\langle ab|$. For two-bit systems, $\langle 0|\otimes\langle 0|=\langle 00|$, $\langle 0|\otimes\langle 1|=\langle 01|$, $\langle 1|\otimes\langle 0|=\langle 10|$, and $\langle 1|\otimes\langle 1|=\langle 11|$.

Properties of tensor product

The tensor product is bilinear — it preserves the familiar rules of linearity in both of its arguments. You can distribute over addition and pull out scalar factors from either side independently.

First argument

$$
(|\phi_1\rangle+|\phi_2\rangle)\otimes|\psi\rangle=|\phi_1\rangle\otimes|\psi\rangle+|\phi_2\rangle\otimes|\psi\rangle
$$

$$
(c\,|\phi\rangle)\otimes|\psi\rangle=c\,(|\phi\rangle\otimes|\psi\rangle)
$$

For example, take $|\phi_1\rangle=|0\rangle$, $|\phi_2\rangle=|1\rangle$, $|\psi\rangle=|0\rangle$:

$$
(|0\rangle+|1\rangle)\otimes|0\rangle=\begin{pmatrix}1\\1\end{pmatrix}\otimes\begin{pmatrix}1\\0\end{pmatrix}=\begin{pmatrix}1\\0\\1\\0\end{pmatrix}=|00\rangle+|10\rangle
$$

For scalar multiplication, take $c=3$, $|\phi\rangle=|0\rangle$, $|\psi\rangle=|1\rangle$:

$$
(3\,|0\rangle)\otimes|1\rangle=\begin{pmatrix}3\\0\end{pmatrix}\otimes\begin{pmatrix}0\\1\end{pmatrix}=\begin{pmatrix}0\\3\\0\\0\end{pmatrix}=3\,|01\rangle
$$

Second argument

$$
|\phi\rangle\otimes(|\psi_1\rangle+|\psi_2\rangle)=|\phi\rangle\otimes|\psi_1\rangle+|\phi\rangle\otimes|\psi_2\rangle
$$

$$
|\phi\rangle\otimes(c\,|\psi\rangle)=c\,(|\phi\rangle\otimes|\psi\rangle)
$$

For example, take $|\phi\rangle=|1\rangle$, $|\psi_1\rangle=|0\rangle$, $|\psi_2\rangle=|1\rangle$:

$$
|1\rangle\otimes(|0\rangle+|1\rangle)=\begin{pmatrix}0\\1\end{pmatrix}\otimes\begin{pmatrix}1\\1\end{pmatrix}=\begin{pmatrix}0\\0\\1\\1\end{pmatrix}=|10\rangle+|11\rangle
$$

For scalar multiplication, take $c=2$, $|\phi\rangle=|0\rangle$, $|\psi\rangle=|1\rangle$:

$$
|0\rangle\otimes(2\,|1\rangle)=\begin{pmatrix}1\\0\end{pmatrix}\otimes\begin{pmatrix}0\\2\end{pmatrix}=\begin{pmatrix}0\\2\\0\\0\end{pmatrix}=2\,|01\rangle
$$

Multiple systems (multilinearity)

Tensor products generalize to three or more systems. If $|\phi_1\rangle,\ldots,|\phi_n\rangle$ are vectors, their tensor product $|\psi\rangle=|\phi_1\rangle\otimes\cdots\otimes|\phi_n\rangle$ is defined by the equation $\langle a_1\cdots a_n|\psi\rangle=\langle a_1|\phi_1\rangle\cdots\langle a_n|\phi_n\rangle$.

The bra $\langle a_1\cdots a_n|$ is shorthand for $\langle a_1|\otimes\langle a_2|\otimes\cdots\otimes\langle a_n|$ (tensor symbols are omitted for concise notation), so $\langle a_1\cdots a_n|\psi\rangle$ means applying that product bra to the tensor-product state. The equation says the larger inner product splits into matching ordinary overlaps:

$$
(\langle a_1|\otimes\cdots\otimes\langle a_n|)(|\phi_1\rangle\otimes\cdots\otimes|\phi_n\rangle)=\langle a_1|\phi_1\rangle\cdots\langle a_n|\phi_n\rangle
$$

For example, take $n=3$ with $|\phi_1\rangle=|0\rangle,\ |\phi_2\rangle=\tfrac{3}{5}\,|0\rangle+\tfrac{4}{5}\,|1\rangle,\ |\phi_3\rangle=|1\rangle$ and let $|\psi\rangle=|\phi_1\rangle\otimes|\phi_2\rangle\otimes|\phi_3\rangle$. To read off the amplitude of the specific outcome $(a_1,a_2,a_3)=(0,1,1)$, apply the formula directly — no need to expand the full tensor product first:

$$
\langle \textcolor{blue}{0}\textcolor{violet}{1}\textcolor{orange}{1}|\psi\rangle=\textcolor{blue}{\langle 0|\phi_1\rangle}\cdot\textcolor{violet}{\langle 1|\phi_2\rangle}\cdot\textcolor{orange}{\langle 1|\phi_3\rangle}=\textcolor{blue}{\langle 0|0\rangle}\cdot\textcolor{violet}{\langle 1|\bigl(\tfrac{3}{5}|0\rangle+\tfrac{4}{5}|1\rangle\bigr)}\cdot\textcolor{orange}{\langle 1|1\rangle}=\textcolor{blue}{1}\cdot\textcolor{violet}{\tfrac{4}{5}}\cdot\textcolor{orange}{1}=\tfrac{4}{5}
$$

The same value can be found by expanding the tensor product first, but that is a bit more work because we build the combined state and then apply the bra:

$$
\begin{aligned}|\psi\rangle&=|0\rangle\otimes\bigl(\tfrac{3}{5}|0\rangle+\tfrac{4}{5}|1\rangle\bigr)\otimes|1\rangle=\textcolor{teal}{\tfrac{3}{5}}\,|001\rangle+\textcolor{violet}{\tfrac{4}{5}}\,|011\rangle\\\langle \textcolor{blue}{0}\textcolor{violet}{1}\textcolor{orange}{1}|\psi\rangle&=\langle \textcolor{blue}{0}\textcolor{violet}{1}\textcolor{orange}{1}|\bigl(\textcolor{teal}{\tfrac{3}{5}}\,|001\rangle+\textcolor{violet}{\tfrac{4}{5}}\,|011\rangle\bigr)=\textcolor{teal}{\tfrac{3}{5}}\cdot0+\textcolor{violet}{\tfrac{4}{5}}\cdot1=\tfrac{4}{5}\end{aligned}
$$

Any other outcome follows the same pattern. For instance, $(a_1,a_2,a_3)=(0,0,1)$:

$$
\langle 001|\psi\rangle=\langle 0|\phi_1\rangle\cdot\langle 0|\phi_2\rangle\cdot\langle 1|\phi_3\rangle=1\cdot\tfrac{3}{5}\cdot1=\tfrac{3}{5}
$$

The direct formula is the cleanest way to get one amplitude, but the recursive view is useful when we want the whole combined vector. It peels off the last factor:

$$
|\phi_1\rangle\otimes\cdots\otimes|\phi_n\rangle=\bigl(|\phi_1\rangle\otimes\cdots\otimes|\phi_{n-1}\rangle\bigr)\otimes|\phi_n\rangle
$$

Using the same example, keep the two coefficients from $|\phi_2\rangle$ visible while peeling off $|\phi_3\rangle=|1\rangle$:

$$
\begin{aligned}
|\psi\rangle
&=\bigl(|\phi_1\rangle\otimes|\phi_2\rangle\bigr)\otimes\textcolor{orange}{|\phi_3\rangle}\\
&=\left(\begin{pmatrix}1\\0\end{pmatrix}\otimes\begin{pmatrix}\textcolor{teal}{\tfrac{3}{5}}\\\textcolor{violet}{\tfrac{4}{5}}\end{pmatrix}\right)\otimes\textcolor{orange}{\begin{pmatrix}0\\1\end{pmatrix}}\\
&=\begin{pmatrix}\textcolor{teal}{\tfrac{3}{5}}\\\textcolor{violet}{\tfrac{4}{5}}\\0\\0\end{pmatrix}\otimes\textcolor{orange}{\begin{pmatrix}0\\1\end{pmatrix}}\\
&=\begin{pmatrix}0\\\textcolor{teal}{\tfrac{3}{5}}\\0\\\textcolor{violet}{\tfrac{4}{5}}\\0\\0\\0\\0\end{pmatrix}\\
&=\textcolor{teal}{\tfrac{3}{5}}\,|001\rangle+\textcolor{violet}{\tfrac{4}{5}}\,|011\rangle
\end{aligned}
$$

Measurement of probabilistic states

Consider the compound system $(X,Y)$ in the probabilistic state: $\tfrac{1}{2}|00\rangle+\tfrac{1}{2}|11\rangle$

Measuring all systems

Measuring the entire compound system at once is equivalent to measuring each subsystem independently — provided *all* systems are measured. The measurement produces a single combined outcome drawn from the joint probability distribution.

For this state the two possible combined outcomes are $|00\rangle$ and $|11\rangle$, each with probability $\tfrac{1}{2}$. Each individual measurement still happens according to those probabilities, but the whole state is measured together — producing exactly one combined result at a time: either $|00\rangle$ or $|11\rangle$.

Measuring some systems

Suppose only a subset of systems is measured — for example, only $X$ from $(X,Y)$. Then, to find the probability that $X$ equals some value $a$, add up probabilities of all outcomes where $X=a$:

$$
\Pr(X=a)=\sum_{b\in\Gamma}\Pr\bigl((X,Y)=(a,b)\bigr)
$$

For the probabilistic state $\tfrac{1}{2}|00\rangle+\tfrac{1}{2}|11\rangle$, the table below expands the Dirac notation into the four possible joint outcomes of $(X,Y)$ and their probabilities.

| X | Y | Pr |
| --- | --- | --- |
| 0 | 0 | $\tfrac{1}{2}$ |
| 0 | 1 | $0$ |
| 1 | 0 | $0$ |
| 1 | 1 | $\tfrac{1}{2}$ |

$$
\Pr(X=0)=\Pr\bigl((X,Y)=(0,0)\bigr)+\Pr\bigl((X,Y)=(0,1)\bigr)=\tfrac{1}{2}+0=\tfrac{1}{2}
$$

$$
\Pr(X=1)=\Pr\bigl((X,Y)=(1,0)\bigr)+\Pr\bigl((X,Y)=(1,1)\bigr)=0+\tfrac{1}{2}=\tfrac{1}{2}
$$

But what happens to our knowledge of $Y$?

After measuring $X$ and getting result $a$, there can still be uncertainty about the state of $Y$. If we already measured $X$ and got $a$, what is the probability that $Y=b$?

$$
\Pr(Y=b\mid X=a)=
\begin{aligned}
&\frac{\Pr\bigl((X,Y)=(a,b)\bigr)}{\Pr(X=a)}
\quad
\begin{array}{l}
\leftarrow\;\text{Probability that both }X=a\text{ and }Y=b\text{ happen}\\[-1pt]
\leftarrow\;\text{Probability that }X=a\text{ happened}
\end{array}
\end{aligned}
$$

Think of the division as rescaling the probabilities so they add up to $1$ again after we focus on only one situation. For example, once we learn that $X=1$, we ignore all outcomes where $X=0$. The probabilities of the remaining outcomes no longer add up to $1$, because part of the original probability space was removed.

To turn the remaining outcomes into a valid probability distribution again, we divide each remaining probability by the total probability of $X=1$. This process is called normalization. So the division means: out of the world where $X=1$ happened, how likely is each remaining outcome?

If we measured $X=0$

$$
\begin{aligned}
\Pr(Y=0\mid X=0)&=\frac{\Pr\bigl((X,Y)=(0,0)\bigr)}{\Pr(X=0)}=\frac{\tfrac{1}{2}}{\tfrac{1}{2}}=1\\
\Pr(Y=1\mid X=0)&=\frac{\Pr\bigl((X,Y)=(0,1)\bigr)}{\Pr(X=0)}=\frac{0}{\tfrac{1}{2}}=0
\end{aligned}
$$

If we measured $X=1$

$$
\begin{aligned}
\Pr(Y=0\mid X=1)&=\frac{\Pr\bigl((X,Y)=(1,0)\bigr)}{\Pr(X=1)}=\frac{0}{\tfrac{1}{2}}=0\\
\Pr(Y=1\mid X=1)&=\frac{\Pr\bigl((X,Y)=(1,1)\bigr)}{\Pr(X=1)}=\frac{\tfrac{1}{2}}{\tfrac{1}{2}}=1
\end{aligned}
$$

Measuring one system, in Dirac notation

Everything above can be expressed directly on the state vector:

$$
\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|ab\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|a\rangle\otimes|b\rangle=\sum_{a\in\Sigma}|a\rangle\otimes\Bigl(\sum_{b\in\Gamma}p_{ab}\,|b\rangle\Bigr)
$$

Reading the chain left to right:

- $\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|ab\rangle$ — a general probabilistic state of $(X,Y)$: the sum over all possible joint outcomes $(a,b)$, each weighted by its probability $p_{ab}$.
- $\sum p_{ab}\,|a\rangle\otimes|b\rangle$ — split every compound basis state into a tensor product, $|ab\rangle=|a\rangle\otimes|b\rangle$, separating the $X$ part from the $Y$ part.
- $\sum_{a\in\Sigma}|a\rangle\otimes\bigl(\sum_{b\in\Gamma}p_{ab}\,|b\rangle\bigr)$
   
   — use bilinearity to pull the shared $|a\rangle$ out of the inner sum. This groups the terms by the value of $X$: one branch per value $a$, with all of $Y$'s weight collected inside the parentheses.

For example, for the probabilistic state $\tfrac{1}{2}|00\rangle+\tfrac{1}{2}|11\rangle$, we can split the compound state $|00\rangle$ into $|0\rangle\otimes|0\rangle$, and $|11\rangle$ into $|1\rangle\otimes|1\rangle$, and group by $X$:

$$
\tfrac{1}{2}|00\rangle+\tfrac{1}{2}|11\rangle=\tfrac{1}{2}\bigl(|0\rangle\otimes|0\rangle\bigr)+\tfrac{1}{2}\bigl(|1\rangle\otimes|1\rangle\bigr)=
$$

$|0\rangle$↑if X=|0⟩

$$
\otimes
$$

$\bigl(\tfrac{1}{2}|0\rangle\bigr)$↑Y has this distribution

$$
+
$$

$|1\rangle$↑if X=|1⟩

$$
\otimes
$$

$\bigl(\tfrac{1}{2}|1\rangle\bigr)$↑Y has this distribution

To get the probability of measuring $X=a$, add the probabilities of all states that start with $a$:

$$
\Pr(X=a)=\sum_{b\in\Gamma}p_{ab}
$$

Where:

- $p_{ab}$ — the probability of the joint outcome $(a,b)$. The first index $a$ is the value of $X$, the second index $b$ is the value of $Y$ — so $p_{01}$ means $\Pr\bigl((X,Y)=(0,1)\bigr)$.
- $\sum_{b\in\Gamma}$ — keep $X$ fixed at $a$ and sweep over every possible $Y$; that is, add up every outcome that starts with $a$.

After measuring $X=a$, all branches with other values of $X$ disappear. The remaining state of $Y$ is the surviving branch, divided by its total weight so the coefficients add back up to $1$:

$$
\frac{\sum_{b\in\Gamma}p_{ab}\,|b\rangle}{\Pr(X=a)}\qquad\text{where}\qquad\Pr(X=a)=\sum_{c\in\Gamma}p_{ac}
$$

The index in the normalizing sum is just a dummy variable: writing $c$ instead of $b$ avoids clashing with the $b$ in the numerator, but both range over all of $\Gamma$. What it really collects is every outcome whose first coordinate is $a$.

Example

Take the probabilistic state of $(X,Y)$:

$$
\tfrac{1}{12}\,|00\rangle+\tfrac{1}{4}\,|01\rangle+\tfrac{1}{3}\,|10\rangle+\tfrac{1}{3}\,|11\rangle
$$

We measure only $X$ (the first bit). Grouping by $X$ as above:

$$
|0\rangle\otimes\Bigl(\tfrac{1}{12}\,|0\rangle+\tfrac{1}{4}\,|1\rangle\Bigr)+|1\rangle\otimes\Bigl(\tfrac{1}{3}\,|0\rangle+\tfrac{1}{3}\,|1\rangle\Bigr)
$$

If we measured $X=0$

$$
\Pr(X=0)=\tfrac{1}{12}+\tfrac{1}{4}=\tfrac{1}{3}
$$

The probabilistic state of $Y$ becomes (normalize by $\tfrac{1}{3}$):

$$
\frac{\tfrac{1}{12}\,|0\rangle+\tfrac{1}{4}\,|1\rangle}{\tfrac{1}{3}}=\tfrac{1}{4}\,|0\rangle+\tfrac{3}{4}\,|1\rangle
$$

If we measured $X=1$

$$
\Pr(X=1)=\tfrac{1}{3}+\tfrac{1}{3}=\tfrac{2}{3}
$$

The probabilistic state of $Y$ becomes (normalize by $\tfrac{2}{3}$):

$$
\frac{\tfrac{1}{3}\,|0\rangle+\tfrac{1}{3}\,|1\rangle}{\tfrac{2}{3}}=\tfrac{1}{2}\,|0\rangle+\tfrac{1}{2}\,|1\rangle
$$

Measuring some: Y

Just like we grouped by $X$, we can apply the same principle and group by $Y$ — regroup the same sum by $|b\rangle$ instead of $|a\rangle$:

$$
\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|ab\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}p_{ab}\,|a\rangle\otimes|b\rangle=\sum_{b\in\Gamma}
$$

$\Bigl(\sum_{a\in\Sigma}p_{ab}\,|a\rangle\Bigr)$↑X has this distribution

$$
\otimes
$$

$|b\rangle$↑if Y=b

To get the probability of measuring $Y=b$, add the probabilities of all states that end with $b$:

$$
\Pr(Y=b)=\sum_{a\in\Sigma}p_{ab}
$$

After measuring $Y=b$, all branches with other values of $Y$ disappear, and the remaining state of $X$ is that branch normalized by its total weight:

$$
\frac{\sum_{a\in\Sigma}p_{ab}\,|a\rangle}{\Pr(Y=b)}\qquad\text{where}\qquad\Pr(Y=b)=\sum_{c\in\Sigma}p_{cb}
$$

Example

Take the same state, but this time measure only $Y$ (the second bit):

$$
\tfrac{1}{12}\,|00\rangle+\tfrac{1}{4}\,|01\rangle+\tfrac{1}{3}\,|10\rangle+\tfrac{1}{3}\,|11\rangle
$$

Grouping by $Y$ this time:

$$
\Bigl(\tfrac{1}{12}\,|0\rangle+\tfrac{1}{3}\,|1\rangle\Bigr)\otimes|0\rangle+\Bigl(\tfrac{1}{4}\,|0\rangle+\tfrac{1}{3}\,|1\rangle\Bigr)\otimes|1\rangle
$$

If we measured $Y=0$

$$
\Pr(Y=0)=\tfrac{1}{12}+\tfrac{1}{3}=\tfrac{5}{12}
$$

The probabilistic state of $X$ becomes (normalize by $\tfrac{5}{12}$):

$$
\frac{\tfrac{1}{12}\,|0\rangle+\tfrac{1}{3}\,|1\rangle}{\tfrac{5}{12}}=\tfrac{1}{5}\,|0\rangle+\tfrac{4}{5}\,|1\rangle
$$

If we measured $Y=1$

$$
\Pr(Y=1)=\tfrac{1}{4}+\tfrac{1}{3}=\tfrac{7}{12}
$$

The probabilistic state of $X$ becomes (normalize by $\tfrac{7}{12}$):

$$
\frac{\tfrac{1}{4}\,|0\rangle+\tfrac{1}{3}\,|1\rangle}{\tfrac{7}{12}}=\tfrac{3}{7}\,|0\rangle+\tfrac{4}{7}\,|1\rangle
$$

Note: for classical, probabilistic systems these conditional probabilities are easier to handle with simple conditional-probability tables. But later, with quantum states, those simple tables stop working — so Dirac notation it is (I am starting to get used to it, but I understand the struggle, deeply!).

Operations on probabilistic states

Probabilistic operations on compound systems, just like for individual systems, are represented by stochastic matrices. But this time the matrices have rows and columns corresponding to the Cartesian product of the individual systems' classical state sets.

A deterministic example: controlled-NOT

Take a controlled-NOT operation on two bits $(X,Y)$:

- if $x=1$ ⇒ apply NOT to $y$;
- if $x=0$ ⇒ do nothing.

Here $X$ is the control bit and $Y$ is the target bit. It is deterministic — each input state maps to exactly one output state:

Input → output

$$
\begin{aligned}|00\rangle&\mapsto|00\rangle\\|01\rangle&\mapsto|01\rangle\\|10\rangle&\mapsto|11\rangle\\|11\rangle&\mapsto|10\rangle\end{aligned}
$$

Matrix representation

$$
M=\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}
$$

For example, applying the matrix operation to $|10\rangle$ flips the target, giving $|11\rangle$:

$$
M\,|10\rangle=\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}\begin{pmatrix}0\\0\\1\\0\end{pmatrix}=\begin{pmatrix}0\\0\\0\\1\end{pmatrix}=|11\rangle
$$

Nothing special is going on — the operation is just a matrix acting on the probability vector.

A probabilistic example

A probabilistic operation makes a random choice. For example:

- with probability $\tfrac{1}{2}$ ⇒ set $y=x$;
- with probability $\tfrac{1}{2}$ ⇒ set $x=y$.

Each branch is itself a deterministic stochastic matrix. Take the input $|10\rangle$ to see how both behave:

Set $y=x$

$$
\begin{pmatrix}1&1&0&0\\0&0&0&0\\0&0&0&0\\0&0&1&1\end{pmatrix}\begin{pmatrix}0\\0\\1\\0\end{pmatrix}=\begin{pmatrix}0\\0\\0\\1\end{pmatrix}=|11\rangle
$$

Set $x=y$

$$
\begin{pmatrix}1&0&1&0\\0&0&0&0\\0&0&0&0\\0&1&0&1\end{pmatrix}\begin{pmatrix}0\\0\\1\\0\end{pmatrix}=\begin{pmatrix}1\\0\\0\\0\end{pmatrix}=|00\rangle
$$

The overall probabilistic operation is the average of these two matrices, weighted by the probabilities $\tfrac{1}{2}$:

$$
\begin{pmatrix}1&\tfrac12&\tfrac12&0\\0&0&0&0\\0&0&0&0\\0&\tfrac12&\tfrac12&1\end{pmatrix}=\tfrac12\begin{pmatrix}1&1&0&0\\0&0&0&0\\0&0&0&0\\0&0&1&1\end{pmatrix}+\tfrac12\begin{pmatrix}1&0&1&0\\0&0&0&0\\0&0&0&0\\0&1&0&1\end{pmatrix}
$$

The resulting matrix does not describe one concrete execution — it describes the distribution over outcomes after the random choice.

Simultaneous operations on a compound system

Suppose instead that two probabilistic operations act on separate systems — $M$ on $X$ and $N$ on $Y$, each its own stochastic matrix on its own probability vector. If we perform both at the same time, how do we describe their combined effect on the $(X,Y)$ compound system?

Tensor product of matrices

Performing $M$ on $X$ and $N$ on $Y$ simultaneously is described by a single matrix on the compound system — the tensor product $M\otimes N$. Writing each operation in Dirac notation,

$$
M=\sum_{a,b\in\Sigma}\alpha_{ab}\,|a\rangle\langle b|\qquad N=\sum_{c,d\in\Gamma}\beta_{cd}\,|c\rangle\langle d|
$$

the tensor product pairs every term of one with every term of the other, multiplying their coefficients:

$$
M\otimes N=\sum_{a,b\in\Sigma}\;\sum_{c,d\in\Gamma}\alpha_{ab}\,\beta_{cd}\,|ac\rangle\langle bd|
$$

- $\alpha_{ab}$ — the entry of $M$ in row $a$, column $b$
- $\beta_{cd}$ — the entry of $N$ in row $c$, column $d$
- $|ac\rangle\langle bd|$ — the compound outer product, using the identity
   
   $|a\rangle\langle b|\otimes|c\rangle\langle d|=|ac\rangle\langle bd|$

Example: bit‑flip $\otimes$ identity

First rewrite each in Dirac notation by reading off its entries: the entry in row $a$, column $b$ is the coefficient of $|a\rangle\langle b|$:

$M$ — bit‑flip on $X$

$$
M=\begin{pmatrix}\textcolor{blue}{0}&\textcolor{orange}{1}\\[2pt]\textcolor{teal}{1}&\textcolor{violet}{0}\end{pmatrix}\qquad\begin{array}{l}\textcolor{blue}{\alpha_{00}=0}\;\longrightarrow\;\textcolor{blue}{0\,|0\rangle\langle 0|}\\[4pt]\textcolor{orange}{\alpha_{01}=1}\;\longrightarrow\;\textcolor{orange}{1\,|0\rangle\langle 1|}\\[4pt]\textcolor{teal}{\alpha_{10}=1}\;\longrightarrow\;\textcolor{teal}{1\,|1\rangle\langle 0|}\\[4pt]\textcolor{violet}{\alpha_{11}=0}\;\longrightarrow\;\textcolor{violet}{0\,|1\rangle\langle 1|}\end{array}
$$

Written out and simplified:

$$
\begin{aligned}M&=\textcolor{blue}{0}\,|0\rangle\langle 0|+\textcolor{orange}{1}\,|0\rangle\langle 1|+\textcolor{teal}{1}\,|1\rangle\langle 0|+\textcolor{violet}{0}\,|1\rangle\langle 1|\\[4pt]&=\textcolor{orange}{|0\rangle\langle 1|}+\textcolor{teal}{|1\rangle\langle 0|}\end{aligned}
$$

$N$ — identity on $Y$

$$
N=\begin{pmatrix}\textcolor{#dc2626}{1}&\textcolor{#15803d}{0}\\[2pt]\textcolor{#b45309}{0}&\textcolor{#db2777}{1}\end{pmatrix}\qquad\begin{array}{l}\textcolor{#dc2626}{\beta_{00}=1}\;\longrightarrow\;\textcolor{#dc2626}{1\,|0\rangle\langle 0|}\\[4pt]\textcolor{#15803d}{\beta_{01}=0}\;\longrightarrow\;\textcolor{#15803d}{0\,|0\rangle\langle 1|}\\[4pt]\textcolor{#b45309}{\beta_{10}=0}\;\longrightarrow\;\textcolor{#b45309}{0\,|1\rangle\langle 0|}\\[4pt]\textcolor{#db2777}{\beta_{11}=1}\;\longrightarrow\;\textcolor{#db2777}{1\,|1\rangle\langle 1|}\end{array}
$$

Written out and simplified:

$$
\begin{aligned}N&=\textcolor{#dc2626}{1}\,|0\rangle\langle 0|+\textcolor{#15803d}{0}\,|0\rangle\langle 1|+\textcolor{#b45309}{0}\,|1\rangle\langle 0|+\textcolor{#db2777}{1}\,|1\rangle\langle 1|\\[4pt]&=\textcolor{#dc2626}{|0\rangle\langle 0|}+\textcolor{#db2777}{|1\rangle\langle 1|}\end{aligned}
$$

Method 1 — Kronecker blocks

Replace each entry of $M$ with the block $M_{ij}\,N$:

$M\otimes N=\begin{pmatrix}\textcolor{blue}{0}&\textcolor{orange}{1}\\\textcolor{teal}{1}&\textcolor{violet}{0}\end{pmatrix}\otimes\begin{pmatrix}1&0\\0&1\end{pmatrix}=\begin{pmatrix}\textcolor{blue}{0\cdot N}&\textcolor{orange}{1\cdot N}\\\textcolor{teal}{1\cdot N}&\textcolor{violet}{0\cdot N}\end{pmatrix}=$$\begin{pmatrix}0&0&1&0\\0&0&0&1\\1&0&0&0\\0&1&0&0\end{pmatrix}$

Method 2 — Dirac expansion

Distribute (bilinearity), collapse each term with the identity, then keep the equality chain going by turning each outer product into its matrix:

$$
|a\rangle\langle b|\otimes|c\rangle\langle d|=|ac\rangle\langle bd|
$$

$$
\begin{aligned}
M\otimes N
&=\bigl(|0\rangle\langle 1|+|1\rangle\langle 0|\bigr)\otimes\bigl(|0\rangle\langle 0|+|1\rangle\langle 1|\bigr)\\[4pt]
&=|0\rangle\langle 1|\otimes|0\rangle\langle 0|+|0\rangle\langle 1|\otimes|1\rangle\langle 1|+|1\rangle\langle 0|\otimes|0\rangle\langle 0|+|1\rangle\langle 0|\otimes|1\rangle\langle 1|\\[4pt]
&=|00\rangle\langle 10|+|01\rangle\langle 11|+|10\rangle\langle 00|+|11\rangle\langle 01|\\[4pt]
&=\begin{pmatrix}0&0&1&0\\0&0&0&0\\0&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&1\\0&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&0\\1&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&0\\0&0&0&0\\0&1&0&0\end{pmatrix}\\[4pt]
&=\begin{pmatrix}0&0&1&0\\0&0&0&1\\1&0&0&0\\0&1&0&0\end{pmatrix}
\end{aligned}
$$

Key idea of Dirac notation

A sandwich $\langle a|M|b\rangle$ is just the single entry of $M$ at row $a$, column $b$. The bra $\langle a|$ is the unit row vector that picks out row $a$, and the ket $|b\rangle$ is the unit column vector that picks out column $b$ — multiplying them on either side of $M$ collapses the whole matrix down to one number. Indices count from $0$, so $\langle 0|M|0\rangle$ is the top-left entry:

$$
M=\begin{pmatrix}\textcolor{#0369a1}{1}&2\\3&4\end{pmatrix}\qquad\langle 0|M|0\rangle=\underbrace{\begin{pmatrix}1&0\end{pmatrix}}_{\langle 0|}\begin{pmatrix}1&2\\3&4\end{pmatrix}\underbrace{\begin{pmatrix}1\\0\end{pmatrix}}_{|0\rangle}=\textcolor{#0369a1}{1}
$$

Equivalent entry rule

Equivalently, the tensor product is the matrix whose compound entry is found by multiplying the matching entry from $M$ with the matching entry from $N$:

$$
\langle ac|M\otimes N|bd\rangle=\langle a|M|b\rangle\,\langle c|N|d\rangle\qquad\text{for all }a,b\in\Sigma\text{ and }c,d\in\Gamma
$$

- Read $\langle ac|M\otimes N|bd\rangle$ from right to left: $|bd\rangle$ is the starting column, $M\otimes N$ is the matrix you apply, and $\langle ac|$ is the output row you read. For example, let's calculate $\langle 10|M\otimes N|00\rangle$. Start with the input column $|00\rangle$, apply $M\otimes N$, then read the value in the output row $\langle 10|$ — the result is $1$.
   
   columns in $|bd\rangle$
   
   $|00\rangle$
   
   $|01\rangle$
   
   $|10\rangle$
   
   $|11\rangle$
   
   $\downarrow$
   
   $M\otimes N=$
   
   $\langle 00|$
   
   $\langle 01|$
   
   $\langle 10|$
   
   $\langle 11|$
   
   0
   
   0
   
   1
   
   0
   
   0
   
   0
   
   0
   
   1
   
   1
   
   0
   
   0
   
   0
   
   0
   
   1
   
   0
   
   0
   
   $\leftarrow$rows out $\langle ac|=\langle 10|$
   
   $\langle 10|M\otimes N|00\rangle=1$
- The right side asks the same question one system at a time: $\langle a|M|b\rangle$ asks how much $M$ sends $|b\rangle$ to $|a\rangle$, and $\langle c|N|d\rangle$ asks how much $N$ sends $|d\rangle$ to $|c\rangle$. For the same example, look up each factor in its own matrix — $\langle 1|M|0\rangle$ is the entry of $M$ in row $\langle 1|$, column $|0\rangle$, and $\langle 0|N|0\rangle$ likewise for $N$ — then multiply them:
   
   $M=$
   
   $|0\rangle$
   
   $|1\rangle$
   
   $\langle 0|$
   
   $\langle 1|$
   
   0
   
   1
   
   1
   
   0
   
   $N=$
   
   $|0\rangle$
   
   $|1\rangle$
   
   $\langle 0|$
   
   $\langle 1|$
   
   1
   
   0
   
   0
   
   1
   
   $\langle 1|M|0\rangle\cdot\langle 0|N|0\rangle=1\cdot 1=1$
- In this example, $M$ flips the first bit and $N$ leaves the second unchanged, so $M\otimes N$ sends $|00\rangle$ straight to $|10\rangle$. The resulting $1$ is the amplitude of that transition: it says all of the input lands on $|10\rangle$ and none on any other basis state — the mapping is exact and deterministic (a coefficient of $0$ would mean $|00\rangle$ never reaches $|10\rangle$).

Why the entry rule holds

We want the entry of $M\otimes N$ at row $\langle ac|$ and column $|bd\rangle$. The rule says: take the matching entry from $M$, take the matching entry from $N$, then multiply them.

$$
\langle ac|M\otimes N|bd\rangle=\langle a|M|b\rangle\,\langle c|N|d\rangle
$$

Derivation

$$
\langle ac|M\otimes N|bd\rangle
$$

The compound entry we want.

$$
=\langle ac|\left(\sum_{i,j\in\Sigma}\alpha_{ij}|i\rangle\langle j|\right)\otimes\left(\sum_{k,l\in\Gamma}\beta_{kl}|k\rangle\langle l|\right)|bd\rangle
$$

Write $M$ and $N$ as sums of weighted outer products.

$$
=\sum_{i,j\in\Sigma}\sum_{k,l\in\Gamma}\alpha_{ij}\beta_{kl}\,\langle ac|ik\rangle\,\langle jl|bd\rangle
$$

Linearity pulls the coefficients out front, and $|i\rangle\langle j|\otimes|k\rangle\langle l|$ becomes $|ik\rangle\langle jl|$.

$$
=\sum_{i,j\in\Sigma}\sum_{k,l\in\Gamma}\alpha_{ij}\beta_{kl}\,\langle a|i\rangle\,\langle c|k\rangle\,\langle j|b\rangle\,\langle l|d\rangle
$$

Each compound overlap breaks into one bracket per system.

$$
=\sum_{i,j\in\Sigma}\sum_{k,l\in\Gamma}\alpha_{ij}\beta_{kl}\,\delta_{ai}\delta_{ck}\delta_{jb}\delta_{ld}
$$

Those brackets vanish unless $i=a$, $j=b$, $k=c$, $l=d$.

$$
=\alpha_{ab}\beta_{cd}
$$

Only that single term survives the double sum.

$$
=\langle a|M|b\rangle\,\langle c|N|d\rangle
$$

Exactly the per-system entry rule we set out to prove.

Selector picture

Focus on the $\alpha_{ij}$ part first. The two brackets are two tests on the same cell:

- $\langle a|i\rangle$ asks: is this cell in row $i=a$?
- $\langle j|b\rangle$ asks: is this cell in column $j=b$?

A cell must pass both tests. If it fails either one, it gets multiplied by $0$.

$$
\alpha_{ij}\langle a|i\rangle\langle j|b\rangle=
\begin{cases}
\alpha_{ab},&i=a\text{ and }j=b\\
0,&\text{otherwise}
\end{cases}
$$

The $\beta_{kl}$ part does the same thing with $k=c$ and $l=d$, so it leaves $\beta_{cd}$.

Example: $a=1$, $b=1$

$$
j=0
$$

$$
j=1
$$

$$
j=2
$$

$$
i=0
$$

$$
\alpha_{00}
$$

$$
\alpha_{01}
$$

$$
\alpha_{02}
$$

$$
i=1
$$

$$
\alpha_{10}
$$

$$
\alpha_{11}
$$

$$
\alpha_{12}
$$

$$
i=2
$$

$$
\alpha_{20}
$$

$$
\alpha_{21}
$$

$$
\alpha_{22}
$$

Pale cells pass one test but fail the other. The highlighted cell passes both, so it is the only one left.

Equivalent action on product states

Equivalently, $M\otimes N$ is the unique matrix that satisfies the equation for all vectors $|\varphi\rangle$ and $|\psi\rangle$:

$$
(M\otimes N)\,\bigl(|\varphi\rangle\otimes|\psi\rangle\bigr)=\bigl(M|\varphi\rangle\bigr)\otimes\bigl(N|\psi\rangle\bigr)
$$

The tensor product is defined by one rule: apply $M$ to the first subsystem and $N$ to the second. Because every state can be built from basis states by adding them together, and linear maps preserve addition, this rule completely determines the full matrix.

Explicit formula

Collecting the entry rule into a single matrix gives the explicit form of the tensor product. Each entry is the product $\alpha_{ab}\,\beta_{cd}=\langle ac|M\otimes N|bd\rangle$, so for $M$ with entries up to $\alpha_{mm}$ and $N$ with entries up to $\beta_{nn}$:

$$
\begin{aligned}M\otimes N&=\begin{pmatrix}\alpha_{00}&\cdots&\alpha_{0m}\\\vdots&\ddots&\vdots\\\alpha_{m0}&\cdots&\alpha_{mm}\end{pmatrix}\otimes\begin{pmatrix}\beta_{00}&\cdots&\beta_{0n}\\\vdots&\ddots&\vdots\\\beta_{n0}&\cdots&\beta_{nn}\end{pmatrix}\\[8pt]&=\begin{pmatrix}\alpha_{00}\beta_{00}&\cdots&\alpha_{00}\beta_{0n}&\cdots&\alpha_{0m}\beta_{00}&\cdots&\alpha_{0m}\beta_{0n}\\[2pt]\vdots&\ddots&\vdots&&\vdots&\ddots&\vdots\\[2pt]\alpha_{00}\beta_{n0}&\cdots&\alpha_{00}\beta_{nn}&\cdots&\alpha_{0m}\beta_{n0}&\cdots&\alpha_{0m}\beta_{nn}\\[2pt]\vdots&&\vdots&\ddots&\vdots&&\vdots\\[2pt]\alpha_{m0}\beta_{00}&\cdots&\alpha_{m0}\beta_{0n}&\cdots&\alpha_{mm}\beta_{00}&\cdots&\alpha_{mm}\beta_{0n}\\[2pt]\vdots&\ddots&\vdots&&\vdots&\ddots&\vdots\\[2pt]\alpha_{m0}\beta_{n0}&\cdots&\alpha_{m0}\beta_{nn}&\cdots&\alpha_{mm}\beta_{n0}&\cdots&\alpha_{mm}\beta_{nn}\end{pmatrix}\end{aligned}
$$

Three or more matrices

Nothing about the entry rule was special to two systems. For $n$ factors $M_1\otimes\cdots\otimes M_n$, a compound entry still splits into one bracket per system — each $M_k$ is asked the same question about its own bits in isolation, and the answers multiply:

$$
\langle a_1\cdots a_n|\,M_1\otimes\cdots\otimes M_n\,|b_1\cdots b_n\rangle=\langle a_1|M_1|b_1\rangle\;\langle a_2|M_2|b_2\rangle\cdots\langle a_n|M_n|b_n\rangle
$$

The product also stays compatible with ordinary matrix multiplication. Running one stack of operators after another is the same as multiplying the matrices system by system (the mixed‑product property):

$$
\bigl(M_1\otimes\cdots\otimes M_n\bigr)\bigl(N_1\otimes\cdots\otimes N_n\bigr)=\bigl(M_1N_1\bigr)\otimes\cdots\otimes\bigl(M_nN_n\bigr)
$$

## Multiple systems: quantum

Quantum states

A quantum state of several systems is represented by a column vector whose indices correspond to the Cartesian product of the individual systems' classical state sets — exactly the same index set as the compound classical system, now carrying complex amplitudes instead of probabilities.

For two systems $X$ with state set $\Sigma$ and $Y$ with state set $\Gamma$, the entries are indexed by $\Sigma\times\Gamma$. If both are bits, the four indices are:

$$
\{0,1\}\times\{0,1\}=\{00,\,01,\,10,\,11\}
$$

So a quantum state of the two-bit system $XY$ is a four-entry column vector. Written as a combination of the standard basis states $|00\rangle,|01\rangle,|10\rangle,|11\rangle$:

$$
|\psi\rangle=\alpha_{00}\,|00\rangle+\alpha_{01}\,|01\rangle+\alpha_{10}\,|10\rangle+\alpha_{11}\,|11\rangle=\begin{pmatrix}\alpha_{00}\\\alpha_{01}\\\alpha_{10}\\\alpha_{11}\end{pmatrix}
$$

As with a single system, the amplitudes are complex numbers and the vector is a unit vector: the squared absolute values sum to one, and $|\alpha_{ab}|^2$ is the probability of measuring the pair $(a,b)$.

$$
\sum_{(a,b)\in\Sigma\times\Gamma}|\alpha_{ab}|^2=1
$$

Definite (basis) state

$$
|10\rangle=\begin{pmatrix}0\\0\\1\\0\end{pmatrix}
$$

$X$ is certainly $|1\rangle$ and $Y$ is certainly $|0\rangle$.

Equal superposition

$$
|\psi\rangle=\tfrac{1}{2}\bigl(|00\rangle+|01\rangle+|10\rangle+|11\rangle\bigr)=\tfrac{1}{2}\begin{pmatrix}1\\1\\1\\1\end{pmatrix}
$$

All four outcomes are equally likely, each with probability $\bigl(\tfrac{1}{2}\bigr)^2=\tfrac{1}{4}$.

Biased superposition

$$
|\psi\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{2}\,|01\rangle+\tfrac{1}{2}\,|11\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{2}\\[4pt]0\\[4pt]\tfrac{1}{2}\end{pmatrix}
$$

Still a unit vector: $\tfrac{1}{2}+\tfrac{1}{4}+\tfrac{1}{4}=1$.

Entangled (Bell) state

$$
|\phi^{+}\rangle=\tfrac{1}{\sqrt{2}}\bigl(|00\rangle+|11\rangle\bigr)=\tfrac{1}{\sqrt{2}}\begin{pmatrix}1\\0\\0\\1\end{pmatrix}
$$

Measuring gives $|00\rangle$ or $|11\rangle$ with equal probability and never $|01\rangle$ or $|10\rangle$, so the two bits always come out the same. Unlike the states above, it cannot be factored into a separate state for each bit — no $|\psi_X\rangle\otimes|\psi_Y\rangle$ equals $|\phi^{+}\rangle$. That inseparability is what entanglement means.

In general, combining $n$ systems with state sets $\Sigma_1,\ldots,\Sigma_n$ gives a state vector indexed by $\Sigma_1\times\cdots\times\Sigma_n$, so its dimension is the product $|\Sigma_1|\cdots|\Sigma_n|$ of the individual sizes — for $n$ qubits, $2^n$ amplitudes.

Tensor products of states

The previous states described one compound system directly. We can also build a compound state by combining states of the parts: the tensor product of two quantum state vectors is again a quantum state vector.

Let $|\phi\rangle$ be a state of system $X$ and $|\psi\rangle$ a state of system $Y$. Their tensor product is a state of the joint system $(X,Y)$:

$$
|\phi\rangle\otimes|\psi\rangle
$$

States of this form are called product states. They describe the two systems acting independently — each part has its own well-defined state, with no correlation between them. (Entangled states like $|\phi^{+}\rangle$ cannot be written this way.)

More generally, if $|\psi_1\rangle,\ldots,|\psi_n\rangle$ are states of systems $X_1,\ldots,X_n$, then their tensor product is a product state of the whole compound system $(X_1,\ldots,X_n)$:

$$
|\psi_1\rangle\otimes\cdots\otimes|\psi_n\rangle
$$

Example: a state that is not a product state

$$
|\psi\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{\sqrt{2}}\,|11\rangle
$$

It is a valid quantum state — a unit vector, since the squared amplitudes sum to one:

$$
\left|\tfrac{1}{\sqrt{2}}\right|^2+\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}+\tfrac{1}{2}=1
$$

But it cannot be written as a tensor product $|\phi\rangle\otimes|\psi\rangle$. Any product $(a|0\rangle+b|1\rangle)\otimes(c|0\rangle+d|1\rangle)$ expands to $ac\,|00\rangle+ad\,|01\rangle+bc\,|10\rangle+bd\,|11\rangle$. Matching our state forces $ad=bc=0$ (no $|01\rangle$ or $|10\rangle$ terms) while $ac$ and $bd$ are both nonzero — which is impossible. So the two bits are entangled, not independent.

The Bell basis

The state from the previous example, $|\phi^{+}\rangle$, is one of four Bell states. Each differs only by which basis pairs are combined and by a sign, and together they form the Bell basis — an orthonormal basis of the two-qubit space made entirely of maximally entangled states.

$$
|\phi^{+}\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{\sqrt{2}}\,|11\rangle
$$

$$
|\phi^{-}\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle-\tfrac{1}{\sqrt{2}}\,|11\rangle
$$

$$
|\psi^{+}\rangle=\tfrac{1}{\sqrt{2}}\,|01\rangle+\tfrac{1}{\sqrt{2}}\,|10\rangle
$$

$$
|\psi^{-}\rangle=\tfrac{1}{\sqrt{2}}\,|01\rangle-\tfrac{1}{\sqrt{2}}\,|10\rangle
$$

None of the four can be written as a tensor product $|\phi\rangle\otimes|\psi\rangle$, so each is entangled. Because they are orthonormal, any two-qubit state can be expressed as a combination of these four.

Three-qubit states

Entanglement is not limited to pairs. Two famous three-qubit states show different ways three systems can be correlated — both are unit vectors and neither is a product state.

GHZ state

$$
|\mathrm{GHZ}\rangle=\tfrac{1}{\sqrt{2}}\,|000\rangle+\tfrac{1}{\sqrt{2}}\,|111\rangle
$$

All three qubits are either $|0\rangle$ or all $|1\rangle$. Measuring any one qubit instantly fixes the other two — the three-qubit analogue of the Bell state $|\phi^{+}\rangle$.

W state

$$
|\mathrm{W}\rangle=\tfrac{1}{\sqrt{3}}\,|001\rangle+\tfrac{1}{\sqrt{3}}\,|010\rangle+\tfrac{1}{\sqrt{3}}\,|100\rangle
$$

Exactly one qubit is $|1\rangle$ and its position is in superposition. Each squared amplitude is $\bigl(\tfrac{1}{\sqrt{3}}\bigr)^2=\tfrac{1}{3}$, and the three sum to one.

Measurements

Measuring a compound quantum system works exactly like measuring a single one — *provided every system is measured*. A standard basis measurement of the whole system returns one combined classical outcome, drawn from the squared amplitudes just as in the single-system case.

If $|\psi\rangle$ is a quantum state of a system $(X_1,\ldots,X_n)$ and all systems are measured, then each $n$-tuple

$$
(a_1,\ldots,a_n)\in\Sigma_1\times\cdots\times\Sigma_n
$$

(or string $a_1\cdots a_n$) is obtained with probability equal to the squared absolute value of its amplitude:

$$
\Pr(\text{outcome}=a_1\cdots a_n)=\bigl|\langle a_1\cdots a_n|\psi\rangle\bigr|^2
$$

The inner product $\langle a_1\cdots a_n|\psi\rangle$ simply picks out the amplitude sitting in front of the basis state $|a_1\cdots a_n\rangle$ — the entry of the state vector indexed by that outcome.

Example 1

Measuring both qubits of the Bell state $|\phi^{+}\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{\sqrt{2}}\,|11\rangle$ gives $00$ or $11$ only — the two qubits always agree:

$$
\Pr(00)=\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}\qquad\Pr(11)=\left|\tfrac{1}{\sqrt{2}}\right|^2=\tfrac{1}{2}\qquad\Pr(01)=\Pr(10)=0
$$

Example 2

Subsystems need not be qubits and amplitudes may be complex. For the pair $(X,Y)$ in the state $\tfrac{3}{5}\,|0\rangle|{\textcolor{#dc2626}{\heartsuit}}\rangle-\tfrac{4i}{5}\,|1\rangle|{\spadesuit}\rangle$:

$$
\Pr(0,\textcolor{#dc2626}{\heartsuit})=\left|\tfrac{3}{5}\right|^2=\tfrac{9}{25}\qquad\Pr(1,\spadesuit)=\left|-\tfrac{4i}{5}\right|^2=\tfrac{16}{25}
$$

The two add to $1$, and the $i$ drops out under $|\cdot|^2$ — only amplitude magnitudes matter.

Measuring some systems

What if two systems $(X,Y)$ share a quantum state but we measure only $X$ and leave $Y$ alone? Write the joint state in the usual form:

$$
|\psi\rangle=\sum_{(a,b)\in\Sigma\times\Gamma}\alpha_{ab}\,|ab\rangle
$$

If both were measured, outcome $(a,b)$ would appear with probability $|\langle ab|\psi\rangle|^2=|\alpha_{ab}|^2$. Measuring only $X$ must give the same probability for $X=a$ as summing over every $Y$ outcome:

$$
\Pr(X=a)=\sum_{b\in\Gamma}|\langle ab|\psi\rangle|^2=\sum_{b\in\Gamma}|\alpha_{ab}|^2
$$

Just as in the probabilistic setting, the state of $Y$ changes as a result. The branch with $X=a$ survives and must be renormalized back to a unit vector — but because these are amplitudes, not probabilities, we divide by the square root of $\Pr(X=a)$:

$$
|\psi_Y\rangle=\frac{\sum_{b\in\Gamma}\alpha_{ab}\,|b\rangle}{\sqrt{\Pr(X=a)}}=\frac{\sum_{b\in\Gamma}\alpha_{ab}\,|b\rangle}{\sqrt{\sum_{c\in\Gamma}|\alpha_{ac}|^2}}
$$

That square-root normalization is the one real difference from the classical case — otherwise partial measurement collapses a quantum compound system exactly the way it collapses a probabilistic one.

Example 3

Suppose $(X,Y)$ is in the state below and we measure only $X$:

$$
|\psi\rangle=\tfrac{1}{\sqrt{2}}\,|00\rangle+\tfrac{1}{2}\,|01\rangle+\tfrac{i}{2\sqrt{2}}\,|10\rangle-\tfrac{1}{2\sqrt{2}}\,|11\rangle
$$

We begin by writing it grouped by $X$, factoring each term into $|a\rangle\otimes|b\rangle$:

$$
|\psi\rangle=|0\rangle\otimes\Bigl(\tfrac{1}{\sqrt{2}}\,|0\rangle+\tfrac{1}{2}\,|1\rangle\Bigr)+|1\rangle\otimes\Bigl(\tfrac{i}{2\sqrt{2}}\,|0\rangle-\tfrac{1}{2\sqrt{2}}\,|1\rangle\Bigr)
$$

The probability of each outcome is the squared norm of its branch. Recall the norm $\|\cdot\|$ is a vector's length, so the squared norm is just the sum of the squared amplitude magnitudes:

$$
\Bigl\|\sum_b \gamma_b\,|b\rangle\Bigr\|^2=\sum_b |\gamma_b|^2
$$

Then keep the surviving branch and divide by $\sqrt{\Pr(X=a)}$ so it is a unit vector again.

Outcome $X=0$

$$
\Pr(X=0)=\Bigl\|\tfrac{1}{\sqrt{2}}|0\rangle+\tfrac{1}{2}|1\rangle\Bigr\|^2=\tfrac{1}{2}+\tfrac{1}{4}=\tfrac{3}{4}
$$

Divide the $X=0$ branch by $\sqrt{\tfrac{3}{4}}=\tfrac{\sqrt{3}}{2}$:

$$
|0\rangle\otimes\frac{\tfrac{1}{\sqrt{2}}|0\rangle+\tfrac{1}{2}|1\rangle}{\sqrt{3/4}}=|0\rangle\otimes\Bigl(\sqrt{\tfrac{2}{3}}\,|0\rangle+\tfrac{1}{\sqrt{3}}\,|1\rangle\Bigr)
$$

Outcome $X=1$

$$
\Pr(X=1)=\Bigl\|\tfrac{i}{2\sqrt{2}}|0\rangle-\tfrac{1}{2\sqrt{2}}|1\rangle\Bigr\|^2=\tfrac{1}{8}+\tfrac{1}{8}=\tfrac{1}{4}
$$

Divide the $X=1$ branch by $\sqrt{\tfrac{1}{4}}=\tfrac{1}{2}$:

$$
|1\rangle\otimes\frac{\tfrac{i}{2\sqrt{2}}|0\rangle-\tfrac{1}{2\sqrt{2}}|1\rangle}{\sqrt{1/4}}=|1\rangle\otimes\Bigl(\tfrac{i}{\sqrt{2}}\,|0\rangle-\tfrac{1}{\sqrt{2}}\,|1\rangle\Bigr)
$$

Each squared amplitude in the post-measurement state now sums back to $1$: for $X=0$, $\tfrac{2}{3}+\tfrac{1}{3}=1$; for $X=1$, $\tfrac{1}{2}+\tfrac{1}{2}=1$ — measuring $X$ has left $Y$ in a valid quantum state.

Measuring some: Y

The same works the other way round. Take the same state but measure only $Y$, grouping each term by $Y$ instead:

$$
|\psi\rangle=\Bigl(\tfrac{1}{\sqrt{2}}\,|0\rangle+\tfrac{i}{2\sqrt{2}}\,|1\rangle\Bigr)\otimes|0\rangle+\Bigl(\tfrac{1}{2}\,|0\rangle-\tfrac{1}{2\sqrt{2}}\,|1\rangle\Bigr)\otimes|1\rangle
$$

Now each branch carries the amplitudes of $X$. The same squared-norm rule gives the outcome probabilities, and the surviving branch is divided by $\sqrt{\Pr(Y=b)}$:

Outcome $Y=0$

$$
\Pr(Y=0)=\Bigl\|\tfrac{1}{\sqrt{2}}|0\rangle+\tfrac{i}{2\sqrt{2}}|1\rangle\Bigr\|^2=\tfrac{1}{2}+\tfrac{1}{8}=\tfrac{5}{8}
$$

Divide the $Y=0$ branch by $\sqrt{\tfrac{5}{8}}$:

$$
\frac{\tfrac{1}{\sqrt{2}}|0\rangle+\tfrac{i}{2\sqrt{2}}|1\rangle}{\sqrt{5/8}}\otimes|0\rangle=\Bigl(\sqrt{\tfrac{4}{5}}\,|0\rangle+\tfrac{i}{\sqrt{5}}\,|1\rangle\Bigr)\otimes|0\rangle
$$

Outcome $Y=1$

$$
\Pr(Y=1)=\Bigl\|\tfrac{1}{2}|0\rangle-\tfrac{1}{2\sqrt{2}}|1\rangle\Bigr\|^2=\tfrac{1}{4}+\tfrac{1}{8}=\tfrac{3}{8}
$$

Divide the $Y=1$ branch by $\sqrt{\tfrac{3}{8}}$:

$$
\frac{\tfrac{1}{2}|0\rangle-\tfrac{1}{2\sqrt{2}}|1\rangle}{\sqrt{3/8}}\otimes|1\rangle=\Bigl(\sqrt{\tfrac{2}{3}}\,|0\rangle-\tfrac{1}{\sqrt{3}}\,|1\rangle\Bigr)\otimes|1\rangle
$$

As before the two probabilities add to $\tfrac{5}{8}+\tfrac{3}{8}=1$, and each collapsed $X$ state is a unit vector again.

Example 4 — three qubits

Nothing changes with more systems. Take the W state of $(X,Y,Z)$ and measure only the first qubit $X$, grouping the rest as $|x\rangle\otimes|yz\rangle$:

$$
|\mathrm{W}\rangle=\tfrac{1}{\sqrt{3}}\bigl(|001\rangle+|010\rangle+|100\rangle\bigr)=|0\rangle\otimes\tfrac{1}{\sqrt{3}}\bigl(|01\rangle+|10\rangle\bigr)+|1\rangle\otimes\tfrac{1}{\sqrt{3}}|00\rangle
$$

Outcome $X=0$

$$
\Pr(X=0)=\Bigl\|\tfrac{1}{\sqrt{3}}|01\rangle+\tfrac{1}{\sqrt{3}}|10\rangle\Bigr\|^2=\tfrac{1}{3}+\tfrac{1}{3}=\tfrac{2}{3}
$$

Divide the branch by $\sqrt{\tfrac{2}{3}}$ — and $(Y,Z)$ is left in the Bell state $|\psi^{+}\rangle$:

$$
\frac{\tfrac{1}{\sqrt{3}}|01\rangle+\tfrac{1}{\sqrt{3}}|10\rangle}{\sqrt{2/3}}=\tfrac{1}{\sqrt{2}}\bigl(|01\rangle+|10\rangle\bigr)
$$

Outcome $X=1$

$$
\Pr(X=1)=\Bigl\|\tfrac{1}{\sqrt{3}}|00\rangle\Bigr\|^2=\tfrac{1}{3}
$$

Divide the branch by $\sqrt{\tfrac{1}{3}}$ — and $(Y,Z)$ is left in the definite state $|00\rangle$:

$$
\frac{\tfrac{1}{\sqrt{3}}|00\rangle}{\sqrt{1/3}}=|00\rangle
$$

This shows the robustness of the W state mentioned earlier: with probability $\tfrac{2}{3}$ losing one qubit still leaves the other two entangled — whereas a GHZ state would collapse to a fully definite $|00\rangle$ or $|11\rangle$ on either outcome.

Key idea: no matter how many systems there are, we can always regroup the state into the part being measured and the part left unmeasured — one $\otimes$ branch per outcome of the measured part. Each outcome's probability is the squared norm of its branch, and the unmeasured part is left in that branch, renormalized by $\sqrt{\Pr(\text{outcome})}$. The split into “measured” vs. “unmeasured” is all that matters — the same recipe handles any subset of any number of systems.

Unitary operations

Just like for a single system, a quantum operation on a compound system is represented by a unitary matrix — but now its rows and columns are indexed by the Cartesian product of the individual classical state sets, the same index set as the compound state vector it acts on.

For instance, if $X$ has state set $\{1,2,3\}$ and $Y$ has state set $\{0,1\}$, the compound system has $3\times 2=6$ classical states, so an operation on $(X,Y)$ is a $6\times 6$ unitary matrix:

$$
U=\begin{pmatrix}
\tfrac{1}{2} & \tfrac{1}{2} & \tfrac{1}{2} & 0 & 0 & \tfrac{1}{2}\\[2pt]
\tfrac{1}{2} & \tfrac{i}{2} & -\tfrac{1}{2} & 0 & 0 & -\tfrac{i}{2}\\[2pt]
\tfrac{1}{2} & -\tfrac{1}{2} & \tfrac{1}{2} & 0 & 0 & -\tfrac{1}{2}\\[2pt]
0 & 0 & 0 & \tfrac{1}{\sqrt{2}} & \tfrac{1}{\sqrt{2}} & 0\\[2pt]
\tfrac{1}{2} & -\tfrac{i}{2} & -\tfrac{1}{2} & 0 & 0 & \tfrac{i}{2}\\[2pt]
0 & 0 & 0 & -\tfrac{1}{\sqrt{2}} & \tfrac{1}{\sqrt{2}} & 0
\end{pmatrix}
$$

Independent operations: tensor product

The general matrix above can entangle the systems it acts on. But often each system is acted on independently — a separate gate on each, with no interaction between them. Just as the tensor product combined separate states into one compound vector, it combines these separate operations into one compound unitary. If $X_1,\ldots,X_n$ carry the operations $U_1,\ldots,U_n$, the combined action on $(X_1,\ldots,X_n)$ is their tensor product:

$$
U_1\otimes\cdots\otimes U_n
$$

Read it slot by slot: the matrix in each position is the operation that system experiences on its own. A tensor product of unitaries is again unitary, so independent operations always assemble into a single valid operation on the whole compound system.

The most common case is acting on just one part and leaving the rest alone — and doing nothing is itself a unitary, the identity matrix $I$. For instance, applying the Hadamard gate $H$ to the first qubit while leaving the second untouched is $H\otimes I$, which passes the second qubit through unchanged:

$$
H\otimes I=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}\otimes\begin{pmatrix}1&0\\0&1\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&0&\tfrac{1}{\sqrt{2}}&0\\[4pt]0&\tfrac{1}{\sqrt{2}}&0&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&0&-\tfrac{1}{\sqrt{2}}&0\\[4pt]0&\tfrac{1}{\sqrt{2}}&0&-\tfrac{1}{\sqrt{2}}\end{pmatrix}
$$

Swapping the order applies $H$ to the second qubit instead — now $H$ appears as two identical blocks along the diagonal:

$$
I\otimes H=\begin{pmatrix}1&0\\0&1\end{pmatrix}\otimes\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\[4pt]\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\\[4pt]0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\[4pt]0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}
$$

Operations that aren't tensor products: SWAP

A tensor product describes systems acted on independently, but not every unitary on a compound system factors that way. The standard example is the SWAP operation, which exchanges the contents of two systems $X$ and $Y$ sharing the same classical state set $\Sigma$:

$$
\mathrm{SWAP}\bigl(|\varphi\rangle\otimes|\psi\rangle\bigr)=|\psi\rangle\otimes|\varphi\rangle
$$

To build that operation, consider one possible input basis state $|a,b\rangle$. The bra $\langle a,b|$ recognizes that input, while the ket $|b,a\rangle$ supplies its swapped output. Their outer product therefore handles one input, and summing over every possible pair handles them all:

$$
\mathrm{SWAP}=\sum_{a,b\in\Sigma}|b,a\rangle\langle a,b|=\sum_{a,b\in\Sigma}|b\rangle\langle a|\otimes|a\rangle\langle b|
$$

Each tensor-product term handles one specific pair: $|b\rangle\langle a|$ changes the first system from $a$ to $b$, while $|a\rangle\langle b|$ changes the second from $b$ to $a$. Together they map $|a,b\rangle$ to $|b,a\rangle$. So every such tensor-product term swaps one specific pair.

For example, let's consider a simple system of two qubits. Each qubit has basis-state set $\Sigma=\{0,1\}$, so both $a$ and $b$ can be either 0 or 1. Writing out all four possible pairs in the sum gives:

$$
\mathrm{SWAP}=\underbrace{|0\rangle\langle0|\otimes|0\rangle\langle0|}_{(a,b)=(0,0)}+\underbrace{|1\rangle\langle0|\otimes|0\rangle\langle1|}_{(a,b)=(0,1)}+\underbrace{|0\rangle\langle1|\otimes|1\rangle\langle0|}_{(a,b)=(1,0)}+\underbrace{|1\rangle\langle1|\otimes|1\rangle\langle1|}_{(a,b)=(1,1)}
$$

$$
\begin{aligned}&=\left[\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}\right]\otimes\left[\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}\right]+\left[\begin{pmatrix}0\\1\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}\right]\otimes\left[\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix}\right]\\[6pt]&\quad+\left[\begin{pmatrix}1\\0\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix}\right]\otimes\left[\begin{pmatrix}0\\1\end{pmatrix}\begin{pmatrix}1&0\end{pmatrix}\right]+\left[\begin{pmatrix}0\\1\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix}\right]\otimes\left[\begin{pmatrix}0\\1\end{pmatrix}\begin{pmatrix}0&1\end{pmatrix}\right]\end{aligned}
$$

$$
=\begin{pmatrix}1&0\\0&0\end{pmatrix}\otimes\begin{pmatrix}1&0\\0&0\end{pmatrix}+\begin{pmatrix}0&0\\1&0\end{pmatrix}\otimes\begin{pmatrix}0&1\\0&0\end{pmatrix}+\begin{pmatrix}0&1\\0&0\end{pmatrix}\otimes\begin{pmatrix}0&0\\1&0\end{pmatrix}+\begin{pmatrix}0&0\\0&1\end{pmatrix}\otimes\begin{pmatrix}0&0\\0&1\end{pmatrix}
$$

$$
=\begin{pmatrix}1&0&0&0\\0&0&0&0\\0&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&0\\0&1&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&1&0\\0&0&0&0\\0&0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0&0\\0&0&0&0\\0&0&0&0\\0&0&0&1\end{pmatrix}
$$

$$
=\begin{pmatrix}1&0&0&0\\0&0&1&0\\0&1&0&0\\0&0&0&1\end{pmatrix}
$$

To apply SWAP, multiply its matrix by the compound state vector. For example, the basis state $|01\rangle$ is the second standard basis vector, so the multiplication moves its 1 into the position for $|10\rangle$:

$$
\mathrm{SWAP}|01\rangle=\begin{pmatrix}1&0&0&0\\0&0&1&0\\0&1&0&0\\0&0&0&1\end{pmatrix}\begin{pmatrix}0\\1\\0\\0\end{pmatrix}=\begin{pmatrix}0\\0\\1\\0\end{pmatrix}=|10\rangle
$$

Controlled operations

Suppose that $X$ is a qubit and $Y$ is an arbitrary quantum system. A controlled-$U$ operation uses $X$ as a switch for a unitary operation $U$ on $Y$. When the control is $|0\rangle$, the target system is left unchanged; when the control is $|1\rangle$, the operation $U$ is applied to it.

The projectors $|0\rangle\langle0|$ and $|1\rangle\langle1|$ select those two branches, giving the following operation on the pair $(X,Y)$:

$$
\operatorname{controlled}\text{-}U=|0\rangle\langle0|\otimes I_Y+|1\rangle\langle1|\otimes U=\begin{pmatrix}I_Y&0\\0&U\end{pmatrix}
$$

The matrix $\begin{pmatrix}I_Y&0\\0&U\end{pmatrix}$ is a block matrix, not an ordinary 2 × 2 matrix. If $Y$ has $d$ basis states, then $I_Y$, $U$, and each zero are all $d\times d$ blocks. The complete controlled-$U$ matrix is therefore $2d\times2d$. In particular, if $Y$ is a qubit, the blocks are $2\times2$ and the full matrix is $4\times4$.

Example: controlled-NOT

Let the target $Y$ also be a qubit and choose $U=\sigma_x$. On one qubit, $\sigma_x$flips the two basis states:

$$
\sigma_x=\begin{pmatrix}0&1\\1&0\end{pmatrix},\qquad \sigma_x|0\rangle=|1\rangle,\qquad \sigma_x|1\rangle=|0\rangle
$$

Now substitute $U=\sigma_x$ and $I_Y=I_2$ into the controlled-$U$ formula. The following equalities turn the general controlled operation into the controlled-NOT, or CNOT, gate:

$$
\mathrm{CNOT}=|0\rangle\langle0|\otimes I_2+|1\rangle\langle1|\otimes\sigma_x
$$

$$
=\begin{pmatrix}1&0\\0&0\end{pmatrix}\otimes\begin{pmatrix}1&0\\0&1\end{pmatrix}+\begin{pmatrix}0&0\\0&1\end{pmatrix}\otimes\begin{pmatrix}0&1\\1&0\end{pmatrix}
$$

$$
=\begin{pmatrix}I_2&0\\0&\sigma_x\end{pmatrix}
$$

$$
=\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}
$$

In this example, the first digit is the control and the second is the target. Therefore $|10\rangle$ means that the control is $|1\rangle$ and the target is $|0\rangle$. A control value of 1 turns the operation on. The control stays at 1, while $\sigma_x$ flips the target from 0 to 1:

$$
\mathrm{CNOT}|10\rangle=|1\rangle\otimes\sigma_x|0\rangle=|1\rangle\otimes|1\rangle=|11\rangle
$$

In the standard basis order $|00\rangle,|01\rangle,|10\rangle,|11\rangle$, the state $|10\rangle$ is the third basis vector. The full matrix multiplication shows the same change from the third basis vector to the fourth:

$$
\mathrm{CNOT}|10\rangle=\begin{pmatrix}1&0&0&0\\0&1&0&0\\0&0&0&1\\0&0&1&0\end{pmatrix}\underbrace{\begin{pmatrix}0\\0\\1\\0\end{pmatrix}}_{|10\rangle}=\underbrace{\begin{pmatrix}0\\0\\0\\1\end{pmatrix}}_{|11\rangle}
$$

The control can be any qubit

The first qubit is the control above only because we chose it that way. A controlled operation can use any qubit as its control. If the second qubit controls an operation $U$ on the first, the projectors move to the second tensor factor:

$$
C_U^{(2\to1)}=I_2\otimes|0\rangle\langle0|+U\otimes|1\rangle\langle1|
$$

With that ordering, the second digit of $|a,b\rangle$ is the control, and the first digit is the target.

Example: Fredkin operation

A controlled-SWAP on three qubits swaps the last two qubits only when the first qubit is 1. This is called the Fredkin operation, or Fredkin gate:

$$
\mathrm{Fredkin}=|0\rangle\langle0|\otimes I_4+|1\rangle\langle1|\otimes\mathrm{SWAP}=\begin{pmatrix}1&0&0&0&0&0&0&0\\0&1&0&0&0&0&0&0\\0&0&1&0&0&0&0&0\\0&0&0&1&0&0&0&0\\0&0&0&0&1&0&0&0\\0&0&0&0&0&0&1&0\\0&0&0&0&0&1&0&0\\0&0&0&0&0&0&0&1\end{pmatrix}
$$

Example: Toffoli operation

A controlled-controlled-NOT on three qubits flips the third qubit only when both the first and second qubits are 1. This is called the Toffoli operation, or Toffoli gate:

$$
\mathrm{Toffoli}=|0\rangle\langle0|\otimes I_2\otimes I_2+|1\rangle\langle1|\otimes\left(|0\rangle\langle0|\otimes I_2+|1\rangle\langle1|\otimes\sigma_x\right)=\begin{pmatrix}1&0&0&0&0&0&0&0\\0&1&0&0&0&0&0&0\\0&0&1&0&0&0&0&0\\0&0&0&1&0&0&0&0\\0&0&0&0&1&0&0&0\\0&0&0&0&0&1&0&0\\0&0&0&0&0&0&0&1\\0&0&0&0&0&0&1&0\end{pmatrix}
$$

## Quantum circuits

Circuits as models of computation

A circuit is a graphical model of computation. It describes how information moves through a computation and which operations are applied along the way.

Wires represent the paths along which values travel from one part of the computation to another. Gates receive one or more input values, perform an operation on them, and pass the resulting values along their outgoing wires.

Interactive Boolean circuit editor

Click X or Y to toggle it. Drag nodes to move them, drag square wire handles to reroute, and drag circular output ports onto input ports to connect. Drag an occupied input port to reconnect its wire. Right-click anywhere on the board for context-specific actions.

The quantum circuit model

In the quantum circuit model, wires represent qubits and gates represent both unitary operations and measurements.

The terminology comes from physical electrical circuits, where wires carry current and components act on it. Here, wires and gates describe information flow, not electricity. The circuits considered here are acyclic: information flows in one direction, conventionally from left to right. A wire never loops back to an earlier gate, so the gates define an unambiguous order of computation.

A one-qubit circuit

From left to right, qubit $X$ passes through $H,\ S,\ H,\ T$. Operators act on a ket from the left, so the rightmost matrix in a product acts first. The output is therefore $|\psi_{\mathrm{out}}\rangle=THSH|\psi_{\mathrm{in}}\rangle$: the combined operation is written $THSH$, the reverse of the visual $HSHT$ order. Hover over, tap or focus a gate to see its full operation.

The three gate matrices are:

$$
H=\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix},\qquad S=\begin{pmatrix}1&0\\0&i\end{pmatrix},\qquad T=\begin{pmatrix}1&0\\0&e^{i\pi/4}\end{pmatrix}=\begin{pmatrix}1&0\\0&\tfrac{1+i}{\sqrt{2}}\end{pmatrix}
$$

Multiply one step at a time, starting at the right:

$$
SH=\begin{pmatrix}1&0\\0&i\end{pmatrix}\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix}=\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\i&-i\end{pmatrix}
$$

$$
HSH=H(SH)=\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix}\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\i&-i\end{pmatrix}=\frac{1}{2}\begin{pmatrix}1+i&1-i\\1-i&1+i\end{pmatrix}
$$

$$
THSH=T(HSH)=\begin{pmatrix}1&0\\0&\tfrac{1+i}{\sqrt{2}}\end{pmatrix}\frac{1}{2}\begin{pmatrix}1+i&1-i\\1-i&1+i\end{pmatrix}=\begin{pmatrix}\tfrac{1+i}{2}&\tfrac{1-i}{2}\\\tfrac{1}{\sqrt{2}}&\tfrac{i}{\sqrt{2}}\end{pmatrix}
$$

A controlled-NOT circuit

The Hadamard gate first acts on the upper qubit $Y$. The filled dot then makes $Y$ the CNOT control, while the circled plus marks $X$ as its target. If $Y=0$, the target is unchanged; if $Y=1$, the target is flipped.

For the matrix calculation, use the register order $(X,Y)$: $|00\rangle_{XY},|01\rangle_{XY},|10\rangle_{XY},|11\rangle_{XY}$. The Hadamard acts on $Y$, the second tensor factor, so its two-qubit matrix is $I_2\otimes H$:

$$
I_2\otimes H=\begin{pmatrix}1&0\\0&1\end{pmatrix}\otimes\frac{1}{\sqrt{2}}\begin{pmatrix}1&1\\1&-1\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\\0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}
$$

With $Y$ as the control, CNOT exchanges $|01\rangle$ and $|11\rangle$ while leaving the other two basis states unchanged:

$$
\mathrm{CNOT}_{Y\to X}=\begin{pmatrix}1&0&0&0\\0&0&0&1\\0&0&1&0\\0&1&0&0\end{pmatrix}
$$

CNOT acts after the Hadamard, so it appears on the left in the product:

$$
U=\mathrm{CNOT}_{Y\to X}(I_2\otimes H)=\begin{pmatrix}1&0&0&0\\0&0&0&1\\0&0&1&0\\0&1&0&0\end{pmatrix}\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\\0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\\0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\end{pmatrix}
$$

Applying that matrix to $|00\rangle_{XY}$ gives the Bell state:

$$
U|00\rangle_{XY}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}&0&0\\0&0&\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}\\0&0&\tfrac{1}{\sqrt{2}}&\tfrac{1}{\sqrt{2}}\\\tfrac{1}{\sqrt{2}}&-\tfrac{1}{\sqrt{2}}&0&0\end{pmatrix}\begin{pmatrix}1\\0\\0\\0\end{pmatrix}=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\0\\0\\\tfrac{1}{\sqrt{2}}\end{pmatrix}=\frac{|00\rangle+|11\rangle}{\sqrt{2}}
$$

For all four computational-basis inputs, the circuit produces the four Bell states below. The final minus sign is a global phase and does not affect measurement probabilities.

$$
\begin{aligned}U|00\rangle_{XY}&=\frac{|00\rangle+|11\rangle}{\sqrt{2}}=|\phi^{+}\rangle,\\[6pt]U|01\rangle_{XY}&=\frac{|00\rangle-|11\rangle}{\sqrt{2}}=|\phi^{-}\rangle,\\[6pt]U|10\rangle_{XY}&=\frac{|01\rangle+|10\rangle}{\sqrt{2}}=|\psi^{+}\rangle,\\[6pt]U|11\rangle_{XY}&=\frac{-|01\rangle+|10\rangle}{\sqrt{2}}=-|\psi^{-}\rangle.\end{aligned}
$$

Follow the state through the circuit

To find the output for one input, we do not need to multiply the full gate matrices. Instead, follow the state from left to right and update it at each gate; the slices below are labeled $|\pi_0\rangle,|\pi_1\rangle,|\pi_2\rangle$.

$$
|0\rangle
$$

$$
|\pi_0\rangle
$$

$$
|\pi_1\rangle
$$

$$
|\pi_2\rangle
$$

$$
\begin{aligned}|\pi_0\rangle&=|0\rangle_X|0\rangle_Y=|00\rangle_{XY},\\[6pt]|\pi_1\rangle&=|0\rangle_X|+\rangle_Y=\frac{|00\rangle+|01\rangle}{\sqrt{2}},\\[6pt]|\pi_2\rangle&=\mathrm{CNOT}_{Y\to X}|\pi_1\rangle=\frac{|00\rangle+|11\rangle}{\sqrt{2}}=|\phi^+\rangle.\end{aligned}
$$

Classical states in a quantum circuit

A measurement connects quantum and classical information: it collapses a qubit to $|0\rangle$ or $|1\rangle$ and writes the corresponding bit onto a double classical wire. Here the measurements store Y in B and X in A. Because the qubits are in $|\phi^+\rangle$, the classical result is $(A,B)=(0,0)$ or $(1,1)$, each with probability one half; those bits can then control later operations or be read as output.

YXBA

Quantum-circuits symbols

Single-qubit gates

Controlled-NOT (CNOT)

SWAP

Toffoli (CCNOT)

Fredkin (controlled-SWAP)

Arbitrary unitary U

Controlled-U

## Quantum states and measurements

Inner products

Suppose that we have two kets $\lvert\psi\rangle$ and $\lvert\phi\rangle$ with complex amplitudes:

$$
\lvert\psi\rangle=\begin{pmatrix}\alpha_1\\\vdots\\\alpha_n\end{pmatrix}\qquad\lvert\phi\rangle=\begin{pmatrix}\beta_1\\\vdots\\\beta_n\end{pmatrix}
$$

To calculate the inner product of $\lvert\psi\rangle$ and $\lvert\phi\rangle$, first turn $\lvert\psi\rangle$ into the bra $\langle\psi\rvert=\lvert\psi\rangle^\dagger$ — its conjugate transpose, whose entries are the complex conjugates $\overline{\alpha_i}$— then multiply it by $\lvert\phi\rangle$ to get a single number:

$$
\langle\psi\vert\phi\rangle=\begin{pmatrix}\overline{\alpha_1}&\cdots&\overline{\alpha_n}\end{pmatrix}\begin{pmatrix}\beta_1\\\vdots\\\beta_n\end{pmatrix}=\overline{\alpha_1}\beta_1+\cdots+\overline{\alpha_n}\beta_n
$$

For example, take the two qubit states:

$$
\lvert\psi\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{i}{\sqrt{2}}\end{pmatrix}\qquad\lvert\phi\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[4pt]\tfrac{1}{\sqrt{2}}\end{pmatrix}
$$

Conjugating the amplitudes of $\lvert\psi\rangle$ flips the $i$ to $-i$:

$$
\langle\psi\vert\phi\rangle=\begin{pmatrix}\tfrac{1}{\sqrt{2}}&-\tfrac{i}{\sqrt{2}}\end{pmatrix}\begin{pmatrix}\tfrac{1}{\sqrt{2}}\\[2pt]\tfrac{1}{\sqrt{2}}\end{pmatrix}=\tfrac{1}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}-\tfrac{i}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}=\frac{1-i}{2}
$$

Alternatively, suppose the same vectors are written as sums over a basis $\Sigma$:

$$
\lvert\psi\rangle=\sum_{a\in\Sigma}\alpha_a\lvert a\rangle\qquad\lvert\phi\rangle=\sum_{b\in\Sigma}\beta_b\lvert b\rangle
$$

This gives the same formula as before, just in a different format. Expanding the product term by term, each $\langle a\vert b\rangle$ is $1$ when $a=b$ and $0$ otherwise, so only the matching terms survive:

$$
\begin{aligned}\langle\psi\vert\phi\rangle&=\left(\sum_{a\in\Sigma}\overline{\alpha_a}\langle a\rvert\right)\left(\sum_{b\in\Sigma}\beta_b\lvert b\rangle\right)\\[2pt]&=\sum_{a\in\Sigma}\sum_{b\in\Sigma}\overline{\alpha_a}\beta_b\langle a\vert b\rangle\\[2pt]&=\sum_{a\in\Sigma}\overline{\alpha_a}\beta_a\end{aligned}
$$

Our example states in this form, over the basis $\Sigma=\{0,1\}$:

$$
\lvert\psi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{i}{\sqrt{2}}\lvert1\rangle\qquad\lvert\phi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{\sqrt{2}}\lvert1\rangle
$$

Substituting our values into the formula, we get (the $\langle 0\vert 1\rangle$ and $\langle 1\vert 0\rangle$ terms are $0$, so they cancel):

$$
\begin{aligned}\langle\psi\vert\phi\rangle&=\left(\tfrac{1}{\sqrt{2}}\langle 0\rvert-\tfrac{i}{\sqrt{2}}\langle 1\rvert\right)\left(\tfrac{1}{\sqrt{2}}\lvert 0\rangle+\tfrac{1}{\sqrt{2}}\lvert 1\rangle\right)\\[4pt]&=\tfrac{1}{2}\langle 0\vert 0\rangle+\cancel{\tfrac{1}{2}\langle 0\vert 1\rangle}-\cancel{\tfrac{i}{2}\langle 1\vert 0\rangle}-\tfrac{i}{2}\langle 1\vert 1\rangle\\[4pt]&=\tfrac{1}{2}-\tfrac{i}{2}=\frac{1-i}{2}\end{aligned}
$$

Inner product as an angle

When the amplitudes are all real, the conjugation does nothing ($\overline{\alpha}=\alpha$), and the inner product has a clean geometric meaning: for two unit vectors it equals the cosine of the angle between them.

$$
\langle\psi\vert\phi\rangle=\alpha_0\beta_0+\alpha_1\beta_1=\cos\theta
$$

|ψ⟩ = 1/√2|0⟩ − 1/√2|1⟩

|ϕ⟩ = 1/2|0⟩ + √3/2|1⟩

θ = 105°

⟨ψ|ϕ⟩ = cos θ = -0.26

Drag either tip around the circle to see how cos changes value

This geometric interpretation only works when all amplitudes are real. With complex amplitudes, the inner product becomes a complex number, so it no longer represents the cosine of an ordinary angle. The quantity that remains physically meaningful is its magnitude, $\lvert\langle\psi\vert\phi\rangle\rvert$, which determines measurement probabilities.

Inner product as an angle in multidimensional space

The idea of the inner product as an angle for real amplitudes generalizes to any number of dimensions: for real unit vectors the inner product is still the cosine of the angle between them — it just lives on a higher-dimensional sphere. Below is the three-dimensional case (a qutrit with real amplitudes, basis $\lvert0\rangle,\lvert1\rangle,\lvert2\rangle$):

$$
\langle\psi\vert\phi\rangle=\alpha_0\beta_0+\alpha_1\beta_1+\alpha_2\beta_2=\cos\theta
$$

|ψ⟩ = 1/√2|0⟩ + 0|1⟩ + 1/√2|2⟩

|ϕ⟩ = 0|0⟩ + 1/√2|1⟩ + 1/√2|2⟩

θ = 60°

⟨ψ|ϕ⟩ = cos θ = 0.50

Drag a tip to move it on the sphere; drag the background to rotate the view.

Properties of the inner product

Relationship to the Euclidean norm

Take the inner product of a vector $\lvert\psi\rangle=\sum_{a\in\Sigma}\alpha_a\lvert a\rangle$ with itself. Each conjugate pair collapses to a squared magnitude, $\overline{\alpha_a}\alpha_a=\lvert\alpha_a\rvert^2$, so the result is the sum of squared amplitudes — exactly the squared Euclidean length of the vector:

$$
\langle\psi\vert\psi\rangle=\sum_{a\in\Sigma}\overline{\alpha_a}\alpha_a=\sum_{a\in\Sigma}\lvert\alpha_a\rvert^2=\bigl\lVert\,\lvert\psi\rangle\,\bigr\rVert^2
$$

When a state is normalized (that is, it represents a valid quantum state), its inner product with itself equals $1$. Equivalently, its Euclidean norm equals $1$. For example, for the vector $\lvert\psi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{i}{\sqrt{2}}\lvert1\rangle$:

$$
\langle\psi\vert\psi\rangle=\left\lvert\tfrac{1}{\sqrt{2}}\right\rvert^2+\left\lvert\tfrac{i}{\sqrt{2}}\right\rvert^2=\tfrac{1}{2}+\tfrac{1}{2}=1
$$

More generally, the Euclidean norm of any vector is the square root of its inner product with itself:

$$
\bigl\lVert\,\lvert\psi\rangle\,\bigr\rVert=\sqrt{\langle\psi\vert\psi\rangle}
$$

Conjugate symmetry

Swapping the order of the two vectors conjugates a different set of amplitudes. For $\lvert\psi\rangle=\sum_{a\in\Sigma}\alpha_a\lvert a\rangle$ and $\lvert\phi\rangle=\sum_{a\in\Sigma}\beta_a\lvert a\rangle$, the two orderings are

$$
\langle\psi\vert\phi\rangle=\sum_{a\in\Sigma}\overline{\alpha_a}\beta_a\qquad\text{and}\qquad\langle\phi\vert\psi\rangle=\sum_{a\in\Sigma}\overline{\beta_a}\alpha_a,
$$

which differ only in which factor carries the bar. In fact the two are complex conjugates of each other:

$$
\overline{\langle\psi\vert\phi\rangle}=\langle\phi\vert\psi\rangle.
$$

Why. Conjugate $\langle\psi\vert\phi\rangle$, pulling the bar inside the sum and onto each factor:

$$
\overline{\langle\psi\vert\phi\rangle}=\overline{\sum_{a\in\Sigma}\overline{\alpha_a}\beta_a}=\sum_{a\in\Sigma}\overline{\overline{\alpha_a}\beta_a}=\sum_{a\in\Sigma}\alpha_a\overline{\beta_a}.
$$

Because complex multiplication is commutative, $\alpha_a\overline{\beta_a}=\overline{\beta_a}\alpha_a$, the final summand is exactly that of $\langle\phi\vert\psi\rangle$, which proves the identity.

Example. Take the two states

$$
\lvert\psi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{i}{\sqrt{2}}\lvert1\rangle\qquad\lvert\phi\rangle=\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{\sqrt{2}}\lvert1\rangle
$$

Computing both orderings yields a conjugate pair:

$$
\langle\psi\vert\phi\rangle=\tfrac{1}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}-\tfrac{i}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}=\frac{1-i}{2}
$$

$$
\langle\phi\vert\psi\rangle=\tfrac{1}{\sqrt{2}}\cdot\tfrac{1}{\sqrt{2}}+\tfrac{1}{\sqrt{2}}\cdot\tfrac{i}{\sqrt{2}}=\frac{1+i}{2}=\overline{\left(\tfrac{1-i}{2}\right)}.
$$

Linearity in the second argument

Suppose that $\lvert\psi\rangle$, $\lvert\phi_1\rangle$, and $\lvert\phi_2\rangle$ are vectors and $\alpha_1$ and $\alpha_2$ are complex numbers. If we define a new vector

$$
\lvert\phi\rangle=\alpha_1\lvert\phi_1\rangle+\alpha_2\lvert\phi_2\rangle,
$$

then the inner product distributes over the combination, with each coefficient pulled out in front:

$$
\langle\psi\vert\phi\rangle=\langle\psi\vert\bigl(\alpha_1\lvert\phi_1\rangle+\alpha_2\lvert\phi_2\rangle\bigr)=\alpha_1\langle\psi\vert\phi_1\rangle+\alpha_2\langle\psi\vert\phi_2\rangle.
$$

Conjugate linearity in the first argument

Suppose that $\lvert\psi_1\rangle$, $\lvert\psi_2\rangle$, and $\lvert\phi\rangle$ are vectors and $\beta_1$ and $\beta_2$ are complex numbers. If we define a new vector

$$
\lvert\psi\rangle=\beta_1\lvert\psi_1\rangle+\beta_2\lvert\psi_2\rangle,
$$

then the inner product is again linear in each term — but forming the bra $\langle\psi\vert$ conjugates every coefficient as it comes out in front:

$$
\langle\psi\vert\phi\rangle=\bigl(\overline{\beta_1}\langle\psi_1\vert+\overline{\beta_2}\langle\psi_2\vert\bigr)\lvert\phi\rangle=\overline{\beta_1}\langle\psi_1\vert\phi\rangle+\overline{\beta_2}\langle\psi_2\vert\phi\rangle.
$$

The Cauchy–Schwarz inequality

For every choice of vectors $\lvert\psi\rangle$ and $\lvert\phi\rangle$, the magnitude of the inner product can never be larger than the product of the vectors' lengths:

$$
\bigl\lvert\langle\psi\vert\phi\rangle\bigr\rvert\le\bigl\lVert\,\lvert\psi\rangle\,\bigr\rVert\;\bigl\lVert\,\lvert\phi\rangle\,\bigr\rVert.
$$

In other words, the overlap between two vectors cannot exceed what their lengths allow.

Equality, $\bigl\lvert\langle\psi\vert\phi\rangle\bigr\rvert=\bigl\lVert\,\lvert\psi\rangle\,\bigr\rVert\;\bigl\lVert\,\lvert\phi\rangle\,\bigr\rVert$, holds only when the two vectors are linearly dependent — that is, one is simply a scalar multiple of the other.

Orthogonality and orthonormality

Two vectors $\lvert\psi\rangle$ and $\lvert\phi\rangle$ are orthogonal if their inner product is zero:

$$
\langle\psi\vert\phi\rangle=0.
$$

An orthogonal set $\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\}$ is one in which every pair is orthogonal:

$$
\langle\psi_j\vert\psi_k\rangle=0\qquad(\text{for all }j\neq k).
$$

An orthonormal set $\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\}$ is an orthogonal set of unit vectors — each pair is orthogonal and each vector has length one:

$$
\langle\psi_j\vert\psi_k\rangle=\begin{cases}1 & j=k\\[2pt]0 & j\neq k\end{cases}\qquad(\text{for all }j,k).
$$

An orthonormal basis $\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\}$ is a set of orthonormal vectors that spans the entire vector space. This means every vector in the space can be written as a linear combination of the basis vectors.

Constructing basis sets

The key idea is that any orthonormal set can always be extended into an orthonormal basis. Suppose that $\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\}$ is an orthonormal set of vectors in an $n$-dimensional space. Because orthonormal sets are always linearly independent, these vectors span a subspace of dimension $m\leq n$.

If $m<n$, then there must exist additional vectors $\lvert\psi_{m+1}\rangle,\dots,\lvert\psi_n\rangle$ so that $\{\lvert\psi_1\rangle,\dots,\lvert\psi_n\rangle\}$ forms an orthonormal basis.

The Gram–Schmidt orthogonalization process can be used to construct these extra vectors. It repeatedly subtracts a vector's projections onto the existing basis vectors, leaving an orthogonal remainder that is then normalized to get the missing basis vectors.

Two dependent vectors — the span is only a line.

Independent vectors span the whole plane.

Orthonormal basis — vectors perpendicular and have unit length.

Orthonormal bases and unitary matrices

A unitary matrix is just a matrix whose columns are an orthonormal basis. More precisely, a square matrix $U$ is unitary if and only if its columns form an orthonormal basis — and equivalently, if and only if its rows do. These conditions are equivalent:

1. $U^\dagger U=I=UU^\dagger$ ($U$ is unitary).
2. The columns of $U$ form an orthonormal basis.
3. The rows of $U$ form an orthonormal basis.

To see why, write the columns of $U$ as vectors and take the conjugate transpose, which turns each column ket into the corresponding row bra. Multiplying $U^\dagger U$ then gives, in row $j$ and column $k$, the inner product $\langle\psi_j\vert\psi_k\rangle$:

$$
U=\left[\begin{array}{cccc}\textcolor{#0284c7}{\rvert}&\textcolor{#7c3aed}{\rvert}& &\textcolor{#db2777}{\rvert}\\[1pt]\textcolor{#0284c7}{\lvert\psi_1\rangle}&\textcolor{#7c3aed}{\lvert\psi_2\rangle}&\cdots&\textcolor{#db2777}{\lvert\psi_n\rangle}\\[1pt]\textcolor{#0284c7}{\rvert}&\textcolor{#7c3aed}{\rvert}& &\textcolor{#db2777}{\rvert}\end{array}\right]\qquad U^\dagger=\left[\begin{array}{ccc}\textcolor{#0284c7}{\rule[0.45ex]{1.3em}{0.5pt}}&\textcolor{#0284c7}{\langle\psi_1\rvert}&\textcolor{#0284c7}{\rule[0.45ex]{1.3em}{0.5pt}}\\[4pt]\textcolor{#7c3aed}{\rule[0.45ex]{1.3em}{0.5pt}}&\textcolor{#7c3aed}{\langle\psi_2\rvert}&\textcolor{#7c3aed}{\rule[0.45ex]{1.3em}{0.5pt}}\\[4pt]&\vdots&\\[4pt]\textcolor{#db2777}{\rule[0.45ex]{1.3em}{0.5pt}}&\textcolor{#db2777}{\langle\psi_n\rvert}&\textcolor{#db2777}{\rule[0.45ex]{1.3em}{0.5pt}}\end{array}\right]\qquad (U^\dagger U)_{jk}=\langle\psi_j\vert\psi_k\rangle
$$

For example, suppose two columns of the unitary matrix are the vectors $\lvert\psi_j\rangle$ and $\lvert\psi_k\rangle$ below. Taking the conjugate transpose of the first turns it into its bra (row) form $\langle\psi_j\rvert$, and the inner product $\langle\psi_j\vert\psi_k\rangle$ is exactly the dot product of the two columns, with complex conjugation applied to the first vector:

$$
\lvert\psi_j\rangle=\begin{pmatrix}a_1\\ a_2\\ \vdots\\ a_n\end{pmatrix},\quad\lvert\psi_k\rangle=\begin{pmatrix}b_1\\ b_2\\ \vdots\\ b_n\end{pmatrix}\qquad\langle\psi_j\rvert=(\overline{a_1},\dots,\overline{a_n})\qquad\langle\psi_j\vert\psi_k\rangle=\overline{a_1}\,b_1+\cdots+\overline{a_n}\,b_n
$$

There are two important cases:

- If $j=k$,
   
   $\langle\psi_j\vert\psi_j\rangle=|a_1|^2+|a_2|^2+\cdots+|a_n|^2=\lVert\psi_j\rVert^2.$
   
   So the diagonal entries of $U^\dagger U$ are the squared lengths of the columns.
- If $j\neq k$, then $\langle\psi_j\vert\psi_k\rangle$ measures how closely the two vectors point in the same direction.
   
   For real vectors, $\langle\psi_j\vert\psi_k\rangle=\lVert\psi_j\rVert\,\lVert\psi_k\rVert\cos\theta$, where $\theta$ is the angle between the vectors. Thus, the inner product measures their overlap:
   
   - large magnitude means they point in similar directions,
   - zero means they are perpendicular (orthogonal).

Therefore $U^\dagger U$ is the matrix of all pairwise inner products between the columns of $U$ — each diagonal entry a squared length, each off-diagonal entry the overlap between two different columns:

$$
U^\dagger U=\begin{pmatrix}\textcolor{#0d9488}{\langle\psi_1\vert\psi_1\rangle}&\textcolor{#d97706}{\langle\psi_1\vert\psi_2\rangle}&\cdots&\textcolor{#d97706}{\langle\psi_1\vert\psi_n\rangle}\\[2pt]\textcolor{#d97706}{\langle\psi_2\vert\psi_1\rangle}&\textcolor{#0d9488}{\langle\psi_2\vert\psi_2\rangle}&\cdots&\textcolor{#d97706}{\langle\psi_2\vert\psi_n\rangle}\\[2pt]\vdots&\vdots&\ddots&\vdots\\[2pt]\textcolor{#d97706}{\langle\psi_n\vert\psi_1\rangle}&\textcolor{#d97706}{\langle\psi_n\vert\psi_2\rangle}&\cdots&\textcolor{#0d9488}{\langle\psi_n\vert\psi_n\rangle}\end{pmatrix}
$$

If $U$ is unitary, then $U^\dagger U=I$, and matching the entries of these two matrices gives

- every diagonal entry equals $1$: $\langle\psi_i\vert\psi_i\rangle=1$, so every column has unit length;
- every off-diagonal entry equals $0$: $\langle\psi_i\vert\psi_j\rangle=0\ (i\neq j)$, so every pair of distinct columns is orthogonal.

So, for a unitary matrix, its columns form an orthonormal set, and since there are exactly $n$ columns in an $n$-dimensional space, an orthonormal set of $n$ vectors is automatically an orthonormal basis.

Applying the same argument to $UU^\dagger=I$ shows that the rows are also an orthonormal basis.

Projections

A square matrix $\Pi$ is called a projection if it satisfies two properties:

1. $\Pi=\Pi^\dagger$ (it is Hermitian);
2. $\Pi^2=\Pi$ (it is idempotent).

Example: a single unit vector

For example, if $\lvert\psi\rangle$ is a unit vector, then its outer product with itself is a projection:

$$
\Pi=\lvert\psi\rangle\langle\psi\rvert.
$$

To confirm this, we check the two defining properties in turn.

1. Hermitian — taking the conjugate transpose reverses the order of a product and daggers each factor, $(AB)^\dagger=B^\dagger A^\dagger$. Since $(\lvert\psi\rangle)^\dagger=\langle\psi\rvert$ and $(\langle\psi\rvert)^\dagger=\lvert\psi\rangle$, the two factors swap back into their original places:
   
   $\Pi^\dagger=(\lvert\psi\rangle\langle\psi\rvert)^\dagger=(\langle\psi\rvert)^\dagger(\lvert\psi\rangle)^\dagger=\lvert\psi\rangle\langle\psi\rvert=\Pi.$
2. Idempotent — applying $\Pi$ twice leaves the inner product $\langle\psi\vert\psi\rangle$ in the middle, and because $\lvert\psi\rangle$ is a unit vector that factor equals $\langle\psi\vert\psi\rangle=1$, so the product collapses back to a single copy:
   
   $\Pi^2=(\lvert\psi\rangle\langle\psi\rvert)^2=\lvert\psi\rangle\langle\psi\vert\psi\rangle\langle\psi\rvert=\lvert\psi\rangle\langle\psi\rvert=\Pi.$

Both properties hold, so $\Pi=\lvert\psi\rangle\langle\psi\rvert$ is indeed a projection.

Visual demo: projecting onto the line

With a single unit vector $\lvert\psi\rangle$, the projection operator $\Pi=\lvert\psi\rangle\langle\psi\rvert$ takes any vector $\lvert v\rangle$ and projects it onto the line spanned by $\lvert\psi\rangle$. The projection is $\Pi\lvert v\rangle=\lvert\psi\rangle\langle\psi\vert v\rangle$, where $\langle\psi\vert v\rangle$ is the scalar giving the component of $\lvert v\rangle$ along $\lvert\psi\rangle$. Thus $\Pi\lvert v\rangle$ always lies on the line, while the remaining vector $\lvert v\rangle-\Pi\lvert v\rangle$ is perpendicular to it. Applying the projection a second time changes nothing, because $\Pi\lvert v\rangle$ is already on the line, so $\Pi^2=\Pi$.

ψ = (0.94, 0.34)

v = (1.55, 1.35)

⟨ψ|v⟩ = 1.92

Πv = ⟨ψ|v⟩·ψ = (1.80, 0.66)

Drag v (violet) or rotate ψ (blue). Πv is the foot of the perpendicular — the closest point on the line.

Generalization: an orthonormal set

This generalizes from a single vector to any orthonormal set. If $\{\lvert\psi_1\rangle,\dots,\lvert\psi_m\rangle\}$ is orthonormal, then the sum of their outer products is again a projection:

$$
\Pi=\sum_{k=1}^{m}\lvert\psi_k\rangle\langle\psi_k\rvert.
$$

The same two checks go through, now using orthonormality $\langle\psi_j\vert\psi_k\rangle=\delta_{jk}$ (equal to $1$ when $j=k$ and $0$ otherwise).

1. Hermitian — the dagger passes through the sum and daggers each term, and every term $\lvert\psi_k\rangle\langle\psi_k\rvert$ is Hermitian by the argument above:
   
   $\Pi^\dagger=\Bigl(\sum_{k=1}^{m}\lvert\psi_k\rangle\langle\psi_k\rvert\Bigr)^\dagger=\sum_{k=1}^{m}(\lvert\psi_k\rangle\langle\psi_k\rvert)^\dagger=\sum_{k=1}^{m}\lvert\psi_k\rangle\langle\psi_k\rvert=\Pi.$
2. Idempotent — multiplying the two sums gives a double sum whose inner factor is $\langle\psi_j\vert\psi_k\rangle$. Orthonormality removes every cross term ($j\neq k$) and leaves $1$ on the diagonal ($j=k$), collapsing the double sum back to a single one:
   
   $\Pi^2=\sum_{j=1}^{m}\sum_{k=1}^{m}\lvert\psi_j\rangle\langle\psi_j\vert\psi_k\rangle\langle\psi_k\rvert=\sum_{k=1}^{m}\lvert\psi_k\rangle\langle\psi_k\rvert=\Pi.$

So any sum of outer products over an orthonormal set is a projection — geometrically, it projects onto the subspace those vectors span.

Projective measurements

A collection of projections $\{\Pi_1,\dots,\Pi_m\}$ that satisfies $\Pi_1+\dots+\Pi_m=I$ describes a projective measurement.

When such a measurement is performed on a system in the state $\lvert\psi\rangle$, two things happen:

1. The outcome $k\in\{1,\dots,m\}$ of the measurement is chosen randomly:
   
   $\begin{aligned}\Pr(\text{outcome is }k)&=\bigl\lVert\Pi_k\lvert\psi\rangle\bigr\rVert^2\\&=\langle\Pi_k\psi\vert\Pi_k\psi\rangle\\&=\langle\psi\rvert\Pi_k^\dagger\Pi_k\lvert\psi\rangle\\&=\langle\psi\rvert\Pi_k\Pi_k\lvert\psi\rangle\\&=\langle\psi\rvert\Pi_k\lvert\psi\rangle.\end{aligned}$
   
   The two forms $\lVert\Pi_k\lvert\psi\rangle\rVert^2$ and $\langle\psi\rvert\Pi_k\lvert\psi\rangle$ are mathematically identical. The first makes the geometry — the squared projection length — much more obvious, while the second connects naturally to the general framework of expectation values.
2. The state of the system becomes
   
   $\dfrac{\Pi_k\lvert\psi\rangle}{\lVert\Pi_k\lvert\psi\rangle\rVert}.$

The outcomes need not be labelled $1,\dots,m$. We are free to name them however is convenient — letters $a,b,c,\dots$, or any index set $\Gamma$, so that a family $\{\Pi_a:a\in\Gamma\}$ with $\sum_{a\in\Gamma}\Pi_a=I$ describes a projective measurement with outcomes in $\Gamma$. The rules are exactly the same.

Visual demo: measuring a qutrit in the standard basis

Take the three rank-one projections $\Pi_k=\lvert k\rangle\langle k\rvert$ onto the basis axes $\lvert0\rangle,\lvert1\rangle,\lvert2\rangle$. They add up to the identity, $\Pi_0+\Pi_1+\Pi_2=I$, so they form a projective measurement:

$$
\Pi_0+\Pi_1+\Pi_2=\begin{pmatrix}1&0&0\\0&0&0\\0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0\\0&1&0\\0&0&0\end{pmatrix}+\begin{pmatrix}0&0&0\\0&0&0\\0&0&1\end{pmatrix}=\begin{pmatrix}1&0&0\\0&1&0\\0&0&1\end{pmatrix}=I.
$$

If outcome $k$ occurs, the state $\lvert\psi\rangle$ is projected onto the corresponding subspace, and the probability $p_k=\langle\psi\rvert\Pi_k\lvert\psi\rangle=\lvert\alpha_k\rvert^2$ is the squared length of the projected vector — so the squared projection lengths (the probabilities) always sum to 1.

‖Π₀|ψ⟩‖

0.59² = 0.35

‖Π₁|ψ⟩‖

0.46² = 0.21

‖Π₂|ψ⟩‖

0.66² = 0.43

‖Π₀|ψ⟩‖² + ‖Π₁|ψ⟩‖² + ‖Π₂|ψ⟩‖² = 1.00

|ψ⟩ = 0.59|0⟩ + 0.46|1⟩ + 0.66|2⟩

Press Measure to sample an outcome. Drag the dark handle to move the state; drag the background to rotate the view.

Measuring one subsystem

Suppose we have two systems, $X$ and $Y$, but we only measure $X$ in the standard basis. We don't need a new mathematical procedure for measuring composite systems — just use the projectors.

$$
\{\textcolor{#2563eb}{\lvert a\rangle\langle a\rvert\otimes I_Y}:a\in\Sigma\}
$$

Each projector acts on the pair $(X,Y)$ in two independent steps:

- $\textcolor{#2563eb}{\lvert a\rangle\langle a\rvert}$ checks whether $X=a$ — it keeps the part of the state carrying that outcome and discards the rest.
- $\textcolor{#2563eb}{I_Y}$ does nothing to $Y$ — the second system is left exactly as it was.

Since these projectors sum to the identity, they form a valid projective measurement:

$$
\sum_{a\in\Sigma}\textcolor{#2563eb}{\bigl(\lvert a\rangle\langle a\rvert\otimes I_Y\bigr)}=\Bigl(\sum_{a\in\Sigma}\lvert a\rangle\langle a\rvert\Bigr)\otimes I_Y=I_X\otimes I_Y=I.
$$

The projective measurement rules apply as before:

$$
\Pr(\text{outcome is }a)=\bigl\lVert\textcolor{#2563eb}{\Pi_a}\lvert\psi\rangle\bigr\rVert^2=\bigl\lVert\textcolor{#2563eb}{(\lvert a\rangle\langle a\rvert\otimes I_Y)}\lvert\psi\rangle\bigr\rVert^2,
$$

and after observing outcome $a$, the state of $(X,Y)$ becomes

$$
\lvert\psi\rangle\;\longrightarrow\;\dfrac{\textcolor{#2563eb}{\Pi_a}\lvert\psi\rangle}{\bigl\lVert\textcolor{#2563eb}{\Pi_a}\lvert\psi\rangle\bigr\rVert}=\dfrac{\textcolor{#2563eb}{(\lvert a\rangle\langle a\rvert\otimes I_Y)}\lvert\psi\rangle}{\bigl\lVert\textcolor{#2563eb}{(\lvert a\rangle\langle a\rvert\otimes I_Y)}\lvert\psi\rangle\bigr\rVert}.
$$

Example — subsystem measurement by projection

For example, suppose $(X,Y)$ is a pair of qubits in the state below, written so that each term is grouped by $X$:

$$
\lvert\psi\rangle=\textcolor{#AB6108}{\lvert0\rangle}\otimes\textcolor{#0A74A9}{\Bigl(\tfrac{1}{\sqrt{2}}\,\lvert0\rangle+\tfrac{1}{2}\,\lvert1\rangle\Bigr)}+\textcolor{#AB6108}{\lvert1\rangle}\otimes\textcolor{#0A74A9}{\Bigl(\tfrac{i}{2\sqrt{2}}\,\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\,\lvert1\rangle\Bigr)}
$$

Outcome $a=0$

The projector for this outcome is

$$
\textcolor{#2563eb}{\Pi_0}=\textcolor{#2563eb}{\lvert0\rangle\langle0\rvert\otimes I_Y}.
$$

Apply it to the full state and expand. Each factor is handled separately, so the $\textcolor{#AB6108}{X}$ kets meet $\textcolor{#2563eb}{\langle0\rvert}$ while the identity $\textcolor{#2563eb}{I_Y}$ leaves the $\textcolor{#0A74A9}{Y}$ part unchanged:

$$
\begin{aligned}\textcolor{#2563eb}{\Pi_0}\lvert\psi\rangle&=\textcolor{#2563eb}{\bigl(\lvert0\rangle\langle0\rvert\otimes I_Y\bigr)}\Bigl[\textcolor{#AB6108}{\lvert0\rangle}\otimes\textcolor{#0A74A9}{\bigl(\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{2}\lvert1\rangle\bigr)}+\textcolor{#AB6108}{\lvert1\rangle}\otimes\textcolor{#0A74A9}{\bigl(\tfrac{i}{2\sqrt{2}}\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\lvert1\rangle\bigr)}\Bigr]\\[6pt]&=\Bigl(\textcolor{#2563eb}{\lvert0\rangle}\underbrace{\textcolor{#2563eb}{\langle0\vert}\textcolor{#AB6108}{0\rangle}}_{\textcolor{#16a34a}{=\,1}}\Bigr)\otimes\underbrace{\textcolor{#2563eb}{I_Y}}_{\text{no effect}}\textcolor{#0A74A9}{\bigl(\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{2}\lvert1\rangle\bigr)}+\Bigl(\textcolor{#2563eb}{\lvert0\rangle}\underbrace{\textcolor{#2563eb}{\langle0\vert}\textcolor{#AB6108}{1\rangle}}_{\textcolor{#dc2626}{=\,0}}\Bigr)\otimes\underbrace{\textcolor{#2563eb}{I_Y}}_{\text{no effect}}\textcolor{#0A74A9}{\bigl(\tfrac{i}{2\sqrt{2}}\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\lvert1\rangle\bigr)}\\[6pt]&=\textcolor{#2563eb}{\lvert0\rangle}\otimes\textcolor{#0A74A9}{\bigl(\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{2}\lvert1\rangle\bigr)}.\end{aligned}
$$

The probability is the squared norm of what survived — the sum of its squared amplitude magnitudes:

$$
\bigl\lVert\textcolor{#2563eb}{\Pi_0}\lvert\psi\rangle\bigr\rVert^2=\left|\tfrac{1}{\sqrt{2}}\right|^2+\left|\tfrac{1}{2}\right|^2=\tfrac{1}{2}+\tfrac{1}{4}=\tfrac{3}{4}
$$

The survivor is not a unit vector, so divide it by its norm $\sqrt{\tfrac{3}{4}}=\tfrac{\sqrt{3}}{2}$:

$$
\frac{\lvert0\rangle\otimes\Bigl(\tfrac{1}{\sqrt{2}}\lvert0\rangle+\tfrac{1}{2}\lvert1\rangle\Bigr)}{\sqrt{3/4}}=\lvert0\rangle\otimes\Bigl(\sqrt{\tfrac{2}{3}}\,\lvert0\rangle+\tfrac{1}{\sqrt{3}}\,\lvert1\rangle\Bigr)
$$

Outcome $a=1$

The other projector does the mirror image: now $\textcolor{#2563eb}{\langle1\rvert}$ keeps the $\lvert1\rangle$ branch and discards the $\lvert0\rangle$ one:

$$
\textcolor{#2563eb}{(\lvert1\rangle\langle1\rvert\otimes I_Y)}\lvert\psi\rangle=\lvert1\rangle\otimes\Bigl(\tfrac{i}{2\sqrt{2}}\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\lvert1\rangle\Bigr)
$$

Its squared norm is the remaining probability. The $i$ drops out under $\lvert\cdot\rvert^2$ — only magnitudes matter:

$$
\bigl\lVert\textcolor{#2563eb}{(\lvert1\rangle\langle1\rvert\otimes I_Y)}\lvert\psi\rangle\bigr\rVert^2=\left|\tfrac{i}{2\sqrt{2}}\right|^2+\left|-\tfrac{1}{2\sqrt{2}}\right|^2=\tfrac{1}{8}+\tfrac{1}{8}=\tfrac{1}{4}
$$

Divide by the norm $\sqrt{\tfrac{1}{4}}=\tfrac{1}{2}$ as before:

$$
\frac{\lvert1\rangle\otimes\Bigl(\tfrac{i}{2\sqrt{2}}\lvert0\rangle-\tfrac{1}{2\sqrt{2}}\lvert1\rangle\Bigr)}{\sqrt{1/4}}=\lvert1\rangle\otimes\Bigl(\tfrac{i}{\sqrt{2}}\,\lvert0\rangle-\tfrac{1}{\sqrt{2}}\,\lvert1\rangle\Bigr)
$$

Implementing projective measurements

So far the projectors have been pure mathematics. Hardware offers only two things: unitary operations and standard basis measurements. That is already enough — any projective measurement can be assembled out of those two.

The trick is to bring in one extra system alongside the one being measured, with a classical state for each possible outcome. A single unitary then files each outcome into its own branch: the extra system carries the outcome's label, and travelling alongside it is the matching projected state of the measured system. Reading the extra system in the standard basis picks one branch — with exactly the probability the projection rule demands — and leaves the measured system in that branch's projected, renormalised state. No projector is ever built as hardware; the projections emerge from the branch structure.

Below, that idea runs on a pair of systems $X$ and $Y$. The bottom wire is the extra qubit — the only thing ever read.

XY$\lvert0\rangle$

The first Hadamard splits the extra qubit into two branches. The controlled-SWAP exchanges $X$ and $Y$ on one branch only, so the two branches now hold the state seen from two different angles. The second Hadamard makes them interfere, which sorts the state into a part unchanged by the exchange and a part that the exchange reverses — and those two parts are precisely the projections. Measuring the extra qubit says which one you landed in, and turns it into a classical bit.

Three gates and a single qubit read out: an abstract pair of projections has become something a machine actually does.

## Limitations of quantum measurements

Irrelevance of global phases

Quantum amplitudes are complex numbers, so each one carries both a magnitude and a phase. Phases show up in two different ways:

- A global phase multiplies every amplitude in a state by the same unit-modulus factor $e^{i\theta}$.
- A relative phase changes the phase *difference* between the amplitudes within a superposition.

The distinction is fundamental. A global phase rotates the entire state vector by the same amount, leaving the relationships between its amplitudes unchanged. A relative phase changes those relationships and therefore affects interference.

A scalar factor in a global phase must preserve the norm, so $\lvert\alpha\rvert=1$. That condition is met exactly by the complex numbers of the form $\alpha=e^{i\theta}$ for some real $\theta$: every such number sits on the unit circle of the complex plane, and every unit-modulus number can be written this way. Two state vectors therefore differ by a global phase when $\lvert\phi\rangle=e^{i\theta}\lvert\psi\rangle$ — every amplitude is multiplied by the same unit-modulus factor $e^{i\theta}$.

For example, the two amplitudes of $\lvert\psi\rangle=\alpha_0\lvert0\rangle+\alpha_1\lvert1\rangle$ are drawn as arrows on the complex plane. Their lengths are the magnitudes, their angles the phases.

Drag global phase — every bar holds still, so the state is physically unchanged. Drag relative phase — the standard basis still won't move, but the ± measurement swings: the relative phase is observable. The amplitude split resizes the arrows and shifts the standard-basis odds.

Global phase (both arrows)0°

Relative phase (α₁ only)90°

Amplitude split (|α₀| vs |α₁|)0.71 · 0.71

Probability of each outcome if we measure $\lvert\psi\rangle$ in that basis:

Standard basis$\{\lvert0\rangle,\lvert1\rangle\}$

$P(0)=\lvert\langle0\vert\psi\rangle\rvert^2$50%

$P(1)=\lvert\langle1\vert\psi\rangle\rvert^2$50%

Hadamard basis$\{\lvert+\rangle,\lvert-\rangle\}$

$P(+)=\lvert\langle+\vert\psi\rangle\rvert^2$50%

$P(-)=\lvert\langle-\vert\psi\rangle\rvert^2$50%

Mathematically, two state vectors differ by a global phase if $\lvert\phi\rangle=e^{i\theta}\lvert\psi\rangle$.

Standard-basis measurement. The probability of outcome $a$ is the squared magnitude of the amplitude $\langle a\vert\phi\rangle$. Substituting $\lvert\phi\rangle=e^{i\theta}\lvert\psi\rangle$, the phase factors out and its magnitude $\lvert e^{i\theta}\rvert^2=1$ disappears:

$$
\bigl\lvert\langle a\vert\phi\rangle\bigr\rvert^2=\bigl\lvert e^{i\theta}\,\langle a\vert\psi\rangle\bigr\rvert^2=\lvert e^{i\theta}\rvert^2\,\bigl\lvert\langle a\vert\psi\rangle\bigr\rvert^2=\bigl\lvert\langle a\vert\psi\rangle\bigr\rvert^2.
$$

Projective measurement. The same cancellation holds for any projective measurement $\{\Pi_1,\dots,\Pi_m\}$. Each outcome probability is the squared norm $\lVert\Pi_k\lvert\phi\rangle\rVert^2$, and pulling the scalar $e^{i\theta}$ out of the norm leaves the factor $\lvert e^{i\theta}\rvert^2=1$ again:

$$
\bigl\lVert\Pi_k\lvert\phi\rangle\bigr\rVert^2=\bigl\lVert e^{i\theta}\,\Pi_k\lvert\psi\rangle\bigr\rVert^2=\lvert e^{i\theta}\rvert^2\,\bigl\lVert\Pi_k\lvert\psi\rangle\bigr\rVert^2=\bigl\lVert\Pi_k\lvert\psi\rangle\bigr\rVert^2.
$$

Therefore all measurements produce exactly the same statistics for $\lvert\psi\rangle$ and $e^{i\theta}\lvert\psi\rangle$. States that differ by a global phase are considered equivalent — they represent the same physical state. A relative phase, by contrast, changes the interference between amplitudes and is observable, as the demo above shows.

A note on representation. This global phase is a degeneracy of the state-vector picture: the same physical state maps to a whole circle of vectors $\{e^{i\theta}\lvert\psi\rangle\}$ that are all indistinguishable. It is an artifact of describing states at this simplified level of generality (it's called “simplified,” though personally I cried at the word “simple”).

The more general formalism uses density matrices: a state vector $\lvert\psi\rangle$ is replaced by the operator $\rho=\lvert\psi\rangle\langle\psi\rvert$. Here the global phase cancels automatically, since $\bigl(e^{i\theta}\lvert\psi\rangle\bigr)\bigl(e^{i\theta}\lvert\psi\rangle\bigr)^{\dagger}=e^{i\theta}e^{-i\theta}\,\lvert\psi\rangle\langle\psi\rvert=\rho$, so equivalent states share exactly one density matrix and the redundancy disappears. Density matrices also describe mixed states (classical uncertainty over several vectors), which no single state vector can capture.

No-cloning theorem

Copying classical information is trivial — you read a bit and write it down twice. It is natural to ask whether a quantum computer can do the same for an unknown state $\lvert\psi\rangle$, producing two identical copies of it. The no-cloning theorem says this is impossible: no single unitary can duplicate an arbitrary unknown quantum state, which is exactly why quantum information cannot simply be copied and why quantum key distribution is secure.

Formally, let $X$ and $Y$ both have the classical state set $\{0,\dots,d-1\}$ with $d\ge 2$. A cloner would be a unitary $U$ on the pair $(X, Y)$ that takes the state in $X$ with a blank $\lvert0\rangle$ in $Y$ and writes a copy into $Y$. The theorem states no such $U$ exists:

$$
\forall\,\lvert\psi\rangle:\quad U\bigl(\lvert\psi\rangle\otimes\lvert0\rangle\bigr)=\lvert\psi\rangle\otimes\lvert\psi\rangle.
$$

Drawn as a circuit, this cloner would feed $\lvert\psi\rangle$ and a blank register $\lvert0\cdots0\rangle$ into $U$ and read out two copies — the box below that cannot exist for every input:

$\lvert0\cdots0\rangle$$\lvert\psi\rangle$$\lvert\psi\rangle$$\lvert\psi\rangle$

No such U exists

Why no such U can exist

Suppose there existed a unitary operator $U$ that could perfectly clone any quantum state. For every state $\textcolor{#2563eb}{\lvert\psi\rangle}$, it would satisfy $U\bigl(\textcolor{#2563eb}{\lvert\psi\rangle}\otimes\lvert0\rangle\bigr)=\textcolor{#2563eb}{\lvert\psi\rangle}\otimes\textcolor{#2563eb}{\lvert\psi\rangle}$, where $\lvert0\rangle$ is a blank qubit that receives the copy.

Applying $U$ to basis states results in: $U\bigl(\textcolor{#2563eb}{\lvert0\rangle}\lvert0\rangle\bigr)=\textcolor{#2563eb}{\lvert0\rangle}\textcolor{#2563eb}{\lvert0\rangle},\quad U\bigl(\textcolor{#2563eb}{\lvert1\rangle}\lvert0\rangle\bigr)=\textcolor{#2563eb}{\lvert1\rangle}\textcolor{#2563eb}{\lvert1\rangle}.$

But now consider a plus (superposition) state $\lvert+\rangle=\dfrac{\lvert0\rangle+\lvert1\rangle}{\sqrt2}$. Because every quantum gate is linear, the cloning operation must satisfy:

$$
\begin{aligned}
U\bigl(\lvert+\rangle\lvert0\rangle\bigr)
&=U\!\left(\dfrac{\lvert0\rangle+\lvert1\rangle}{\sqrt2}\otimes\lvert0\rangle\right)\\[4pt]
&=\dfrac{1}{\sqrt2}\,U\bigl(\textcolor{#2563eb}{\lvert0\rangle}\lvert0\rangle\bigr)+\dfrac{1}{\sqrt2}\,U\bigl(\textcolor{#2563eb}{\lvert1\rangle}\lvert0\rangle\bigr)\\[4pt]
&=\dfrac{1}{\sqrt2}\,\textcolor{#2563eb}{\lvert0\rangle}\textcolor{#2563eb}{\lvert0\rangle}+\dfrac{1}{\sqrt2}\,\textcolor{#2563eb}{\lvert1\rangle}\textcolor{#2563eb}{\lvert1\rangle}\\[4pt]
&=\dfrac{\lvert00\rangle+\lvert11\rangle}{\sqrt2}.
\end{aligned}
$$

But if $U$ were truly a cloning machine, the output should instead be

$$
U\bigl(\textcolor{#2563eb}{\lvert+\rangle}\lvert0\rangle\bigr)=\textcolor{#2563eb}{\lvert+\rangle}\textcolor{#2563eb}{\lvert+\rangle}.
$$

This expands to

$$
\lvert+\rangle\lvert+\rangle=\left(\dfrac{\lvert0\rangle+\lvert1\rangle}{\sqrt2}\right)\otimes\left(\dfrac{\lvert0\rangle+\lvert1\rangle}{\sqrt2}\right)=\dfrac{\lvert00\rangle+\lvert01\rangle+\lvert10\rangle+\lvert11\rangle}{2}.
$$

These two states are different:

$$
\dfrac{\lvert00\rangle+\lvert11\rangle}{\sqrt2}\;\neq\;\dfrac{\lvert00\rangle+\lvert01\rangle+\lvert10\rangle+\lvert11\rangle}{2}.
$$

The contradiction arises because linearity forces one output, while perfect cloning requires another. Therefore, no unitary operation can perfectly clone an arbitrary unknown quantum state.

Remarks

- Approximate forms of the cloning theorem are known.
- Copying a standard basis state is possible — the no-cloning theorem does not contradict this.
   
   For example, a $\mathrm{CNOT}$ controlled by a basis value $\lvert a\rangle$ copies it into a blank $\lvert0\rangle$, giving $\lvert a\rangle\lvert a\rangle$:
   
   $\lvert0\rangle$$\lvert a\rangle$$\lvert a\rangle$$\lvert a\rangle$
- Cloning a probabilistic state (classically) is also impossible.
- Perfect clones *are* possible if they are all encrypted. A 2026 protocol, *encrypted cloning*, deterministically produces any number of perfect copies of an unknown state — as long as the copies are simultaneously locked with a single-use quantum decryption key. Decrypting one clone consumes the key and renders every other clone indecipherable, so no two usable copies ever coexist and the theorem still holds. The real constraint is not the copying, but that the decryption mechanism must be single-use. The payoff is redundancy, like keeping backups: you hold many encrypted copies and recover the original from any one that survives. It has been demonstrated on IBM Heron-R2 hardware with up to 154 qubits ([Yamaguchi et al., 2026](https://arxiv.org/abs/2602.10695)).

Discriminating non-orthogonal states

It is not possible to perfectly discriminate two non-orthogonal quantum states. Equivalently, if we can discriminate two quantum states perfectly, then they must be orthogonal.

Two states $\lvert\psi\rangle$ and $\lvert\phi\rangle$ can be discriminated perfectly if there is a unitary operation $U$ that works like this:

$\lvert\pi_0\rangle$0

$$
U\bigl(\lvert0\cdots0\rangle\lvert\psi\rangle\bigr)=\lvert\pi_0\rangle\lvert0\rangle
$$

$\lvert\pi_1\rangle$1

$$
U\bigl(\lvert0\cdots0\rangle\lvert\phi\rangle\bigr)=\lvert\pi_1\rangle\lvert1\rangle
$$

Suppose a unitary operator $U$ perfectly distinguishes two states $\lvert\psi\rangle$ and $\lvert\phi\rangle$. By definition,

$$
\begin{aligned}
U\bigl(\lvert0\cdots0\rangle\lvert\psi\rangle\bigr)&=\lvert\pi_0\rangle\lvert0\rangle,\\[4pt]
U\bigl(\lvert0\cdots0\rangle\lvert\phi\rangle\bigr)&=\lvert\pi_1\rangle\lvert1\rangle,
\end{aligned}
$$

where the final qubit stores the measurement result and the remaining qubits $\lvert\pi_0\rangle$ and $\lvert\pi_1\rangle$ represent arbitrary ancilla states.

The overlap of two states is their inner product $\langle a\vert b\rangle$— a number that measures how similar they are. It is $1$ for identical states and $0$ for orthogonal (perfectly distinguishable) ones.

Now take two states $\lvert a\rangle$ and $\lvert b\rangle$ and apply $U$ to each. To form the overlap of the outputs, the first ket $U\lvert a\rangle$ becomes a bra by taking its conjugate transpose, which flips $U$ into $U^\dagger$:

$$
\langle Ua\vert Ub\rangle=\langle a\rvert\,U^\dagger U\,\lvert b\rangle=\langle a\rvert\,I\,\lvert b\rangle=\langle a\vert b\rangle,
$$

where the middle $U^\dagger U$ collapses to the identity $I$ because $U$ is unitary. So passing both states through the same $U$ leaves their overlap unchanged — the input overlap equals the output overlap.

For the discriminating unitary, the two input vectors are $\lvert0\cdots0\rangle\lvert\psi\rangle$ and $\lvert0\cdots0\rangle\lvert\phi\rangle$. Their overlap factorizes across the ancilla and state registers, and because the ancilla register is the same in both inputs $\langle0\cdots0\vert0\cdots0\rangle=1$:

$$
\bigl(\langle0\cdots0\rvert\langle\psi\rvert\bigr)\bigl(\lvert0\cdots0\rangle\lvert\phi\rangle\bigr)=\langle0\cdots0\vert0\cdots0\rangle\cdot\langle\psi\vert\phi\rangle=\langle\psi\vert\phi\rangle.
$$

The output overlap factorizes the same way. The measurement outcomes are different, so $\langle0\vert1\rangle=0$, and the whole overlap collapses to zero:

$$
\bigl(\langle\pi_0\rvert\langle0\rvert\bigr)\bigl(\lvert\pi_1\rangle\lvert1\rangle\bigr)=\langle\pi_0\vert\pi_1\rangle\,\langle0\vert1\rangle=\langle\pi_0\vert\pi_1\rangle\cdot0=0.
$$

Equating the input and output overlaps, and recalling the output overlap is zero, gives

$$
\boxed{\;\langle\psi\vert\phi\rangle=0.\;}
$$

In other words, perfect discrimination is possible only for orthogonal quantum states.

The demo illustrates this result visually. Drag either state to change their overlap $\langle\psi\vert\phi\rangle$. As the overlap decreases, the states become easier to distinguish. When $\langle\psi\vert\phi\rangle=0$, they are orthogonal and can be distinguished perfectly.

θ = 60°

⟨ψ|ϕ⟩ = cos θ = 0.50

best distinguishing probability0.93

0.5 · coin flip1.0 · perfect

Non-orthogonal — the overlap is nonzero, so no measurement can tell the two apart with certainty. Drag the arrows 90° apart.

## Entanglement

Two qubits are entangled when their joint state cannot be written as $|a\rangle\otimes|b\rangle$. Measure both systems several times and compare the results: the separable state behaves like two independent coin flips, while the entangled Bell state always produces matching outcomes.

Separable$|{+}\rangle\otimes|{+}\rangle=\tfrac12|00\rangle+\tfrac12|01\rangle+\tfrac12|10\rangle+\tfrac12|11\rangle$

✗ Different

|00⟩

0.00

|01⟩

0.00

|10⟩

0.00

|11⟩

0.00

CorrelationIndependent • 0%

Entangled$|\phi^{+}\rangle=\tfrac{1}{\sqrt2}|00\rangle+\tfrac{1}{\sqrt2}|11\rangle$

✗ Different

|00⟩

0.00

|01⟩

0.00

|10⟩

0.00

|11⟩

0.00

CorrelationPerfectly correlated • 0%

0 measurements

These correlations are stronger than anything independent classical systems can share, so entanglement is treated as a resource. One maximally entangled pair $|\phi^{+}\rangle$ is one unit of it — an e-bit.

Quantum teleportation

Quantum teleportation is a protocol that allows a sender to send quantum information to a receiver using only entanglement and classical communication to accomplish that transmission.

Setup

- Alice holds a qubit $Q$ in an unknown state $|\psi\rangle$ that she wants to transfer to Bob.
- Alice and Bob share an entangled pair (an e-bit) in the state $|\phi^{+}\rangle$. Alice holds qubit $A$, and Bob holds qubit $B$. How or when they established this shared entanglement—for example, during an earlier meeting—is irrelevant to the protocol.
- Alice can communicate with Bob only by sending classical bits.
- An unknown quantum state cannot be completely described by classical bits, so classical communication alone is insufficient.
- Because of the no-cloning theorem, once the protocol is complete and Bob's qubit is in state $|\psi\rangle$, Alice no longer has a copy of that quantum state.

$|\psi\rangle$QABAliceBob

$$
|\psi\rangle
$$

Protocol

1. 1Alice performs a controlled-NOT operation, where $Q$ is the control and $A$ is the target.
2. 2Alice performs a Hadamard operation on $Q$.
3. 3Alice measures $A$ and $Q$, obtaining binary outcomes $m_A$ and $m_Q$, respectively.
4. 4Alice sends $m_A$ and $m_Q$ to Bob.
5. 5
   
   Bob performs these two steps on qubit $B$:
   
   - If $m_A = 1$, Bob applies an $X$ operation.
   - If $m_Q = 1$, Bob applies a $Z$ operation.

Superdense coding

Superdense coding is a protocol that allows a sender to transmit two classical bits to a receiver by sending only a single qubit, using one shared e-bit of entanglement to accomplish that transmission.

Scenario

- Alice has two classical bits $(a,b)$ that she wishes to transmit to Bob.
- Alice is able to send only a $single\ qubit$ to Bob.
- Alice and Bob already share an entangled pair (an e-bit) in the state $|\phi^{+}\rangle$.
- Without the e-bit the task would be impossible: by Holevo's theorem, two classical bits cannot be reliably transmitted by a single qubit alone.

baAliceBob

Protocol

1. 1Alice applies $X^{a}Z^{b}$ to her qubit — an $X$ when $a=1$ and a $Z$ when $b=1$.
2. 2Alice sends her qubit to Bob.
3. 3Bob applies a controlled-NOT, with the qubit received from Alice as the control and his own qubit as the target.
4. 4Bob applies a Hadamard to the qubit received from Alice.
5. 5Bob measures both qubits, reading off $a$ and $b$.

CHSH game

A nonlocal game is a mathematical and physical framework modeling two or more cooperating players who try to win a game against a referee. The defining rule is that once the game begins, players cannot communicate.

Set-up

- The players Alice and Bob cooperate as a team against a referee.
- The referee runs the game: it sends each player a question and checks their answers against a fixed rule.
- Alice and Bob may prepare a strategy together beforehand — including sharing entanglement.
- But once the game starts they are forbidden from communicating: neither learns the other's question or answer.

Alice

Referee

$x$$a$$b$$y$

One round

The referee asks. The referee picks two questions using randomness and sends one to each player — $x$ to Alice and $y$ to Bob. Each player sees only their own question.

The CHSH referee

The CHSH game is one example of a nonlocal game, in which the referee follows these rules. Questions and answers are all bits $x,y,a,b \in \{0,1\}$, the questions $x$ and $y$ are chosen uniformly at random, and the team wins exactly when $a \oplus b = x \wedge y$.

| $(x,y)$ | $x \wedge y$ | Winning condition |
| --- | --- | --- |
| (0,0) | 0 | $a = b$ |
| (0,1) | 0 | $a = b$ |
| (1,0) | 0 | $a = b$ |
| (1,1) | 1 | $a \neq b$ |

CHSH — Deterministic strategy

Program Alice and Bob before the game begins. For each possible question, choose the answer they will always give. Once the game starts, they cannot communicate or change their strategy. Can you find a strategy that wins all four rounds?

Strategy

Alice

If asked $x=0$, answer

If asked $x=1$, answer

If asked $y=0$, answer

If asked $y=1$, answer

CHSH game results

(0,0)$x \wedge y$= 0need$a = b$a=0, b=0→$a = b$

✓ Win

(0,1)$x \wedge y$= 0need$a = b$a=0, b=0→$a = b$

✓ Win

(1,0)$x \wedge y$= 0need$a = b$a=0, b=0→$a = b$

✓ Win

(1,1)$x \wedge y$= 1need$a \neq b$a=0, b=0→$a = b$

✗ Lose

3 / 4 Wins

75%

🏆 This is the best possible deterministic strategy.

CHSH — Probabilistic strategy

Now let Alice and Bob answer at random. For each question, set how often they reply with a 1. The game is unchanged — only the strategy is now a coin flip. Can randomness push them past 75%?

Strategy

Alice

$P(a{=}1 \mid x{=}0)$50%

$P(a{=}1 \mid x{=}1)$50%

$P(b{=}1 \mid y{=}0)$50%

$P(b{=}1 \mid y{=}1)$50%

CHSH game results

0 wins2,000 losses

Only a few thousand rounds, so the result is noisy — a lucky sample can drift above or below the true rate. Re-run it a few times to see it bounce around the expected value.

0.0%

win rate over 2,000 rounds

expected 50.0% · best possible 75%

CHSH — Quantum strategy

Now Alice and Bob share an entangled pair prepared before the game. Each question picks a measurement angle instead of a fixed answer. Same game, same rule — but can entanglement beat the classical 75% ceiling?

Measurements are angles

Measuring a qubit is not simply reading a fixed bit — the player first chooses a direction to measure along. That choice is an angle $\theta$: outcome 0 corresponds to the direction $|\psi_\theta\rangle$, and outcome 1 to the perpendicular direction $|\psi_{\theta+\pi/2}\rangle$. The measurement asks: which of these two is the qubit closer to?

Measurement basis

$$
|\psi_\theta\rangle = \cos(\theta)\,|0\rangle + \sin(\theta)\,|1\rangle
$$

$$
\angle\,\theta = 67.5^\circ
$$

$\cos\theta = 0.383$$\sin\theta = 0.924$

| θ | deg | cos θ | sin θ |
| --- | --- | --- | --- |
| $0$ | 0° | $1$ | $0$ |
| $\tfrac{\pi}{8}$ | 22.5° | $\tfrac{\sqrt{2+\sqrt{2}}}{2}$ | $\tfrac{\sqrt{2-\sqrt{2}}}{2}$ |
| $\tfrac{\pi}{4}$ | 45° | $\tfrac{1}{\sqrt{2}}$ | $\tfrac{1}{\sqrt{2}}$ |
| $\tfrac{3\pi}{8}$ | 67.5° | $\tfrac{\sqrt{2-\sqrt{2}}}{2}$ | $\tfrac{\sqrt{2+\sqrt{2}}}{2}$ |
| $\tfrac{\pi}{2}$ | 90° | $0$ | $1$ |

So how does the angle determine the probability?

The qubit is also represented by a direction on the same circle. Suppose it points at angle $\varphi$, so its state is $|\psi_\varphi\rangle$.

A measurement compares the qubit's direction with the measurement direction $\theta$. The closer they are, the more likely the measurement returns outcome 0. If they point in exactly the same direction, the result is always 0. If they are perpendicular, outcome 0 is impossible.

Quantum mechanics quantifies this “closeness” using the inner product (also called the overlap). For a qubit pointing at angle $\varphi$ and a measurement at angle $\theta$, the overlap between the qubit state $|\psi_\varphi\rangle$ and the measurement direction $|\psi_\theta\rangle$ is

$$
\langle \psi_\theta | \psi_\varphi \rangle = \cos(\theta - \varphi).
$$

The probability of obtaining a measurement outcome is the square of the overlap with the corresponding measurement direction. Since outcome 0 corresponds to $|\psi_\theta\rangle$,

$$
\begin{aligned}\Pr[\text{outcome }0] &= \big|\langle \psi_\theta | \psi_\varphi \rangle\big|^2 \\ &= \cos^2(\theta - \varphi).\end{aligned}
$$

$$
\Pr[\text{outcome }0] = \big|\langle \psi_\theta | \psi_\varphi \rangle\big|^2 = \cos^2(\theta - \varphi).
$$

Likewise, outcome 1 corresponds to the perpendicular direction $|\psi_{\theta+\pi/2}\rangle$, so

$$
\Pr[\text{outcome }1] = \sin^2(\theta - \varphi).
$$

The simplest choice is $\theta = 0^\circ$, where the measurement direction lines up with the horizontal axis. In this case,

$$
|\psi_0\rangle = |0\rangle, \qquad |\psi_{\pi/2}\rangle = |1\rangle.
$$

So the two measurement outcomes are simply the familiar states $|0\rangle$ and $|1\rangle$. This is called the standard (or computational, $Z$) basis. Measuring in this basis is the familiar question: “Is the qubit 0 or 1?”

Of course, nothing requires us to measure in this basis. We can rotate the measurement direction to any angle, creating a different pair of measurement states. A few of these angles are used so often that they have their own names. Each basis is simply a different measurement angle — choose one below, or drag the arrows to rotate the measurement basis and watch the outcome probabilities change while the qubit itself stays fixed.

$\theta = 0^\circ$standard basis

outcome 0 · cos²(θ−φ)33%

outcome 1 · sin²(θ−φ)67%

From one qubit to an entangled pair

So far we've measured a single qubit. Now imagine Alice and Bob each receive one qubit from a shared entangled pair prepared before the game.

Just as before, each player independently chooses a measurement angle — Alice uses $\alpha$, Bob uses $\beta$. Each individual measurement still looks completely random: Alice sees 0 or 1 with equal probability, and so does Bob.

The surprise is that their outcomes are correlated. The probability that they obtain the same result depends only on the angle between their measurements:

$$
\Pr[\text{Alice} = \text{Bob}] = \cos^2(\alpha - \beta).
$$

Just like for a single qubit, only the difference between the two angles matters — not their absolute positions.

- If $\alpha = \beta$, Alice and Bob always obtain the same result.
- If the measurement directions are $90^\circ$ apart, they always obtain opposite results.
- Between these extremes, the probability changes smoothly as the angle changes.

α = 0°β = 30°|α−β| = 30.0°

same outcome · cos²(α−β)75%

different · sin²(α−β)25%

The CHSH game circuit

Alice and Bob begin with the shared Bell state $|\phi^{+}\rangle$. Their CHSH questions don't determine the answers — they determine which measurement basis each player uses. In the circuit below, question $x$ selects Alice's rotation and question $y$ selects Bob's. After applying these rotations, both qubits are measured in the standard basis.

$|\phi^{+}\rangle$$y$$x$BobAlice

$b$$a$

Strategy

Alice

angle Alice measures at for each question x

question $x = 0$→$A_0$0°

question $x = 1$→$A_1$45°

angle Bob measures at for each question y

question $y = 0$→$B_0$22.5°

question $y = 1$→$B_1$157.5°

CHSH game results

0 wins2,000 losses

0.0%

win rate over 2,000 rounds

expected 85.4% · maximum 85.36% (Tsirelson bound)

⚛️ Optimal — this reaches the Tsirelson bound, the quantum maximum.

These notes are based on John Watrous's [Understanding Quantum Information and Computation](https://www.youtube.com/playlist?list=PLOFEBzvs-VvqKKMXX4vbi4EB1uaErFMSO), made with IBM Quantum, with additional explanations, derivations and interactive demos.
