Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Introduction to quantum mechanics

I will start with a quasi-historical, phenomenological discussion of some basic aspects of quantum mechanics:

  1. The need for a new dimensionful quantity ℏ\hbar with units [Energy ×\times time].

  2. The quantization of the energy of light with frequency ω\omega into parcels (“photons”) with energy ℏω\hbar \omega,

  3. The fact that quantum states can be written as complex vectors, and that the sum of two such vectors is another legitimate quantum state.

  4. The probabilistic nature of the outcome of quantum measurements.

These are all important; point (3) motivates a serious survey/review of linear algebra, which is also needed to discuss (4) more precisely.

Blackbody radiation

Consider a cavity with walls at some temperature TT, such that the electromagnetic field inside of the cavity is at equilibrium with the walls

black body model

What is the expected energy density inside the cavity? We could (following many textbooks) try to comute this from classical electromagnetism, but we will appeal to dimensional analysis.

The total energy density should be an integral over all allowed frequencies:

E(T)=∫0∞dωE(T,ω)\CE(T) = \int_0^{\infty}d\omega \CE(T,\omega)

Here we are assuming that the volume of the box only appears in the total energy E=VE(T)E = V \CE(T), which is a thermodynamically extensive quantity. This makes sense if the box is very large compared to the wavelength of the light inside.

We have the following dimensionful quantities available to construct E(T,ω)\CE(T,\omega):

  1. The frequency ω\omega.

  2. The temperature writen as a thermal energy kBTk_B T.

  3. The speed of light cc.

To get an energy density we can make an energy from kBTk_B T and a length scale c/ωc/\omega (proportional to the wavelength). E(T,ω)\CE(T,\omega) should have units of energy denisty per frequency. The only combination of the above yielding a quantity with the right dimensions is:

E(T,ω)=AkBT(ωc)31ω=AkBTc3ω2\CE(T, \omega) = A k_B T \left(\frac{\omega}{c}\right)^3 \frac{1}{\omega} = A \frac{k_B T}{c^3} \omega^2

where AA is a dimensionless constant that requires a first-principles calculation to obtain. This is called the “Rayleigh” law and AA was calculated by Rayleigh in 1900, and it matches the observed spectrum at sufficiently low frequencies.

The total energy is thus

E(T)=(AkBTc3)∫0∞dωω2=∞\CE(T) = \left(\frac{A k_B T}{c^3}\right) \int_0^{\infty} d\omega \omega^2 = \infty

This is sometimes termed the “ultraviolet catastrophe”. What is needed is some new dimensionful scale, so that E(T,ω)\CE(T,\omega) can have a functional dependence on ω\omega that dies off quickly enough at large ω\omega that the integral is better behaved. One possibility is to just assume that the electromagnetic field has a largest possible frequency ωmax\omega_{max} and the integral is cut off, which amounts to multiplying E(T,ω)\CE(T,\omega) by θ(ωmax−ω)\theta(\omega_{max} - \omega).

However, the real issue is that the black body spectrum does not look like this. E(ω,T)\CE(\omega,T) is observable, for example by poking a small hole in the side of the black body and measuring the spectrum of teh emitted power (as illustrated above). It was already known that at high frequencies the black body spectrum follows the (phenomenological) Wien law E(ω,T)∼ω3e−cω/(kBT\CE(\omega,T) \sim \omega^3 e^{-c \omega/(k_B T}, where cc is some constant (which must have units of [energy ×\times time].)

Planck deduced a functional form that he later justifies with a hypothesis. We will cheat and start with that hypothesis. Let us fix the polarization state and the wavenumber k⃗{\vec k} of a mode of the electromagnetic field (ignoring boundary conditions and such, this is all very handwavy to make a point). We assume that each mode has an energy that is an integer multiple of ω\omega:

E(k⃗,n^pol,ω)=nℏω;  n={0,1,2,…}\CE({\vec k}, {\hat n}_{pol}, \omega) = n \hbar \omega ; \ \ n = \{0, 1,2,\ldots\}

where n^pol{\hat n}_{pol} denotes one of the two independent polarization states (linear, circular, etc). To get an expression appropriate for thermal equilibrium, we must do statistical mechanics. Boltzmann’s hypopthesis states that the probability of the mode being in the state labelled by nn is:

pn=e−nℏω/(kBT)∑m=0∞e−nℏω/(kBT)=e−nℏω/(kBT)(1−e−ℏω/(kBT))p_n = \frac{e^{-n \hbar \omega/(k_B T)}}{\sum_{m = 0}^{\infty} e^{-n \hbar \omega/(k_B T)}} = e^{-n\hbar\omega/(k_B T)}\left(1 - e^{-\hbar \omega/(k_B T)}\right)

where we have used the classic formula ∑N=0∞xN=11−x\sum_{N = 0}^{\infty} x^N = \frac{1}{1 - x} for x<1x < 1. To get the thermodynamic behavior, we need to compute the average energy per mode using the above probability distribution:

⟨En⟩=∑nEnpn=∑nnℏω(1−e−ℏω/(kBT))e−nℏω/(kBT)=(1−e−ℏω/(kBT))(kBT)2∂∂(kBT)∑ne−nβω/(kBT)=ℏωe−ℏω/(kBT)1−e−ℏω/(kBT)=ℏωeℏω/(kBT)−1\begin{align} \vev{\CE_n} & = \sum_n \CE_n p_n \\ & = \sum_n n \hbar \omega \left(1 - e^{-\hbar\omega/(k_B T)}\right) e^{-n\hbar\omega/(k_B T)} \\ & = \left(1 - e^{-\hbar\omega/(k_B T)}\right) (k_B T)^2 \frac{\del}{\del (k_B T)} \sum_n e^{-n\beta \omega/(k_B T)}\\ & = \frac{\hbar\omega e^{-\hbar\omega/(k_B T)}}{1 - e^{-\hbar\omega/(k_B T)}} \\ & = \frac{\hbar\omega}{e^{\hbar\omega/(k_B T)} - 1} \end{align}

If we carefully sum over all wavenumbers (with fixed frequency) and polarizations, we get:

E(ω,T)=ℏω3π2c31eℏω/(kBT)−1\CE(\omega, T) = \frac{\hbar \omega^3}{\pi^2 c^3} \frac{1}{e^{\hbar \omega/(k_B T)} - 1}

One can show that this takes the Rayleigh form for ℏω≪kBT\hbar\omega \ll k_B T and the Wien form for ℏω≫kBT\hbar\omega \gg k_B T. Finally, to tie this back to our earlier dimensional analysis, we can rewrite it as:

E(ω,T)=kBTω2c3f(ℏωkBT)\CE(\omega,T) = k_B T\frac{\omega^2}{c^3} f\left(\frac{\hbar\omega}{k_B T}\right)

where

f(x)=1π2xex−1\begin{align*} f(x) & = \frac{1}{\pi^2} \frac{x}{e^x - 1} \end{align*}

In other words, the fact that we can create a new energy scale ℏω\hbar\omega is what allows us to write a functional form for E(ω,T)\CE(\omega,T) consistent with the observed high- and low-frequency limits. Note that we now know the soure of this quantization (and we will say more about it below): light comes in “packets” consisting of individual particles known as photons with energy ℏω\hbar\omega.

CMBR spectrum from COBE and others

Image from Lawrence Berkeley Labs

This black body spectrum has been observed with exquisite prediction. As an example, a prediction of the hot big bang theory is that the early universe had a phase in which electrons, protons, and photons were in a state of thermal equilibrium with temperature T∼3000KT \sim 3000K. The photons would then have the black body spectrum at this temperature. At t∼105t \sim 10^5 years after the big bang, the electrons and protons combined to form neutral hydrogen. The photons retained the blackbody spectrum; however, as the universe expanded, the photons redshifted. The functional form of the blackboddy spectrum is known to be retained, with the temperature appearing as a parameter that decreases with redshift. Thus, we should see a “cosmic microwave background radiation” (CMBR) today, a blackbody spectrum parametyerized by a temperature T∼2.73KT \sim 2.73K. This has been observed to exquisite precision by the Cosmic Background Explorer (COBE) satellite which announced their results in 1992. (Different parts of the spectrum had been observed by earlier ground-based and balloon-borne experiments. This is a long and fascinating story with some interesting wrong turns. COBE showed that the spctrum was truly a blackbody spectrum from the Rayleigh through to the Wien range). There are small deviations from this, represented by local deviations in temperature, to a part in 105. This is consistent with the hot big bang theory, and in fact these small flucutuations are the seed for the cosmic structure we see today.

The photoelectric effect

The next phenomenon was famously discussed by Einstein during his annus mirabilis of 1905. Consider a beam of light with frequency ω\omega shining on an electrode comprised of some particular metal. A cathode can be reached by electons in the metal if they can travel across some voltage drop VV.

Hertz observed the following in 1887:

  1. The plates emit electrons (and no positvely charged particles)

  2. Whether the plate emits electrons depends only on the frequency of the incoming light, and not on its intensity.

  3. The magnitude of tte current in the cathode is proportional to the frequency

  4. The energy of each photoelectron (as measured by mmeasuring for what VV an electron hits the cathode) is independent of the intensity of the light, and is linear in frequency, above some critical frequency ωc\omega_c which depends on the metal (and not on the photon frequency).

Photelectric effect

Einstein’s interpretation was that light consisted of single particles or photons, each of which carries an energy ℏω\hbar \omega. Electrons acquire kinetic energy by absorbing a single photon. To be ejected from the metal, the electron must cross a potential barrier W=ℏωcW = \hbar \omega_c which is intrinsic to that metal (WW is called the work function of the metal). The kinetic energy for frequencies higher than ωc\omega_c is thus

12mv2=ℏ(ω−ωc)θ(ω−ωc)\half m v^2 = \hbar(\omega - \omega_c)\theta(\omega - \omega_c)

where θ(x)\theta(x) is the Heaviside step function. The point here is that the packets of energy Placnk suggested could be interpreted as a single quantum or particle of light called a photon whose energy is related to their frequency by the dimensionfuil constant ℏ\hbar. Increasing the intensity of light does not change the energy of each photon, but the number of photons. If there is not a single photon with enough energy to eject an electron, no electron will be ejected.

Photon polarization

Classical description

Consider an electromagnetic field propagating in a vacuum along the zz direction. The most general solution to Maxwell’s equations in this case is:

E⃗=Re {[Exx^+Eyy^]ei(kx−ωt)}=Re{E⃗cei(kx−ωt)}{\vec E} = \text{Re}\ \left\{ \left[ E_x {\hat x} + E_y {\hat y} \right]e^{i (kx - \omega t)}\right\} = \text{Re}\left\{ {\vec E}_c e^{i (kx - \omega t)} \right\}

where Ex,y∈CE_{x,y} \in \mathbb{C} so that Ei=∣Ei∣eiδiE_i = |E_i| e^{i\delta_i}, and ω=ck\omega = c k. We can rewrite this as

E⃗=∣Ex∣cos⁡(kz−ωt+δx)x^+∣Ey∣cos⁡(kx−ωt+δy){\vec E} = |E_x|\cos(kz - \omega t+ \delta_x) {\hat x} + |E_y| \cos(kx - \omega t + \delta_y)

The magnetic field H⃗{\vec H} can be determined from the Maxwell equation

∇×E⃗+1c∂∂tH⃗\nabla \times {\vec E} + \frac{1}{c}\frac{\del}{\del t}\vec{H}

This supports various polarization states such as

  1. Linear/plane polarization: E⃗x=Eeiδn^{\vec E}_x = E e^{i\delta} {\hat n} where EE is a positive real number and n^{\hat n} a unit vector in the x−yx-y plane. In particular n^=x^,y^{\hat n} = {\hat x}, {\hat y} correspond to plane polarization along the xx- and yy-axis respectively.

  2. Circular polarization:

E⃗LCP=E(cos⁡(kz−ωt+δ)x^+sin⁡(kz−ωt+δ)y^){\vec E}_{LCP} = E\left(\cos(kz - \omega t + \delta){\hat x} + \sin (kz - \omega t + \delta) {\hat y}\right)
E⃗RCP=E(cos⁡(kz−ωt+δ)x^+sin⁡(kz−ωt+δ)y^){\vec E}_{RCP} = E\left(\cos(kz - \omega t + \delta){\hat x} + \sin (kz - \omega t + \delta) {\hat y}\right)

Note that for any plane wave, E⃗c{\vec E}_c can be written as a linear combination of left- and right-circular polarizations, or as xx- and yy-plane polarizations.

The energy density is

E=18π[E⃗2+H⃗2]=14π[∣Ex∣2cos⁡2(kz−ωt+δx)+∣Ey∣2cos⁡2(kz−ωt+δy)]\begin{align} \cal{E} & = \frac{1}{8\pi}\left[ {\vec E}^2 + {\vec H}^2\right]\\ & = \frac{1}{4\pi} \left[|E_x|^2\cos^2(kz - \omega t + \delta_x) + |E_y|^2\cos^2(kz - \omega t + \delta_y)\right] \end{align}

Assuming the field propagates in some volume VV, the total energy is

U=∫d2xzE=V2×4π[∣Ex∣2+∣Ey∣2]U = \int d^2 x z {\cal E} = \frac{V}{2\times 4\pi} \left[|E_x|^2 + |E_y|^2\right]

the factor of 2 in the denominator comes from integrating the cos⁡2\cos^2 factors over zz.

Now let us pass our beam through a polarizer. We call an “x-polarizer” one which admits only light polarized along the x^{\hat x} direction, so that ExE_x remains unchanged by EyE_y is set to zero. In general, we can consider a polarizer aligned along any complex unit vector n^{\hat n} so that any initial light wave described by (11) becomes

E⃗after=Re{E⃗c⋅n^∗n^ei(kz−ωt)}{\vec E}_{after} = \text{Re}\left\{ {\vec E}_c \cdot {\hat n}^{\ast} {\hat n} e^{i(kz - \omega t)} \right\}

For a polarizer aligned along x^{\hat x}, n^=x^{\hat n} = {\hat x}; for a polarizer admitting LCP, n^LCP=12(x^−iy^){\hat n}_{LCP} = \frac{1}{\sqrt{2}} \left({\hat x} - i {\hat y}\right); for a polarizer admitting RCP, n^RCP=12(x^+iy^){\hat n}_{RCP} = \frac{1}{\sqrt{2}}\left({\hat x} + i {\hat y}\right).

This generally reduces the energy of the beam of light. For example, we can show that if we consider light polarized along n^=12(x^+y^){\hat n} = \frac{1}{\sqrt{2}} \left(\hat{x} + {\hat y}\right), that is at a 45∘45^{\circ} angle fromthe xx-axis, and pass it through a polatizer aligned along the xx axis,

E⃗=Ecos⁡(kx−ωt+δ)(x^+y^)→Ecos⁡(kx−ωt+δ)x^{\vec E} = E\cos(kx - \omega t + \delta)\left({\hat x} + {\hat y}\right) \rightarrow E\cos(kx - \omega t + \delta){\hat x}

The energy in this case is cut in half:

Uinit=V8π×2×∣E∣2→Uafter=V8π∣E∣2=12UinitU_{init} = \frac{V}{8\pi}\times 2 \times |E|^2 \rightarrow U_{after} = \frac{V}{8\pi}|E|^2 = \half U_{init}

Quantum description

If we apply Planck and Einstein’s insights, the beam of light consists of photons with energy U=NℏωU = N\hbar \omega. Consider the case above of light polarized at 45∘45^{\circ} from the xx-axis passing through a polarizer aligned along the xx axos. Then Ufinal=12NℏωU_{final} = \half N \hbar \omega and the number of photons is cut in half. But now we should be asking:

Our best available interpretation is that each photon has a 50%50\% chance of passing through this polarizer. For typical classical beams one can see that N≫1N \gg 1, so that the law of large numbers/central limit theorem tells us

N→N(12+O(1N))N \to N\left(\half + {\cal O}\left(\frac{1}{\sqrt{N}}\right)\right)

For more general polarizations,

E⃗=Re [Exei(kx−ωt)x^+Eyei(kx−ωt)y^]Uinitial=V8π(∣Ex∣2+∣Ey∣2)Ufinal=V8π∣Ex∣2\begin{align} {\vec E} = \text{Re}\ \left[ E_x e^{i(kx - \omega t)} {\hat x} + E_y e^{i(kx - \omega t)} {\hat y}\right]\\ U_{initial} & = \frac{V}{8\pi}\left(|E_x|^2 + |E_y|^2\right)\\ U_{final} & = \frac{V}{8\pi} |E_x|^2 \end{align}

So the probability of a single photon passing through the polarizer, given the polarization state listed above, is

p=NfinalNinitℏωℏω=UfinalUinit=∣Ex∣2∣Ex∣2+∣Ey∣2p = \frac{N_{final}}{N_{init}} \frac{\hbar\omega}{\hbar\omega} = \frac{U_{final}}{U_{init}} = \frac{|E_x|^2}{|E_x|^2 + |E_y|^2}

For more general polarization states, the probability of the photon passing through the polarizer aligned along the xx-axis is

p=∣Ex∣2∣Ex∣2+∣Ey∣2∼NfNp = \frac{|E_x|^2}{|E_x|^2 + |E_y|^2} \sim \frac{N_f}{N}

For a single photon with fixed frequency and wavenumber,

Utot=V8π(∣Ex∣2+∣Ey∣2)=ℏωU_{tot} = \frac{V}{8\pi} \left(|E_x|^2 + |E_y|^2\right) = \hbar \omega

For an electric field satisfying this relation, we can write a “state vector” that describes the photon polarization:

∣ψ⟩≡(ψxψy)ψi=V8πℏωEi\begin{align} & \ket{\psi} \equiv \begin{pmatrix} \psi_x \\ \psi_y \end{pmatrix}\\ & \psi_i = \sqrt{\frac{V}{8\pi \hbar \omega}} E_i \end{align}

The notation ∣ψ⟩\ket{\psi} should here be understood in this way, as shorthand for a vector (here a two-dimensional vector with complex entries). Using (25) we can see that ∣ψx∣2+∣ψy∣2=1|\psi_x|^2 + |\psi_y|^2 = 1, apnd that px=∣ψx∣2p_x = |\psi_x|^2. Similarly, the probability of the same photon passing through a polarizer aligned alongthe yy axis is py=∣ψy∣2p_y = |\psi_y|^2. It is as if the two possible outcomes are that the photon is polarized along the xx or yy axis. On the other hand, we could rotate both polarizers by (say) 17∘17^{\circ} and do the same experiment and the probability of the photon passing through each of the two perpendicular polarizers will sum to 1.

Finally, not that there is a linear structure to the space of photon polarizations, in that we can add two polarization vectors and get another polarization vector, up to the overall normalization constraint (25) for a single photon. In particular, if we define

∣ψx⟩=(10)∣ψy⟩=(01)\begin{align*} & \ket{\psi_x} = \begin{pmatrix} 1 \\ 0 \end{pmatrix}\\ & \ket{\psi_y} = \begin{pmatrix} 0 \\ 1 \end{pmatrix} \end{align*}

then any polarization can be described by the linear combination

α∣ψx⟩+β∣ψy⟩=(αβ)\alpha \ket{\psi_x} + \beta \ket{\psi_y} = \begin{pmatrix} \alpha \\ \beta \end{pmatrix}

with α,β∈C\alpha,\beta \in \mathbb{C}, and ∣α∣2+∣β∣2=1|\alpha|^2 + |\beta|^2 = 1. Similarly, we can consider the circular polarization states:

∣ψRCP⟩=12(1i)∣ψLCP⟩=12(1−i)\begin{align*} & \ket{\psi_{RCP}} = \frac{1}{\sqrt{2}} \begin{pmatrix} 1 \\ i \end{pmatrix}\\ & \ket{\psi_{LCP}} = \frac{1}{\sqrt{2}} \begin{pmatrix} 1 \\ -i \end{pmatrix} \end{align*}

You can show that any polarization state satisfying (25) can be written as

α~∣ψRCP⟩+β~∣ψLCP⟩=12(α~+β~i(α~−β~)){\tilde\alpha} \ket{\psi_{RCP}} + {\tilde\beta} \ket{\psi_{LCP}} = \frac{1}{\sqrt{2}}\begin{pmatrix} {\tilde \alpha} + {\tilde \beta} \\ i ({\tilde\alpha} - {\tilde\beta}) \end{pmatrix}

with α~,β~∈C{\tilde\alpha}, {\tilde\beta} \in \mathbb{C}, if ∣α~∣2+∣β~∣2=1|{\tilde\alpha}|^2 + |{\tilde\beta}|^2 = 1.

The essential point here is that the states of the photon are described by a 2-component vector; these vectors can be added to describe other polarization states; and to get the probability of a given experiment yielding a specific outcome in the form of a specific polarization, one takes the absolute value squared of the projection of the vector along that direction.

This hints at a very general structure for describing physical states and teh reults of mesurement in quantum mechanics. To set this up we need to lay down the correct mathematical language for describing such vectors, namely linear algebra.