Source: Friedberg, Insel, Spence -- Linear Algebra 4th ed., Ch. 6.1-6.2
Tags: inner product, norm, Cauchy-Schwarz inequality, triangle inequality, orthogonal, orthonormal, Gram-Schmidt, orthogonal complement, Fourier coefficients, projection, Bessel's inequality, MATH 301, UIUC, abstract linear algebra
Difficulty: Intermediate-Advanced Prerequisites: Vector spaces over R and C, bases and dimension, linear transformations, eigenvalues and eigenvectors (Chapters 1-5).
This material introduces the idea of "geometry" into abstract vector spaces. Up to this point in the course, you have worked with vector spaces that have no built-in concept of length, angle, or distance. Inner product spaces fix that: they equip a vector space with a function that generalises the dot product from R^n, and from it you get norms (lengths), orthogonality (perpendicularity), and projections (closest points). These two sections lay the foundation for everything that follows in Chapter 6, including the Spectral Theorem. If you are comfortable with the dot product in R^n, the abstraction here is the same idea lifted to arbitrary vector spaces over R or C.
An inner product is a generalisation of the dot product to any vector space over R or C, letting you define length (norm), angle, and orthogonality. The Gram-Schmidt process takes any basis and produces an orthonormal one, which makes almost every computation in inner product spaces drastically simpler, from finding coordinates to computing projections.
Inner product
A function ⟨x, y⟩ from V × V to F (where F = R or C) satisfying four axioms: additivity in the first argument, homogeneity in the first argument, conjugate symmetry, and positive definiteness. In simple terms, it is a machine that eats two vectors and outputs a scalar, generalising the dot product.
Standard inner product on F^n
For x = (a₁, ..., aₙ) and y = (b₁, ..., bₙ) in F^n, ⟨x, y⟩ = Σ aᵢ b̄ᵢ. Over R, this is just the dot product x · y. Think of it as the default inner product you already know, extended to allow complex entries.
Conjugate transpose (adjoint) of a matrix, A*
The n × m matrix obtained by transposing A and taking the complex conjugate of every entry: (A*)ᵢⱼ = Āⱼᵢ. For real matrices, A* is just the transpose Aᵗ. In simple terms, flip the matrix across the diagonal and conjugate each entry.
Frobenius inner product
For n × n matrices A, B: ⟨A, B⟩ = tr(B*A). This turns the space of matrices into an inner product space. Think of it as treating the matrix entries as a vector and taking the standard inner product.
Inner product space
A vector space V over F equipped with a specific inner product. The same vector space can become different inner product spaces under different inner products.
Norm (length)
For x in an inner product space, ‖x‖ = √⟨x, x⟩. Generalises Euclidean length.
Orthogonal (perpendicular)
Two vectors x, y are orthogonal if ⟨x, y⟩ = 0.
Orthogonal set
A set of vectors in which every pair is orthogonal.
Orthonormal set
An orthogonal set in which every vector is a unit vector (‖v‖ = 1). Equivalently, ⟨vᵢ, vⱼ⟩ = δᵢⱼ (the Kronecker delta).
Normalizing a vector
Replacing a nonzero vector x with (1/‖x‖)x, giving a unit vector in the same direction.
Orthonormal basis
An ordered basis that is also an orthonormal set. The standard basis for F^n is the prototypical example.
Fourier coefficients
Given an orthonormal set β and a vector x, the scalars ⟨x, y⟩ for each y ∈ β are the Fourier coefficients of x relative to β.
Gram-Schmidt process
An algorithm that takes a linearly independent set {w₁, ..., wₙ} and produces an orthogonal set {v₁, ..., vₙ} spanning the same subspace. Normalising the output gives an orthonormal set.
Orthogonal complement, S⊥
The set of all vectors in V orthogonal to every vector in S. It is always a subspace.
Orthogonal projection of y on W
The unique vector u in a finite-dimensional subspace W that is closest to y. Computed as u = Σ ⟨y, vᵢ⟩ vᵢ where {v₁, ..., vₖ} is an orthonormal basis for W.
The four axioms for an inner product ⟨ · , · ⟩ on a vector space V over F are:
(a) ⟨x + z, y⟩ = ⟨x, y⟩ + ⟨z, y⟩ (additivity in the first component)
(b) ⟨cx, y⟩ = c⟨x, y⟩ (homogeneity in the first component)
(c) ⟨x, y⟩ = conjugate of ⟨y, x⟩ (conjugate symmetry; reduces to symmetry over R)
(d) ⟨x, x⟩ > 0 whenever x ≠ 0 (positive definiteness)
From these axioms, Theorem 6.1 gives important consequences:
The inner product is conjugate-linear in the second component: ⟨x, cy⟩ = c̄ ⟨x, y⟩
⟨x, 0⟩ = ⟨0, x⟩ = 0
⟨x, x⟩ = 0 if and only if x = 0
If ⟨x, y⟩ = ⟨x, z⟩ for all x in V, then y = z
Standard inner product on F^n: ⟨x, y⟩ = Σ aᵢ b̄ᵢ
Function space C([0,1]): ⟨f, g⟩ = ∫₀¹ f(t)g(t) dt
Frobenius inner product on Mₙₓₙ(F): ⟨A, B⟩ = tr(B*A)
The space H of continuous complex-valued functions on [0, 2π]: ⟨f, g⟩ = (1/2π) ∫₀²π f(t) ḡ(t) dt
Theorem 6.2 establishes:
‖cx‖ = |c| · ‖x‖
‖x‖ = 0 iff x = 0
Cauchy-Schwarz inequality: |⟨x, y⟩| ≤ ‖x‖ · ‖y‖
Triangle inequality: ‖x + y‖ ≤ ‖x‖ + ‖y‖
The Cauchy-Schwarz proof works by setting c = ⟨x, y⟩/⟨y, y⟩ in the expansion of ‖x - cy‖² ≥ 0. The Triangle inequality follows from Cauchy-Schwarz.
Equality holds in Cauchy-Schwarz if and only if one vector is a scalar multiple of the other.
Other useful identities:
Pythagorean theorem: if ⟨x, y⟩ = 0, then ‖x + y‖² = ‖x‖² + ‖y‖²
Parallelogram law: ‖x + y‖² + ‖x - y‖² = 2‖x‖² + 2‖y‖²
An orthogonal set of nonzero vectors is always linearly independent (Corollary 2 to Theorem 6.3). This is a very clean way to verify independence.
If S = {v₁, ..., vₖ} is orthogonal with nonzero vectors and y ∈ span(S), then:
y = Σ (⟨y, vᵢ⟩ / ‖vᵢ‖²) vᵢ
If S is orthonormal, this simplifies to:
y = Σ ⟨y, vᵢ⟩ vᵢ
This is why orthonormal bases are so useful: the coefficients of any vector are just inner products, no matrix inversion required.
Given a linearly independent set {w₁, ..., wₙ}:
v₁ = w₁
For k ≥ 2: vₖ = wₖ - Σⱼ₌₁ᵏ⁻¹ (⟨wₖ, vⱼ⟩ / ‖vⱼ‖²) vⱼ
The result {v₁, ..., vₙ} is an orthogonal set with span(S') = span(S). Normalise each vₖ to get an orthonormal set.
Every nonzero finite-dimensional inner product space has an orthonormal basis (Theorem 6.5). You get one by applying Gram-Schmidt to any basis and normalising.
For any subset S of V, S⊥ is a subspace.
Theorem 6.6 (the projection theorem): Let W be a finite-dimensional subspace of V and y ∈ V. Then there exist unique vectors u ∈ W and z ∈ W⊥ such that y = u + z. If {v₁, ..., vₖ} is an orthonormal basis for W:
u = Σᵢ₌₁ᵏ ⟨y, vᵢ⟩ vᵢ
The vector u is the orthogonal projection of y on W, and it is the unique closest vector to y in W: ‖y - x‖ ≥ ‖y - u‖ for all x ∈ W, with equality only when x = u.
Theorem 6.7 gives the dimension formula: dim(V) = dim(W) + dim(W⊥).
Bessel's inequality: if S = {v₁, ..., vₙ} is orthonormal in V, then for any x ∈ V:
‖x‖² ≥ Σ |⟨x, vᵢ⟩|²
Equality holds if and only if x ∈ span(S).
Parseval's identity (finite-dimensional): if {v₁, ..., vₙ} is an orthonormal basis for V, then for any x, y ∈ V:
⟨x, y⟩ = Σ ⟨x, vᵢ⟩ · conjugate(⟨y, vᵢ⟩)
Standard inner product on F^n: ⟨x, y⟩ = Σᵢ₌₁ⁿ aᵢ b̄ᵢ
Norm: ‖x‖ = √⟨x, x⟩
Gram-Schmidt formula: vₖ = wₖ - Σⱼ₌₁ᵏ⁻¹ (⟨wₖ, vⱼ⟩ / ‖vⱼ‖²) vⱼ
Orthogonal projection of y onto W (with orthonormal basis {v₁, ..., vₖ}): u = Σᵢ₌₁ᵏ ⟨y, vᵢ⟩ vᵢ
Orthogonal projection is the mathematical backbone of least squares approximation, which is used everywhere from fitting curves to experimental data to training machine learning models. The Fourier coefficients computed via inner products on function spaces are the basis of signal processing, audio compression (MP3), and image compression (JPEG).
Students often assume the inner product is linear in both components. It is not: it is linear in the first and conjugate-linear in the second (over C). This means ⟨x, cy⟩ = c̄ ⟨x, y⟩, not c⟨x, y⟩.
The Gram-Schmidt process requires a linearly independent set as input, not just any set of vectors. Applying it to a linearly dependent set will produce a zero vector at some step.
Orthogonality depends on the choice of inner product. Two vectors orthogonal under one inner product may not be orthogonal under another, even on the same vector space.
The projection formula u = Σ ⟨y, vᵢ⟩ vᵢ requires the basis to be orthonormal. Using a non-orthonormal basis in this formula gives a wrong answer.
⚠️ You will almost certainly be asked to carry out the Gram-Schmidt process on a specific set of vectors. Practise the computation until it is mechanical.
⚠️ Know the Cauchy-Schwarz and Triangle inequalities, including the conditions for equality.
⚠️ Be able to compute orthogonal projections and the distance from a point to a subspace. This is a very common exam question.
⚠️ Know the four inner product axioms and be able to verify or disprove that a given function is an inner product (check each axiom, especially positive definiteness).
⚠️ The dimension formula dim(V) = dim(W) + dim(W⊥) is frequently tested in short-answer or true/false questions.
True or False: An inner product is linear in both of its arguments. False. It is linear in the first and conjugate-linear in the second (over C).
True or False: Every orthogonal set is linearly independent. False. The set must consist of nonzero vectors. An orthogonal set containing the zero vector is not linearly independent.
True or False: Every nonzero finite-dimensional inner product space has an orthonormal basis. True.
Fill in the blank: The Cauchy-Schwarz inequality states |⟨x, y⟩| ≤ ____. ‖x‖ · ‖y‖
Fill in the blank: The orthogonal projection of y onto a subspace W is the unique vector in W that ____ y. is closest to
Q: Let x = (1, 2, 3) and y = (3, -1, 2) in R³ with the standard inner product. Compute ⟨x, y⟩, ‖x‖, and verify the Cauchy-Schwarz inequality.
A: ⟨x, y⟩ = 3 - 2 + 6 = 7. ‖x‖ = √(1+4+9) = √14. ‖y‖ = √(9+1+4) = √14. Check: |7| = 7 ≤ √14 · √14 = 14. ✓
Q: Apply the Gram-Schmidt process to {w₁ = (1, 0, 1, 0), w₂ = (1, 1, 1, 1)} in R⁴ with the standard inner product.
A: v₁ = (1, 0, 1, 0). Then ⟨w₂, v₁⟩ = 1+0+1+0 = 2 and ‖v₁‖² = 2. So v₂ = (1,1,1,1) - (2/2)(1,0,1,0) = (0, 1, 0, 1). Normalising: u₁ = (1/√2)(1,0,1,0), u₂ = (1/√2)(0,1,0,1).
Q: Let W = span{e₁, e₂} in R³. Find W⊥ and the orthogonal projection of y = (3, 4, 5) onto W.
A: W⊥ = span{e₃} = {(0, 0, c) : c ∈ R}. Projection of y onto W: u = ⟨y, e₁⟩e₁ + ⟨y, e₂⟩e₂ = 3e₁ + 4e₂ = (3, 4, 0). Distance from y to W: ‖y - u‖ = ‖(0, 0, 5)‖ = 5.
Q: Why can two different inner products on the same vector space give different orthogonal complements?
A: Orthogonality is defined by ⟨x, y⟩ = 0, so changing the inner product changes which pairs of vectors satisfy this condition. The same subspace W can have different orthogonal complements under different inner products because the notion of "perpendicular" itself has changed.
Q: State Bessel's inequality and the condition for equality.
A: If {v₁, ..., vₙ} is an orthonormal set in V and x ∈ V, then ‖x‖² ≥ Σ |⟨x, vᵢ⟩|². Equality holds if and only if x ∈ span{v₁, ..., vₙ}.
This material connects directly to the Spectral Theorem (Section 6.6), which decomposes a normal or self-adjoint operator into a sum of orthogonal projections, each computed using the tools from these two sections. The Gram-Schmidt process also underlies the QR factorisation used in numerical linear algebra. The Fourier coefficients introduced here generalise to the full Fourier series and Fourier transform studied in analysis and signal processing courses.
inner product space, dot product, standard inner product, Frobenius inner product, Hilbert space, norm, length, Cauchy-Schwarz, Schwarz inequality, Cauchy inequality, triangle inequality, orthogonal vectors, perpendicular, orthonormal basis, Gram-Schmidt orthogonalisation, Gram-Schmidt orthogonalization, orthogonal complement, S perp, orthogonal projection, best approximation, closest vector, Fourier coefficients, Bessel inequality, Parseval identity, conjugate transpose, adjoint matrix, parallelogram law, Pythagorean theorem inner product, normalising, unit vector, Kronecker delta, linear algebra final exam, MATH 301, UIUC