Orthogonal Sets, Orthonormal Bases, and the Gram-Schmidt Process – MATH 416, Lecture 31 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Inner product spaces, linear independence, span, basis, norms


Big Picture

This is the first half of a lecture that builds the machinery for working with "nice" bases in inner product spaces. In earlier lectures you learned that any finite-dimensional vector space has a basis, and that inner products let you measure angles and lengths. Here you learn that you can always upgrade an arbitrary basis into one where every pair of vectors is perpendicular (orthogonal) or even perpendicular and unit-length (orthonormal). The tool that does this is the Gram-Schmidt process. Once you have an orthonormal basis, computing coordinates becomes trivially easy, and that payoff carries through the rest of the course.


TL;DR

An orthogonal set of nonzero vectors is automatically linearly independent, and coordinates in such a set are given by simple inner-product formulas. The Gram-Schmidt process takes any linearly independent set and produces an orthogonal set that spans the same subspace. Normalising gives an orthonormal basis. This works in finite and countably infinite dimensions.


Key Terms

Orthogonal set

A set S = {v₁, …, vₖ} in an inner product space (V, ⟨·,·⟩) where ⟨vᵢ, vⱼ⟩ = 0 for every i ≠ j. In simple terms, every pair of distinct vectors in the set is perpendicular.

Orthonormal set

An orthogonal set where each vector also has unit norm: ‖vᵢ‖ = 1 for all i. Think of it as an orthogonal set that has been "rescaled" so every vector has length 1.

Orthonormal basis

An orthonormal set that is also a basis for the space (or subspace) in question. In simple terms, it is both a complete spanning set and a set of mutually perpendicular unit vectors.

Gram-Schmidt process

An algorithm that takes a linearly independent set {w₁, …, wₖ} and outputs an orthogonal set {v₁, …, vₖ} spanning the same subspace. Think of it as a systematic way to "peel off" the non-perpendicular parts, one vector at a time.

Legendre polynomials

The orthonormal polynomials you get when you apply Gram-Schmidt to {1, x, x², …} in the inner product space C⁰([−1, 1]) with ⟨f, g⟩ = ∫₋₁¹ f(x)g(x) dx. They are a classical family of orthogonal polynomials that appear across applied mathematics and physics.


Core Content

Properties of Orthogonal Sets

  • If S = {v₁, …, vₖ} is orthogonal and every vᵢ ≠ 0, then S is linearly independent.

    • The proof uses the inner product to isolate each coefficient in a supposed linear dependence relation; orthogonality forces every coefficient to zero.

  • If y ∈ Span(S) and S is orthogonal with all nonzero vectors, the coordinates of y are given by a clean formula:

    y = (⟨y, v₁⟩ / ‖v₁‖²) v₁ + ⋯ + (⟨y, vₖ⟩ / ‖vₖ‖²) vₖ

    • No row reduction or matrix inversion needed. Each coefficient is computed independently via a single inner product divided by a squared norm.

    • When the set is orthonormal (‖vᵢ‖ = 1 for all i), this simplifies further to y = ⟨y, v₁⟩v₁ + ⋯ + ⟨y, vₖ⟩vₖ.

The Gram-Schmidt Process – Statement

Theorem (Gram-Schmidt). If {w₁, …, wₖ} is a linearly independent set in an inner product space (V, ⟨·,·⟩), then there exists an orthogonal set {v₁, …, vₖ} ⊂ V such that

Span{v₁, …, vₖ} = Span{w₁, …, wₖ}.

Corollary. Every finite-dimensional inner product space (Vⁿ, ⟨·,·⟩) has an orthogonal (and therefore an orthonormal) basis.

The Gram-Schmidt Process – Construction

The algorithm is iterative. At each step you subtract off the components along the previously constructed orthogonal vectors.

  • Step 1. Set v₁ = w₁.

  • Step 2. Set v₂ = w₂ − (⟨w₂, v₁⟩ / ‖v₁‖²) v₁.

    • You are removing the projection of w₂ onto v₁, leaving only the component perpendicular to v₁.

  • Step 3. Set v₃ = w₃ − (⟨w₃, v₁⟩ / ‖v₁‖²) v₁ − (⟨w₃, v₂⟩ / ‖v₂‖²) v₂.

  • General step. For step k:

    vₖ = wₖ − Σⱼ₌₁ᵏ⁻¹ (⟨wₖ, vⱼ⟩ / ‖vⱼ‖²) vⱼ

The Gram-Schmidt Process – Verification

The proof proceeds by induction on k.

  • Base case (k = 1). v₁ = w₁, so there is nothing to check.

  • Inductive step. Assume {v₁, …, vₖ₋₁} is orthogonal. For any j ∈ {1, …, k−1}:

    ⟨vₖ, vⱼ⟩ = ⟨wₖ − Σᵢ₌₁ᵏ⁻¹ (⟨wₖ, vᵢ⟩ / ‖vᵢ‖²) vᵢ, vⱼ⟩

    Expanding by linearity and applying the inductive hypothesis (⟨vᵢ, vⱼ⟩ = 0 for i ≠ j, and ⟨vⱼ, vⱼ⟩ = ‖vⱼ‖²):

    = ⟨wₖ, vⱼ⟩ − (⟨wₖ, vⱼ⟩ / ‖vⱼ‖²)·‖vⱼ‖² = 0.

  • Span equality. By construction each vᵢ is a linear combination of w₁, …, wᵢ, so Span{v₁, …, vₖ} ⊆ Span{w₁, …, wₖ}. Since {v₁, …, vₖ} is orthogonal (hence linearly independent) and has the same number of vectors, the two spans have the same dimension and must be equal.

From Orthogonal to Orthonormal

Once you have an orthogonal set {v₁, …, vₖ}, normalise each vector:

{v₁/‖v₁‖, …, vₖ/‖vₖ‖}

This gives an orthonormal set spanning the same subspace.

Remark on Infinite Sets

Gram-Schmidt can be applied to countably infinite linearly independent sets {w₁, w₂, w₃, …} as well. The process runs the same way, one step at a time, producing a countably infinite orthogonal (or orthonormal) sequence.

Example – Legendre Polynomials

Consider P(ℝ) ⊂ C⁰([−1, 1]) with the inner product ⟨f, g⟩ = ∫₋₁¹ f(x)g(x) dx.

Starting with the standard monomial basis {1, x, x², x³, x⁴, …} and applying Gram-Schmidt:

  • Orthogonal output: {1, x, x² − 1/3, x³ − (3/5)x, x⁴ − (6/7)x² + 3/35, …}

  • Orthonormal output: {1/√2, √(3/2) x, √(5/8)(x² − 1/3), …}

These normalised polynomials are (scalar multiples of) the Legendre polynomials, which appear throughout mathematical physics, numerical integration (Gauss-Legendre quadrature), and approximation theory.


Formulas and Diagrams

Coordinate formula (orthogonal basis):

y = Σᵢ₌₁ᵏ (⟨y, vᵢ⟩ / ‖vᵢ‖²) vᵢ

Gram-Schmidt recursion:

v₁ = w₁

vₖ = wₖ − Σⱼ₌₁ᵏ⁻¹ (⟨wₖ, vⱼ⟩ / ‖vⱼ‖²) vⱼ, for k ≥ 2

Normalisation:

uᵢ = vᵢ / ‖vᵢ‖


Real-World Applications

Gram-Schmidt underpins the QR factorisation used in numerical linear algebra to solve least-squares problems, which is how GPS receivers triangulate your position and how regression models fit data. The Legendre polynomials produced in the example are the basis of Gauss-Legendre quadrature, a standard method for high-accuracy numerical integration in engineering and scientific computing.


Common Misconceptions

  • Students sometimes believe Gram-Schmidt changes the span. It does not. The output set spans exactly the same subspace as the input set; only the directions of the individual vectors change.

  • A common error is forgetting to divide by ‖vⱼ‖² in the projection terms and writing ⟨wₖ, vⱼ⟩vⱼ instead of (⟨wₖ, vⱼ⟩ / ‖vⱼ‖²)vⱼ. This only works when the vⱼ are already unit vectors.

  • Students often confuse "orthogonal" (perpendicular) with "orthonormal" (perpendicular and unit-length). Gram-Schmidt produces an orthogonal set; a separate normalisation step is needed for orthonormality.

  • The coordinate formula y = Σ(⟨y, vᵢ⟩ / ‖vᵢ‖²)vᵢ requires y to actually be in Span(S). If y is not in the span, this formula gives the projection of y, not y itself. That distinction becomes critical in Part 2.


Why It Matters / Exam Flags

⚠️ The Gram-Schmidt formula is a near-certainty on exams. Be able to carry it out by hand for 2 or 3 vectors in ℝⁿ.

⚠️ Know the coordinate formula for orthogonal and orthonormal bases cold. It is the payoff for all this machinery and appears in projection and least-squares problems.

⚠️ Be ready to explain why an orthogonal set of nonzero vectors is linearly independent. This is a common short-proof exam question.

⚠️ The Legendre polynomial example connects Gram-Schmidt to function spaces. If your course covers approximation or Fourier-like expansions, expect a question that asks you to orthogonalise low-degree monomials.


Quick Self-Test

True or False: An orthogonal set can contain the zero vector and still be linearly independent.

A: False. If any vector in the set is zero, the set is linearly dependent.

Fill in the blank: In the Gram-Schmidt process, v₂ = w₂ − ______ · v₁.

A: ⟨w₂, v₁⟩ / ‖v₁‖²

True or False: Gram-Schmidt can only be applied to finite sets of vectors.

A: False. It extends to countably infinite linearly independent sets.


Practice Q&A

Q: Let {w₁, w₂} be linearly independent in an inner product space. Write out the Gram-Schmidt procedure to produce an orthogonal set {v₁, v₂}.

A: Set v₁ = w₁. Then v₂ = w₂ − (⟨w₂, v₁⟩ / ‖v₁‖²) v₁. This ensures ⟨v₁, v₂⟩ = 0 and Span{v₁, v₂} = Span{w₁, w₂}.

Q: Why does the coordinate formula y = Σ (⟨y, vᵢ⟩ / ‖vᵢ‖²) vᵢ not require solving a system of equations?

A: Because orthogonality decouples the coordinates. Taking ⟨y, vⱼ⟩ isolates the j-th coefficient since all cross terms ⟨vᵢ, vⱼ⟩ vanish for i ≠ j.

Q: Apply Gram-Schmidt to {1, x} in C⁰([−1,1]) with ⟨f,g⟩ = ∫₋₁¹ f(x)g(x) dx.

A: v₁ = 1. Then ⟨x, 1⟩ = ∫₋₁¹ x dx = 0 and ‖1‖² = ∫₋₁¹ 1 dx = 2, so v₂ = x − (0/2)·1 = x. The monomials 1 and x are already orthogonal on [−1, 1].

Q: Suppose {v₁, v₂, v₃} is orthogonal in ℝ³ with all nonzero vectors. Must it be a basis for ℝ³?

A: Yes. An orthogonal set of nonzero vectors is linearly independent, and three linearly independent vectors in a three-dimensional space form a basis.


Connections to Other Topics

This material connects directly to orthogonal projections and the best approximation theorem (covered in Part 2 of these notes), which rely on having an orthonormal basis for a subspace. It also connects to QR factorisation in matrix algebra, where Gram-Schmidt is applied to the columns of a matrix. Later in the course, the same ideas extend to Fourier series, where you orthogonalise trigonometric functions rather than polynomials.


Related Terms / Search Tags

orthogonal set, orthonormal set, orthogonal basis, orthonormal basis, Gram-Schmidt process, Gram-Schmidt orthogonalisation, inner product space, coordinate formula orthogonal basis, Legendre polynomials, QR factorisation, orthogonal vectors, perpendicular vectors, normalisation, linear independence from orthogonality, MATH 416, abstract linear algebra, UIUC