Source: Friedberg, Insel, Spence – Linear Algebra, 4th Ed.
Tags: T-invariant subspace, cyclic subspace, Cayley-Hamilton theorem, restriction of operator, characteristic polynomial, f(T), direct sum of invariant subspaces
Difficulty: Advanced | Prerequisites: Sections 5.1 and 5.2 (eigenvalues, diagonalizability), determinants, null space, range.
This section introduces two ideas that go beyond diagonalization. First, T-invariant subspaces: subspaces W where T maps W back into itself, letting you study T "restricted to W" as a smaller, simpler operator. Second, the Cayley-Hamilton theorem, which says every operator satisfies its own characteristic equation: if f(t) is the characteristic polynomial of T, then f(T) = 0 (the zero transformation). This result is a bridge to Chapter 7, where you study non-diagonalizable operators via the Jordan form. You should be comfortable with eigenspaces, characteristic polynomials, and (optionally) direct sums from Section 5.2.
A T-invariant subspace is one that T maps into itself. The T-cyclic subspace generated by a vector v is the smallest such subspace containing v, and its characteristic polynomial divides that of T. The Cayley-Hamilton theorem follows: plug T into its own characteristic polynomial and you get the zero operator.
T-invariant subspace
A subspace W of V such that T(W) ⊆ W, meaning T(v) ∈ W for every v ∈ W.
Think of it as: a subspace that T "cannot escape." Once you are in W, applying T keeps you in W.
T-cyclic subspace generated by x
The subspace W = span({x, T(x), T^2(x), ...}). It is the smallest T-invariant subspace of V containing x.
In simple terms, start with a vector x and keep applying T. Collect everything you can reach: that is the cyclic subspace. It always has a neat basis {x, T(x), ..., T^{k-1}(x)}.
Restriction T_W
If W is a T-invariant subspace, then T_W is the linear operator on W defined by T_W(v) = T(v) for v ∈ W. Because W is T-invariant, the output T(v) stays in W, so T_W is well-defined.
Think of it as: zooming in on T's behaviour within the smaller space W, ignoring everything outside.
Cayley-Hamilton theorem
If f(t) is the characteristic polynomial of a linear operator T, then f(T) = T_0 (the zero transformation). Equivalently, for a matrix A, f(A) = O (the zero matrix).
In simple terms, every square matrix (or linear operator) satisfies its own characteristic equation. This is a deep structural fact that constrains what T can do.
Five subspaces of V are always T-invariant:
{0} (trivially)
V itself
R(T), the range of T
N(T), the null space of T
E_λ, the eigenspace for any eigenvalue λ
These follow directly from the definitions.
If W is a T-invariant subspace, extend a basis for W to a basis for V. The matrix of T in this basis has block upper triangular form:
[T]_β = [[B_1, B_2],[O, B_3]]
where B_1 = [T_W]_γ. The characteristic polynomial of T_W divides the characteristic polynomial of T.
This is powerful: studying T on a smaller invariant subspace gives partial information about T's characteristic polynomial.
Let W be the T-cyclic subspace generated by nonzero v, with dim(W) = k. Then:
(a) {v, T(v), T^2(v), ..., T^{k-1}(v)} is a basis for W. (You stop at k-1 because T^k(v) becomes linearly dependent on the earlier vectors.)
(b) If a_0 v + a_1 T(v) + ... + a_{k-1} T^{k-1}(v) + T^k(v) = 0, then the characteristic polynomial of T_W is f(t) = (-1)^k (a_0 + a_1 t + ... + a_{k-1} t^{k-1} + t^k).
The matrix of T_W in this basis has the companion form:
[[0, 0, ..., 0, -a_0], [1, 0, ..., 0, -a_1], [0, 1, ..., 0, -a_2], [...], [0, 0, ..., 1, -a_{k-1}]]
Statement: If f(t) is the characteristic polynomial of T, then f(T) = T_0.
Proof outline: For any nonzero v ∈ V, let W be the T-cyclic subspace generated by v, with characteristic polynomial g(t) for T_W. By construction, g(T)(v) = 0. By Theorem 5.21, g(t) divides f(t), so f(t) = q(t)g(t). Then f(T)(v) = q(T)(g(T)(v)) = q(T)(0) = 0. Since this holds for every v, f(T) = T_0.
Corollary for matrices: If f(t) is the characteristic polynomial of A, then f(A) = O (the zero matrix).
For T(a, b) = (a + 2b, -2a + b) on R^2 with matrix A = [[1, 2],[-2, 1]]:
Characteristic polynomial: f(t) = t^2 - 2t + 5.
Check: A^2 - 2A + 5I = [[-3, 4],[-4, -3]] + [[-2, -4],[4, -2]] + [[5, 0],[0, 5]] = [[0, 0],[0, 0]]. Confirmed.
If V = W_1 ⊕ W_2 ⊕ ... ⊕ W_k where each W_i is T-invariant, and f_i(t) is the characteristic polynomial of T_{W_i}, then the characteristic polynomial of T is f_1(t) · f_2(t) · ... · f_k(t).
This decomposes the characteristic polynomial into factors associated with invariant subspaces, which is the starting point for the theory in Chapter 7.
T is diagonalizable if and only if V can be written as a direct sum of one-dimensional T-invariant subspaces (each spanned by an eigenvector).
Block form of [T]_β when W is T-invariant:
[T]_β = [[B_1, B_2],[O, B_3]]
where B_1 = [T_W]_γ.
Cayley-Hamilton: f(A) = O, where f(t) = det(A - tI).
Characteristic polynomial factoring over direct sums: f(t) = f_1(t) · f_2(t) · ... · f_k(t).
The Cayley-Hamilton theorem is used in control theory to express high powers of a system matrix as a linear combination of lower powers (since A^n can be written in terms of A^0, A^1, ..., A^{n-1}). This is the basis of the "minimal polynomial" approach to computing matrix functions. Invariant subspaces underpin modal analysis in mechanical engineering, where each vibration mode corresponds to an invariant subspace of the system operator.
"The Cayley-Hamilton theorem says det(A - AI) = 0." This is a common symbolic confusion. The theorem says: compute f(t) = det(A - tI) as a polynomial in t, then substitute the matrix A for t. You are not substituting A into the determinant formula; you are evaluating the polynomial f at the matrix A.
"Only diagonalizable operators satisfy their characteristic polynomial." Every operator satisfies its characteristic polynomial, diagonalizable or not. The Cayley-Hamilton theorem has no diagonalizability assumption.
"A T-invariant subspace must be an eigenspace." Eigenspaces are T-invariant, but not all T-invariant subspaces are eigenspaces. The range R(T), null space N(T), cyclic subspaces, and V itself are all T-invariant without necessarily being eigenspaces.
"The T-cyclic subspace generated by v is always all of V." It can be strictly smaller. In Example 3, the T-cyclic subspace generated by e_1 is the xy-plane, not all of R^3.
⚠️ The Cayley-Hamilton theorem is a classic "state and verify" exam problem. Be able to compute f(t) and then confirm f(A) = O for a given matrix.
⚠️ Know the five standard T-invariant subspaces ({0}, V, R(T), N(T), E_λ) and be ready to prove any of them is invariant.
⚠️ T-cyclic subspaces and their basis {v, T(v), ..., T^{k-1}(v)} may appear in exam problems. Practise computing them for small examples.
⚠️ The block triangular form [[B_1, B_2],[O, B_3]] and the fact that the characteristic polynomial of T_W divides that of T is a structural result that exam questions may test conceptually.
⚠️ The Cayley-Hamilton theorem connects to Exercise 23 of Section 5.1: for diagonalizable T, f(T) = T_0 was provable directly. Theorem 5.23 generalises this to all operators.
True or false: N(T) is always a T-invariant subspace. A: True. If v ∈ N(T), then T(v) = 0 ∈ N(T).
True or false: The Cayley-Hamilton theorem requires T to be diagonalizable. A: False. It holds for all linear operators on finite-dimensional spaces.
Fill in the blank: The T-cyclic subspace generated by v is the ______ T-invariant subspace containing v. A: Smallest.
True or false: If f(t) = t^3 - 2t^2 + t is the characteristic polynomial of A, then A^3 - 2A^2 + A = O. A: True, by the Cayley-Hamilton theorem.
Q: Let T(a, b, c) = (a + b, b + c, 0) on R^3. Show that the xy-plane W = {(x, y, 0)} is T-invariant.
A: Take any (a, b, 0) ∈ W. Then T(a, b, 0) = (a + b, b + 0, 0) = (a + b, b, 0), which is in W. So T(W) ⊆ W.
Q: Let T(a, b, c) = (-b + c, a + c, 3c) on R^3. Find the T-cyclic subspace generated by e_1.
A: T(e_1) = (0, 1, 0) = e_2. T^2(e_1) = T(e_2) = (-1, 0, 0) = -e_1. Since T^2(e_1) is in span({e_1, e_2}), the cyclic subspace is span({e_1, e_2}), the xy-plane. Dimension 2, with basis {e_1, e_2}.
Q: Verify the Cayley-Hamilton theorem for A = [[1, 2],[-2, 1]].
A: f(t) = det(A - tI) = (1-t)^2 + 4 = t^2 - 2t + 5. Compute: A^2 = [[-3, 4],[-4, -3]]. Then A^2 - 2A + 5I = [[-3, 4],[-4, -3]] + [[-2, -4],[4, -2]] + [[5, 0],[0, 5]] = [[0, 0],[0, 0]]. Confirmed.
Q: If the characteristic polynomial of T on V (dim 5) is f(t) = -(t-2)^3(t+1)^2, and W is a 3-dimensional T-invariant subspace with T_W having characteristic polynomial g(t), what can you say about g(t)?
A: By Theorem 5.21, g(t) divides f(t). Since dim(W) = 3, g(t) has degree 3. The possible forms are g(t) = c(t-2)^a(t+1)^b where a + b = 3, a ≤ 3, b ≤ 2.
The Cayley-Hamilton theorem is the stepping stone to the minimal polynomial (the lowest-degree polynomial annihilating T), which is central to the Jordan canonical form in Chapter 7. Invariant subspaces and their direct sum decompositions are the organising principle of Chapter 7: when V cannot be decomposed into one-dimensional invariant subspaces (i.e. T is not diagonalizable), it can still be decomposed into cyclic invariant subspaces, leading to the Jordan form. The block triangular structure from Theorem 5.21 also connects to Schur's theorem (Chapter 6), which guarantees an upper triangular representation.
T-invariant subspace, invariant subspace, cyclic subspace, T-cyclic subspace, restriction of linear operator, Cayley-Hamilton theorem, characteristic polynomial, f(T) equals zero, companion matrix, minimal polynomial, Jordan canonical form, direct sum of invariant subspaces, block triangular matrix