How does the orthogonality of the residual lead to the normal equations ATAx* = ATb?
The residual vector r=Ax∗−b is orthogonal to the column space C(A) because Ax∗ is the orthogonal projection of b onto C(A). Orthogonality to C(A) means r is in the orthogonal complement C(A)⊥. Using the identity C(A)⊥=N(AT), we know r∈N(AT), which implies ATr=0. Substituting r gives AT(Ax∗−b)=0, which expands to the normal equations ATAx∗=ATb.
Conditions
Ax∗ is the orthogonal projection of b onto C(A)
C(A)⊥=N(AT) (Fundamental Theorem of Linear Algebra)
Matrix multiplication distributes over subtraction
Reasoning, step by step
Establish that the residual Ax∗−b is orthogonal to every vector in C(A).
Translate geometric orthogonality to algebraic membership: Ax∗−b∈C(A)⊥.
Apply the subspace identity C(A)⊥=N(AT) to conclude Ax∗−b∈N(AT).
Use the definition of the null space: AT(Ax∗−b)=0.
Distribute AT to get ATAx∗−ATb=0.
Rearrange to obtain the normal equations: ATAx∗=ATb.
Example
The board shows the chain: Ax∗−b∈C(A)⊥, then C(A)⊥=N(AT), then AT(Ax∗−b)=0, leading to ATAx∗=ATb.
Common misconceptions
Confusing the left null space N(AT) with the right null space N(A).
Thinking that the residual is orthogonal to the rows of A rather than the columns.
The least-squares fit Ax∗ is identified with the orthogonal projection of b onto C(A) because the goal of least squares is to find the vector in the subspace C(A) that is closest to b. A fundamental geometric property of subspaces states that the unique closest point in a subspace to an external vector is its orthogonal projection.
Conditions: C(A) is a subspace of Rn; b∈Rn; Distance is measured by the Euclidean norm
The solution to the normal equations is considered the least-squares solution because it satisfies the necessary and sufficient condition for minimizing the residual norm ∥b−Ax∥. The derivation shows that minimizing this norm is equivalent to requiring the residual Ax∗−b to be orthogonal to the column space C(A).
Conditions: The original system Ax=b may be inconsistent; ATAx∗=ATb has a solution
Every product Ax is a member of the column space C(A) because matrix-vector multiplication is defined as a linear combination of the columns of A. Specifically, if A=[a1a2⋯ak] and x=[x1,…,xk]T, then Ax=x1a1+⋯+xkak.
Conditions: A is an n×k matrix; x∈Rk; C(A) is the span of the columns of A
The equation Ax=b has no solution because solving it is equivalent to finding weights x1,…,xk such that x1a1+⋯+xkak=b. The column space C(A) is defined as the set of all possible linear combinations of the columns of A.
Conditions: A is an n×k matrix; x∈Rk and b∈Rn; C(A) denotes the column space of A
Minimizing the Euclidean norm ∥b−Ax∗∥ is equivalent to minimizing its square, ∥b−Ax∗∥2. The squared norm expands algebraically into the sum of the squared differences of corresponding components: (b1−v1)2+(b2−v2)2+⋯+(bn−vn)2, where v=Ax∗.
Conditions: b,v∈Rn; Standard Euclidean norm is used; v=Ax∗