Why is the least-squares fit Ax* identified with the orthogonal projection of b onto C(A)?
The least-squares fit Ax∗ is identified with the orthogonal projection of b onto C(A) because the goal of least squares is to find the vector in the subspace C(A) that is closest to b. A fundamental geometric property of subspaces states that the unique closest point in a subspace to an external vector is its orthogonal projection. Therefore, Ax∗=projC(A)b.
Conditions
C(A) is a subspace of Rn
b∈Rn
Distance is measured by the Euclidean norm
Reasoning, step by step
Define the least-squares objective: minimize ∥b−Ax∗∥ subject to Ax∗∈C(A).
Invoke the geometric theorem: For any subspace W and vector b, the vector w∈W minimizing ∥b−w∥ is projWb.
Apply this theorem with W=C(A).
Conclude that the optimal image vector Ax∗ must equal projC(A)b.
Example
The video boxes the equation Ax∗=projC(A)b after explaining that the closest vector in the subspace is the projection.
Common misconceptions
Thinking that the projection is onto the null space instead of the column space.
Believing that any vector in C(A) is equally close to b.
The solution to the normal equations is considered the least-squares solution because it satisfies the necessary and sufficient condition for minimizing the residual norm ∥b−Ax∥. The derivation shows that minimizing this norm is equivalent to requiring the residual Ax∗−b to be orthogonal to the column space C(A).
Conditions: The original system Ax=b may be inconsistent; ATAx∗=ATb has a solution
Every product Ax is a member of the column space C(A) because matrix-vector multiplication is defined as a linear combination of the columns of A. Specifically, if A=[a1a2⋯ak] and x=[x1,…,xk]T, then Ax=x1a1+⋯+xkak.
Conditions: A is an n×k matrix; x∈Rk; C(A) is the span of the columns of A
The equation Ax=b has no solution because solving it is equivalent to finding weights x1,…,xk such that x1a1+⋯+xkak=b. The column space C(A) is defined as the set of all possible linear combinations of the columns of A.
Conditions: A is an n×k matrix; x∈Rk and b∈Rn; C(A) denotes the column space of A
Minimizing the Euclidean norm ∥b−Ax∗∥ is equivalent to minimizing its square, ∥b−Ax∗∥2. The squared norm expands algebraically into the sum of the squared differences of corresponding components: (b1−v1)2+(b2−v2)2+⋯+(bn−vn)2, where v=Ax∗.
Conditions: b,v∈Rn; Standard Euclidean norm is used; v=Ax∗
When Ax=b has no exact solution, the least-squares solution x∗ is defined as the vector that minimizes the Euclidean norm of the residual, ∥b−Ax∗∥. Geometrically, this means choosing x∗ such that Ax∗ is the closest possible vector to b within the column space C(A).
Conditions: The system Ax=b is inconsistent (no exact solution exists); Distance is measured by the standard Euclidean norm