Skip to content
← All questions

How does minimizing the norm ||b - Ax*|| relate to minimizing the sum of squared residuals?

Minimizing the Euclidean norm ∥b⃗−Ax⃗∗∥\|\vec{b} - A\vec{x}^*\| is equivalent to minimizing its square, ∥b⃗−Ax⃗∗∥2\|\vec{b} - A\vec{x}^*\|^2. The squared norm expands algebraically into the sum of the squared differences of corresponding components: (b1−v1)2+(b2−v2)2+⋯+(bn−vn)2(b_1 - v_1)^2 + (b_2 - v_2)^2 + \cdots + (b_n - v_n)^2, where v⃗=Ax⃗∗\vec{v} = A\vec{x}^*. This explicit sum-of-squares form is the origin of the term 'least squares'.

Conditions

  • b⃗,v⃗∈Rn\vec{b}, \vec{v} \in \mathbb{R}^n
  • Standard Euclidean norm is used
  • v⃗=Ax⃗∗\vec{v} = A\vec{x}^*

Reasoning, step by step

  1. Start with the objective to minimize the length (norm) of the residual vector b⃗−v⃗\vec{b} - \vec{v}.
  2. Square the norm to simplify the optimization (since x\sqrt{x} is monotonic, minimizing the norm is equivalent to minimizing the squared norm).
  3. Expand the squared norm ∥b⃗−v⃗∥2\|\vec{b} - \vec{v}\|^2 using the definition of the Euclidean norm.
  4. Write the expansion as ∑i=1n(bi−vi)2\sum_{i=1}^n (b_i - v_i)^2.
  5. Identify this sum of squared component errors as the 'least squares' objective.

Example

The board shows the progression from 'minimize ||b - Ax*||' to the vector [b1−v1,…,bn−vn]T[b_1-v_1, \dots, b_n-v_n]^T and finally to the scalar expression (b1−v1)2+⋯+(bn−vn)2(b_1-v_1)^2 + \dots + (b_n-v_n)^2.

Common misconceptions

  • Thinking that squaring the norm changes the location of the minimum.
  • Confusing the residual vector with the error in the coefficients x.

Watch the explanation

Connected concepts

Explore next

Related questions

Understand why

↗
Understand why

↗
Understand why

↗
Understand why

↗
Find a method

↗

Answers are generated from source material and independently checked. Consult the original video or creator if something is unclear.