Skip to content
Back to exploration
Algebra / English

Least squares approximation | Linear Algebra | Khan Academy

Khan Academy · YouTube · 15:32

Open original
READ & KEEP

The explanation, unpacked.

Reviewed learning material · Video analysis · English
Read the full overview

This 180-second whiteboard segment sets up the motivation for least squares by studying when Ax⃗=b⃗A\vec{x}=\vec{b} has no solution. It first fixes dimensions for an n×kn\times k matrix, then rewrites the equation as a linear combination of the columns of AA. From there it identifies inconsistency with the statement b⃗∉C(A)\vec{b}\notin C(A) and illustrates that geometrically with a plane for the column space and a vector outside it. The final seconds pivot toward approximation by introducing a candidate x⃗∗\vec{x}^*, but the exact criterion is not completed within the supplied clip. This 180-second whiteboard segment motivates the least-squares approximation for an inconsistent linear system Ax=b. It starts from the geometric fact that b is not in the column space C(A)C(A), so no exact solution exists. The lecturer then defines x* by minimizing the distance ||b - Ax*||, observes that Ax* = v must lie in C(A)C(A), expands the residual b - v into components, and rewrites the squared length as a sum of squared errors. This algebraic form directly explains the terminology 'least squares estimate', 'least squares solution', or 'least squares approximation'. This 180-second whiteboard segment develops the geometric meaning of a least-squares solution when Ax=b is inconsistent. It begins by replacing exact solvability with the objective of minimizing ||b - A x^*||, then uses the fact that the closest vector in a subspace is the orthogonal projection to characterize the solution by A x^* = proj_{C(A)C(A)} b. The speaker notes that computing the projection matrix directly is cumbersome, so the clip ends by subtracting b from both sides to set up an easier derivation. This 180-second whiteboard segment explains least squares geometrically and algebraically. Starting from the unsolvable system Ax=b, it defines x* so that Ax* is the closest point to b in the column space C(A)C(A), identifies Ax* with proj_{C(A)C(A)} b, and uses the perpendicular residual to show Ax∗−bA x* - b lies in C(A)⊥C(A)^{\perp}. It then invokes C(A)⊥=N(AT)C(A)^{\perp}=N(A^T), multiplies by ATA^T, and derives the normal equations ATAxA^T A x* = ATbA^T b. This 180-second whiteboard segment reviews the least-squares problem for an inconsistent linear system Ax=b. It first restates the definition of a least-squares solution x^* as the vector minimizing ||b - A x^*||, then connects that minimization to the geometry of projecting b onto the column space C(A)C(A). After noting that directly computing the projection is cumbersome, the clip introduces the standard algebraic shortcut: multiply the original equation on the left by ATA^T to obtain the normal equations ATAx=ATbA^T A x = A^T b. The speaker emphasizes that this new equation always has a solution and that its solution is the desired least-squares solution, not an exact solution of the original inconsistent system. This 32-second clip is a whiteboard-style summary of the least-squares method for an inconsistent linear system Ax=b. The speaker explains that if one can solve the displayed matrix-vector equation, then one has made the best possible attempt at solving Ax=b by minimizing the error between b and A x^*. The board ties together three central viewpoints: minimizing ||b - A x^*||, identifying A x^* with proj_{C(A)C(A)} b, and deriving the normal equations ATAxA^T A x^* = ATbA^T b from the orthogonality condition ATA^T(b - A x^*) = 0. In the last seconds, the camera zooms out to show the larger conceptual map, including the column-combination form of Ax=b, the geometric picture of projecting b onto the plane C(A)C(A), and the subspace identity C(A)⊥=N(AT)C(A)^\perp = N(A^T). The speaker closes by saying the material is still somewhat abstract but will become clearly useful in the next video.

Use the learning inspector for key ideas and moments, or open the reading tabs for the complete notes.

Chapters

0:00Set up Ax⃗=b⃗A\vec{x}=\vec{b} with A∈Rn×kA\in\mathbb{R}^{n\times k}0:24Assume there is no solution0:40Expand AA into column vectors1:14Interpret no solution as b⃗∉C(A)\vec{b}\notin C(A)1:48Draw the column space and b⃗\vec{b} outside it2:27From inconsistency toward approximation with x⃗∗\vec{x}^*3:00Setup: Ax=b has no solution because b is outside C(A)C(A)3:25Defining least squares by minimizing ||b - Ax*||3:52Interpreting Ax* as a vector v in the column space4:48Expanding the residual into componentwise differences5:08Squared norm becomes a sum of squares5:25Naming x* the least squares estimate6:00No exact solution, but a best approximation6:14Closest vector in a subspace is the projection6:40Least-squares condition A x^* = proj_{C(A)C(A)} b7:47Why the direct projection formula is inconvenient8:13Subtract b to prepare the next step9:00Least-squares setup and projection picture9:40Residual orthogonality to the column space10:20Using C(A)⊥=N(AT)C(A)^{\perp}=N(A^T)11:00Multiplying by ATA^T and deriving the normal equations12:00Board recap and algebraic cleanup12:27Recall the least-squares problem13:10Projection onto the column space14:00Derive the normal equations14:40Identify the least-squares solution15:00Best approximate solution to Ax=b15:09Matrix-versus-vector structure and minimizing the error15:18Least-square solution and abstraction note15:24Zoom-out to full geometric and subspace picture

Learning script

Generated from the video's visuals and explanation; not verbatim speech.

The lesson opens by writing a general matrix equation on the board: Ax⃗=b⃗A\vec{x}=\vec{b}, with AA specified as an n×kn\times k matrix. The speaker immediately uses dimension compatibility to place the unknowns correctly: because AA has kk columns, x⃗\vec{x} must lie in Rk\mathbb{R}^k, and because the product has nn rows, b⃗\vec{b} must lie in Rn\mathbb{R}^n. This establishes the setting for the whole discussion.

Next, the speaker adds a hypothesis that matters: suppose this system has no solution. That statement is written explicitly as "No solution to Ax⃗=b⃗A\vec{x}=\vec{b}". The purpose is not yet to solve the system, but to understand what failure of solvability means structurally.

To expose that structure, the matrix AA is rewritten in terms of its columns, [a⃗1 a⃗2 ⋯ a⃗k][\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k], and multiplied by the component vector [x1x2⋮xk]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}. The board then shows the equivalent linear-combination form x1a⃗1+x2a⃗2+⋯+xka⃗k=b⃗x_1\vec{a}_1+x_2\vec{a}_2+\cdots+x_k\vec{a}_k=\vec{b}. This step translates the abstract matrix equation into a concrete statement about weights on the columns of AA.

With that reformulation in hand, the speaker interprets the assumption of no solution: there is no choice of weights on the columns of AA that produces b⃗\vec{b}. Equivalently, no linear combination of the columns of AA equals b⃗\vec{b}. The clip then names this geometrically by writing that b⃗\vec{b} is not in the column space C(A)C(A).

The algebraic claim is then visualized. A purple plane-like region is drawn and labeled C(A)C(A), standing for the collection of all linear combinations of the columns of AA. An origin is marked, and a cyan arrow labeled b⃗\vec{b} is drawn starting from the origin but pointing outside that plane. The picture makes the previous sentence intuitive: if b⃗\vec{b} lies outside the column space, then it cannot be reached by any combination of the columns.

After establishing that exact solvability fails, the speaker contrasts this with an older habit of stopping there: form an augmented matrix, row-reduce, encounter a contradiction such as 0=10=1, and conclude there is nothing more to do. The lesson instead asks whether one can do better by finding something close to a solution.

That question leads directly to a new symbol, x⃗∗\vec{x}^*, written at the lower left together with the partial phrase "where Ax⃗∗A\vec{x}^*". Within this supplied segment, the intended role of x⃗∗\vec{x}^* is clear at the level of motivation: it is being introduced as a candidate associated with an approximate solution. However, the exact condition it is supposed to satisfy is cut off before the clip ends.

The board opens on an inconsistent linear system. We have Ax=bA x = b with x∈Rkx \in R^k and b∈Rnb \in R^n, and the lecturer has already written that there is no solution to Ax=bA x = b. On the left, the same equation is expanded in two equivalent ways: as a matrix times a column vector, [a1a2a_1 a_2 ... aka_k][x1x2x_1 x_2 ... xkx_k]^T=bT = b, and as a linear combination of columns, x1a1+x2a2x_1 a_1 + x_2 a_2 + ... + xkak=bx_k a_k = b. On the right, a purple plane labeled C(A)C(A) represents the column space, while a cyan arrow labeled b points outside that plane. The key point is that exact solvability would require b to be expressible as a combination of the columns of A, but the picture explicitly says b is not in C(A)C(A).

From that failed exact solution, the lecture pivots to approximation. The speaker says that when he talks about getting 'close,' he means close in length. He therefore introduces x* by the requirement that A x* be as close as possible to b, and writes the objective as minimize ||b - A x*||. Mathematically, this changes the question from 'solve Ax=bA x = b exactly' to 'choose x* so that the residual vector b - A x* has smallest possible Euclidean norm.'

To make the geometry clearer, the lecturer renames the candidate output vector. He says A x is always a member of the column space, because multiplying any vector in RkR^k by A forms a linear combination of the columns of A. He writes v=Axv = A x* and places v inside the plane C(A)C(A). Thus the optimization is no longer abstract: among all vectors v that lie in C(A)C(A), we want the one nearest to the external point b. The invariant constraint is v∈C(A)v \in C(A); the quantity being reduced is the distance from b to v.

Next the residual is written in coordinates. Since v=Axv = A x*, the vector b - A x* becomes b - v. The lecturer expands this as the column vector [b1−v1b_1 - v_1, b2−v2b_2 - v_2, ..., bn−vnb_n - v_n]^T. This step shows that the single geometric distance ||b - v|| is built from all coordinate mismatches between the target vector b and the chosen column-space vector v.

The lecture then squares the residual norm. He states that the length squared of [b1−v1b_1 - v_1, ..., bn−vnb_n - v_n]^T is (b1−v1b_1 - v_1)^2+(b2−v2)22 + (b_2 - v_2)^2 + ... + (bn−vnb_n - v_n)^2. This is the crucial algebraic reformulation: minimizing the Euclidean distance is represented by minimizing an explicit sum of squared errors. The reason for squaring here is not to change the geometric target but to expose the additive quadratic structure that will name the method.

Finally, the terminology is attached to the construction. Because the quantity being minimized is a sum of squares of the residuals, the lecturer calls x* the least squares estimate for Ax=bA x = b, also referring to it as the least squares solution or least squares approximation. The segment closes with the conceptual chain fully visible: b outside C(A)C(A) prevents an exact solution; x* is chosen so that A x* = v in C(A)C(A) is nearest to b; and the nearest-distance condition is expressed algebraically as minimizing the sum of squared componentwise errors.

The board starts from an inconsistent system: there is no solution to Ax=b. Instead of stopping there, the lesson introduces a least-squares goal: find x^* so that A x^* is as close as possible to b.

That closeness is expressed by minimizing ||b - A x^*||. Geometrically, A x^* must lie in the column space C(A)C(A), because multiplying A by any vector produces a linear combination of the columns of A.

The speaker then recalls a general geometric fact: among all vectors in a subspace, the one closest to an external vector is the orthogonal projection of that external vector onto the subspace.

Applying that fact with subspace C(A)C(A) and external vector b gives the key characterization of the least-squares solution: A x^* = proj_{C(A)C(A)} b. This equation is written and boxed on the board.

However, the speaker points out that using the explicit projection formula A(ATA)−1ATbA(A^T A)^{-1}A^T b is still hard computationally, so the lesson looks for an easier path.

To begin that easier derivation, b is subtracted from both sides of the boxed equation, producing A x^* - b = proj_{C(A)C(A)} b - b, which reframes the problem in terms of residuals and prepares the next argument.

The board opens with the approximate-system context: a matrix equation [a1a_1 ... aka_k][x1x_1 ... xkx_k]^T=bT = b is shown, together with the note that the goal is least squares. The speaker frames x* as the choice for which Ax* is as close as possible to b, and the objective is written as minimizing ||b-Ax*||, expanded into a sum of squared component differences.

At the center, a geometric picture supports the algebra. A purple plane labeled C(A)C(A) represents the column space. The vector b is drawn outside that plane, while the yellow vector Ax* lies inside it. The speaker identifies Ax* with the orthogonal projection of b onto C(A)C(A), making explicit that the least-squares fit is the nearest point in the column space.

The explanation then focuses on the error vector between b and its projection. Using the right side of the board, the speaker writes Ax∗−bA x* - b and notes that this is the same as proj_{C(A)C(A)} b - b once the projection identity is used. The key claim is that this difference is perpendicular to the plane C(A)C(A), so it belongs to C(A)⊥C(A)^{\perp}.

Next, the video translates the geometric orthogonality condition into a standard matrix identity. The board states C(A)⊥=N(AT)C(A)^{\perp}=N(A^T), and the speaker calls N(AT)N(A^T) the left null space of A. From this, the residual condition is rewritten as Ax∗−b∈N(AT)A x* - b \in N(A^T).

Once the residual is known to lie in the left null space, the speaker applies the defining property of that null space: multiplying it by ATA^T gives the zero vector. This produces AT(Ax∗−b)=0A^T(A x* - b)=0.

Finally, the expression is expanded by distributivity to ATAx∗−ATb=0A^T A x* - A^T b = 0, which is the normal-equation form derived in this clip. The segment closes on this algebraic condition, linking the original projection geometry to the practical linear system used to compute least-squares solutions.

The clip opens on a dense blackboard summary of the least-squares topic. Visible relations include the minimization statement for ||b - A x^*||, the projection identity A x^* = proj_{C(A)C(A)} b, an orthogonality relation involving ATA^T(Ax^*-b), and the normal equation ATAxA^T A x^* = ATbA^T b. The speaker is finishing an algebraic simplification that leaves the normal-equation form on the board.

The explanation then steps back to the origin of the problem. The system Ax=bA x = b is inconsistent, so there is no exact solution. Instead of solving that impossible equation, the goal is to choose a vector x^* that makes A x^* as close as possible to b. This is expressed by minimizing the residual norm ||b - A x^*||. The term 'least squares' is justified because minimizing the length of the residual is equivalent to minimizing the sum of squared component differences.

Next, the video gives the geometric interpretation. Among all vectors in the column space C(A)C(A), the one closest to b is the orthogonal projection of b onto C(A)C(A). Therefore the optimal image A x^* must satisfy A x^* = proj_{C(A)C(A)} b. If some x^* in RkR^k produces that projected vector, then that x^* is the least-squares solution.

The speaker then notes that computing the projection directly is conceptually clear but algebraically cumbersome. A simpler route is to transform the original inconsistent equation. Starting from Ax=bA x = b, multiply both sides on the left by ATA^T. This gives ATAx=ATbA^T A x = A^T b, the normal equations.

Finally, the clip stresses the logical status of this new equation. Its solution is not an exact solution of the original system Ax=bA x = b. Rather, the normal equations are guaranteed to have a solution, and that solution is precisely the least-squares solution x^* sought at the start.

The clip opens on a dense handwritten board centered on the least-squares problem for a linear system Ax=b. The speaker begins by pointing out the structural difference between the objects in the equation being discussed: one side is a matrix and the other is a vector. This sets up the practical claim that if we can solve this displayed equation, then we have done the best possible job of approximating a solution to the original system.

The board makes that “best possible job” precise through the minimization statement minimize ||b - A x^*||. In words, x^* is chosen so that A x^* is as close as possible to b. The speaker then states the consequence directly: we have minimized the error, we obtain A x^*, and the difference between A x^* and b is as small as possible. This is exactly what the board labels as the least-square solution.

Alongside the minimization formulation, the board presents an equivalent geometric identity: A x^* = proj_{C(A)C(A)} b. That is, the image of the least-squares solution is not arbitrary; it is the orthogonal projection of b onto the column space of A. The board also records the residual condition in several compatible forms, including A x^* - b∈C(A)⊥b \in C(A)^\perp and ATA^T(b - A x^*) = 0. From this orthogonality condition, the displayed algebra rewrites the problem into the boxed normal equations ATAxA^T A x^* = ATbA^T b.

In the final seconds, the camera zooms out to reveal the wider conceptual layout. On the left, Ax=b is rewritten as a column-combination problem, x1a1+x2a2+⋯+xkak=bx_1 a_1 + x_2 a_2 + \cdots + x_k a_k = b, clarifying that solving the system means asking whether b lies in the span of the columns of A. In the middle, a geometric sketch shows b, its projection onto the plane representing C(A)C(A), and the perpendicular residual direction. On the right, the board states C(A)⊥=N(AT)C(A)^\perp = N(A^T), linking the orthogonal-complement description of the residual to the null space of A transpose. Over this full-board view, the speaker remarks that the topic still feels abstract in this clip but promises that the next video will show why the idea is very useful.

Knowledge cards

01

Dimension setup for Ax⃗=b⃗A\vec{x}=\vec{b}

For an n×kn\times k matrix AA, the equation Ax⃗=b⃗A\vec{x}=\vec{b} forces x⃗∈Rk\vec{x}\in\mathbb{R}^k and b⃗∈Rn\vec{b}\in\mathbb{R}^n. This is the starting framework used in the clip before discussing solvability.

A∈Rn×k, x⃗∈Rk, b⃗∈RnA\in\mathbb{R}^{n\times k},\ \vec{x}\in\mathbb{R}^k,\ \vec{b}\in\mathbb{R}^n
02

Matrix-vector product as a column combination

Writing AA by columns turns Ax⃗=b⃗A\vec{x}=\vec{b} into x1a⃗1+x2a⃗2+⋯+xka⃗k=b⃗x_1\vec{a}_1+x_2\vec{a}_2+\cdots+x_k\vec{a}_k=\vec{b}. Thus solving the system means asking whether b⃗\vec{b} can be formed from the columns of AA.

[a⃗1 a⃗2 ⋯ a⃗k][x1x2⋮xk]=b⃗  ⟺  x1a⃗1+⋯+xka⃗k=b⃗[\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}=\vec{b}\iff x_1\vec{a}_1+\cdots+x_k\vec{a}_k=\vec{b}
03

No solution means b⃗∉C(A)\vec{b}\notin C(A)

The clip states that if Ax⃗=b⃗A\vec{x}=\vec{b} has no solution, then no linear combination of the columns of AA equals b⃗\vec{b}. In column-space language, this is exactly the assertion that b⃗\vec{b} is not in C(A)C(A).

Ax⃗=b⃗ has no solution   ⟺  b⃗∉C(A)A\vec{x}=\vec{b}\text{ has no solution }\iff \vec{b}\notin C(A)
04

Geometric picture of the column space

A plane-like region labeled C(A)C(A) represents all attainable linear combinations of the columns of AA. Drawing b⃗\vec{b} outside that region visualizes why the equation cannot be solved exactly.

05

From inconsistency to approximation

Instead of ending at “no solution,” the lesson pivots toward finding a vector x⃗∗\vec{x}^* that gets close to the target. The symbol x⃗∗\vec{x}^* is introduced, but the precise defining condition is not completed within this segment.

06

Inconsistent system Ax=b when b is outside C(A)C(A)

The segment begins with the equation Ax=bA x = b, where A is n×kn\times k, x∈Rkx \in R^k, and b∈Rnb \in R^n. The board states that there is no solution because b is not in the column space C(A)C(A). The same equation is rewritten as x1a1+x2a2x_1 a_1 + x_2 a_2 + ... + xkak=bx_k a_k = b, showing that an exact solution would require b to be a linear combination of the columns of A.

Ax⃗=b⃗,x1a⃗1+x2a⃗2+⋯+xka⃗k=b⃗A\vec{x}=\vec{b},\quad x_1\vec{a}_1+x_2\vec{a}_2+\cdots+x_k\vec{a}_k=\vec{b}
07

Least-squares objective as distance minimization

Since exact equality is impossible, the lecturer defines x* by making A x* as close as possible to b. 'Close' is made precise using vector length, so the problem becomes minimizing the norm of the residual b - A x*.

minimize ∥b⃗−Ax⃗∗∥\text{minimize }\|\vec{b}-A\vec{x}^*\|
08

Why Ax* must lie in the column space

The speaker introduces v=Axv = A x* and notes that any product A x is a linear combination of the columns of A, hence belongs to C(A)C(A). Geometrically, the least-squares problem is to choose the vector v in the plane C(A)C(A) that is nearest to the external vector b.

v⃗=Ax⃗∗∈C(A)\vec{v}=A\vec{x}^*\in C(A)
09

Residual vector in coordinates

Once v=Axv = A x* is substituted, the residual b - A x* is rewritten as b - v and expanded componentwise. This displays the error as the vector of coordinate differences between b and the chosen column-space vector v.

b⃗−v⃗=[b1−v1b2−v2⋮bn−vn]\vec{b}-\vec{v}=\begin{bmatrix}b_1-v_1\\b_2-v_2\\\vdots\\b_n-v_n\end{bmatrix}
10

Squared length becomes a sum of squares

The lecturer then squares the residual norm. The squared Euclidean length of the residual vector equals the sum of the squared componentwise errors. This algebraic form is what motivates the name of the method.

∥[b1−v1b2−v2⋮bn−vn]∥2=(b1−v1)2+(b2−v2)2+⋯+(bn−vn)2\left\|\begin{bmatrix}b_1-v_1\\b_2-v_2\\\vdots\\b_n-v_n\end{bmatrix}\right\|^2=(b_1-v_1)^2+(b_2-v_2)^2+\cdots+(b_n-v_n)^2
11

Meaning of least squares estimate / solution / approximation

Because the objective is a sum of squared residuals, x* is called the least squares estimate for Ax=bA x = b, also the least squares solution or least squares approximation. The terminology refers to minimizing the total squared error between b and the best attainable vector A x* in C(A)C(A).

x∗ is the least squares estimate for Ax⃗=b⃗.x^*\text{ is the least squares estimate for }A\vec{x}=\vec{b}.
12

Least-squares setup for an inconsistent system

When Ax=b has no exact solution, the video defines a least-squares approach by seeking x^* such that A x^* is as close as possible to b. This is formalized as minimizing the residual norm ||b - A x^*||.

min⁡  ∥b−Ax∗∥\min \; \|b - A x^*\|
13

Closest vector in a subspace is the projection

For a vector outside a subspace, the nearest point inside the subspace is its orthogonal projection. In this lesson the relevant subspace is the column space C(A)C(A).

closest vector in C(A) to b=proj⁡C(A)b\text{closest vector in } C(A) \text{ to } b = \operatorname{proj}_{C(A)} b
14

Characterization of the least-squares solution

Combining the minimization objective with the projection principle yields the central equation of the clip: the least-squares solution satisfies A x^* = proj_{C(A)C(A)} b.

Ax∗=proj⁡C(A)bA x^* = \operatorname{proj}_{C(A)} b
15

Direct projection formula is cumbersome here

The speaker notes that computing the projection via A(ATA)−1ATbA(A^T A)^{-1}A^T b is difficult in practice, motivating a different derivation route.

proj⁡C(A)b=A(ATA)−1ATb\operatorname{proj}_{C(A)} b = A(A^T A)^{-1}A^T b
16

Residual rearrangement of the least-squares equation

To search for an easier method, the boxed equation is rewritten by subtracting b from both sides, giving a residual form that sets up the next step of the lesson.

Ax∗−b=proj⁡C(A)b−bA x^* - b = \operatorname{proj}_{C(A)} b - b
17

Least-squares solution as closest point in C(A)C(A)

The clip defines the least-squares solution x* by requiring Ax* to be as close as possible to b. On the board this is expressed through the minimization of ||b-Ax*|| and visually through the column-space plane C(A)C(A). The important point is that x* is not presented as an exact solution to Ax=b, but as the coefficient vector producing the best approximation inside the range of A.

minimize ∥b⃗−Ax⃗∗∥\text{minimize } \|\vec b-A\vec x^{*}\|
18

Projection identity for least squares

The geometric core of the lesson is the boxed identity A x* = proj_{C(A)C(A)} b. This says that the least-squares image of x* is exactly the orthogonal projection of b onto the column space. The diagram reinforces this by placing Ax* in the plane C(A)C(A) and b outside it.

Ax⃗∗=proj⁡C(A)b⃗A\vec x^{*}=\operatorname{proj}_{C(A)}\vec b
19

Orthogonality of the residual

Because Ax* is the projection of b onto C(A)C(A), the difference vector from b to Ax* is perpendicular to the entire column space. The board records this as Ax∗−b∈C(A)⊥A x* - b \in C(A)^{\perp}. This is the bridge from the geometric picture to the later algebraic derivation.

Ax⃗∗−b⃗∈C(A)⊥A\vec x^{*}-\vec b \in C(A)^{\perp}
20

Column-space complement equals left null space

The speaker then uses the fundamental-subspaces identity C(A)⊥=N(AT)C(A)^{\perp}=N(A^T), naming N(AT)N(A^T) the left null space of A. This converts the orthogonality statement into a null-space membership statement for the residual.

C(A)⊥=N(AT)C(A)^{\perp}=N(A^T)
21

Deriving the normal equations

From Ax∗−b∈N(AT)A x* - b \in N(A^T), multiplying by ATA^T gives zero. Expanding the product yields the normal equations in the form shown at the end of the clip. This is the main algebraic payoff of the geometric argument.

ATAx⃗∗=ATb⃗A^T A\vec x^{*}=A^T\vec b
22

Least-squares solution definition

For an inconsistent system Ax=bA x = b, the least-squares solution x^* is defined as the vector that minimizes the residual size ||b - A x^*||. The name comes from equivalently minimizing the sum of squared component errors.

x⃗∗ minimizes ∥b⃗−Ax⃗∗∥\vec x^* \text{ minimizes } \|\vec b - A\vec x^*\|
23

Projection characterization

The best approximant to b inside the column space C(A)C(A) is the orthogonal projection of b onto C(A)C(A). Hence the least-squares condition can be written as A x^* = proj_{C(A)C(A)} b.

Ax⃗∗=proj⁡C(A)b⃗A\vec x^*=\operatorname{proj}_{C(A)}\vec b
24

Normal equations

A practical way to compute the least-squares solution is to left-multiply the original equation Ax=bA x = b by ATA^T, producing ATAx=ATbA^T A x = A^T b. Solving this system gives the least-squares solution x^*.

ATAx⃗∗=ATb⃗A^{\mathrm T}A\vec x^*=A^{\mathrm T}\vec b

Detailed learning notes

Explore conditions, steps and evidence. Supplementary explanations are labeled separately from content shown in the video.

Symbols · 60

A

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "Let's say I have some matrix A."

  2. Formula
    Observation

    Yellow handwritten AA is written at the upper left.

Symbol

A

Meaning

An n×kn \times k matrix whose columns are used to form linear combinations.

Domain

A∈Rn×kA \in \mathbb{R}^{n \times k}

x⃗\vec{x}

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "I have the equation Ax is equal to b" and later "x would have to be a member of Rk."

  2. Formula
    Observation

    Yellow handwritten x⃗\vec{x} appears in Ax⃗=b⃗A\vec{x}=\vec{b}, with x⃗∈Rk\vec{x}\in\mathbb{R}^k written nearby.

Symbol

x⃗\vec{x}

Meaning

Unknown vector in the matrix equation Ax⃗=b⃗A\vec{x}=\vec{b}.

Domain

x⃗∈Rk\vec{x}\in\mathbb{R}^k

b⃗\vec{b}

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "Ax is equal to b" and "b is a member of Rn."

  2. Formula
    Observation

    Yellow handwritten b⃗\vec{b} appears in Ax⃗=b⃗A\vec{x}=\vec{b}, with b⃗∈Rn\vec{b}\in\mathbb{R}^n written nearby.

Symbol

b⃗\vec{b}

Meaning

Right-hand-side vector in the equation Ax⃗=b⃗A\vec{x}=\vec{b}.

Domain

b⃗∈Rn\vec{b}\in\mathbb{R}^n

n, k

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "it's an n by k matrix" and explains that xx has kk entries because there are kk columns.

  2. Formula
    Observation

    Yellow handwritten n×kn \times k is placed under AA.

Symbol

n, k

Meaning

Dimensions of AA: nn rows and kk columns.

Domain

Positive integers

a⃗1\vec{a}_1,a⃗2\vec{a}_2,…\ldots,a⃗k\vec{a}_k

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "If I write A like this, A1, A2 ... all the way through Ak" and refers to them as column vectors.

  2. Formula
    Observation

    Pink handwritten bracketed expression [a⃗1 a⃗2 ⋯ a⃗k][\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k] is written below the initial equation.

Symbol

a⃗1\vec{a}_1,a⃗2\vec{a}_2,…\ldots,a⃗k\vec{a}_k

Meaning

Column vectors of AA, indexed from 11 to kk.

Domain

Each a⃗j\vec{a}_j is a column of AA; hence a⃗j∈Rn\vec{a}_j\in\mathbb{R}^n

x1x_1,x2x_2,…\ldots,xkx_k

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "multiply it times x1, x2 all the way through xk."

  2. Formula
    Observation

    Pink handwritten column vector [x1x2⋮xk]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix} is written next to the column expansion of AA.

Symbol

x1x_1,x2x_2,…\ldots,xkx_k

Meaning

Scalar components of the unknown vector x⃗\vec{x}, used as weights on the columns of AA.

Domain

xj∈Rx_j\in\mathbb{R} for j=1,…,kj=1,\ldots,k

C(A)C(A)

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "b is not in the column space of A."

  2. Formula
    Observation

    Yellow handwritten text states "b⃗\vec{b} is not in the C(A)C(A)".

Symbol

C(A)C(A)

Meaning

Column space of AA, i.e. the set of all linear combinations of the columns of AA.

Domain

Subspace of Rn\mathbb{R}^n

x⃗∗\vec{x}^*

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "I want to find some x star for now where A times x star is..."

  2. Formula
    Observation

    Yellow handwritten x⃗∗\vec{x}^* appears at the lower left, followed by "where Ax⃗∗A\vec{x}^*".

Uncertainties
  1. The sentence and displayed condition after Ax⃗∗A\vec{x}^* are cut off before completion.

Symbol

x⃗∗\vec{x}^*

Meaning

A candidate vector introduced for a better approximate solution when Ax⃗=b⃗A\vec{x}=\vec{b} has no exact solution.

Domain

Intended to lie in Rk\mathbb{R}^k by analogy with x⃗\vec{x}, but the clip does not explicitly restate the domain

A

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    Matrix A appears in Ax=bA x = b and is described as n×kn\times k.

  2. Audio
    Observation

    The speaker refers to multiplying a vector by the matrix A and getting a member of its column space.

Symbol

A

Meaning

An n×kn\times k matrix whose columns span C(A)C(A).

Domain

n×kn\times k matrix

x

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    x∈Rkx \in R^k is written beside Ax=bA x = b.

  2. Audio
    Observation

    The speaker says any vector in RkR^k times the matrix A gives a member of the column space.

Symbol

x

Meaning

Unknown vector in RkR^k.

Domain

RkR^k

b

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    b∈Rnb \in R^n is written beside Ax=bA x = b.

  2. Audio
    Observation

    The speaker repeatedly refers to making Ax* as close as possible to b.

Symbol

b

Meaning

Target vector in RnR^n that generally lies outside C(A)C(A) in this setup.

Domain

RnR^n

x^*

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    x* is introduced in the line 'x* where Ax* is as close to b as possible'.

  2. Audio
    Observation

    The speaker defines x* through minimizing ||b - Ax*||.

Symbol

x^*

Meaning

Vector chosen so that Ax* is the closest point in C(A)C(A) to b.

Domain

RkR^k

Knowledge points · 30

Matrix equation setup and dimension matching

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker introduces a matrix AA that is nn by kk and the equation Ax=bAx=b, then states the dimensions of xx and bb.

  2. Formula
    Observation

    Yellow handwritten AA, n×kn\times k, Ax⃗=b⃗A\vec{x}=\vec{b}, x⃗∈Rk\vec{x}\in\mathbb{R}^k, and b⃗∈Rn\vec{b}\in\mathbb{R}^n appear across the top of the board.

Definition
Explanation

The clip begins with a general linear system written as Ax⃗=b⃗A\vec{x}=\vec{b}, where AA is an n×kn\times k matrix. Dimension compatibility forces x⃗\vec{x} to have kk entries and b⃗\vec{b} to have nn entries, so x⃗∈Rk\vec{x}\in\mathbb{R}^k and b⃗∈Rn\vec{b}\in\mathbb{R}^n.

Formula
Ax⃗=b⃗,A∈Rn×k,x⃗∈Rk,b⃗∈RnA\vec{x}=\vec{b},\quad A\in\mathbb{R}^{n\times k},\quad \vec{x}\in\mathbb{R}^k,\quad \vec{b}\in\mathbb{R}^n
Conditions
  1. AA has nn rows and kk columns

  2. The product Ax⃗A\vec{x} is defined only when x⃗\vec{x} has kk entries

  3. The resulting vector has nn entries, matching b⃗\vec{b}

Prerequisites
  1. A
  2. x⃗\vec{x}
  3. b⃗\vec{b}
  4. n, k

Matrix-vector product as a linear combination of columns

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker expands AA into its column vectors and multiplies by the components of xx, saying this is the same equation rewritten.

  2. Formula
    Observation

    Pink handwritten [a⃗1 a⃗2 ⋯ a⃗k][x1x2⋮xk]=b⃗[\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}=\vec{b} is written, followed by x1a⃗1+x2a⃗2+⋯+xka⃗k=b⃗x_1\vec{a}_1+x_2\vec{a}_2+\cdots+x_k\vec{a}_k=\vec{b}.

Formula
Explanation

The equation Ax⃗=b⃗A\vec{x}=\vec{b} is rewritten by expressing AA in terms of its columns. Multiplying the column-expanded matrix by the component vector gives the weighted sum of the columns of AA, showing that solving Ax⃗=b⃗A\vec{x}=\vec{b} is equivalent to finding weights that express b⃗\vec{b} as a linear combination of the columns of AA.

Formula
[a⃗1 a⃗2 ⋯ a⃗k][x1x2⋮xk]=b⃗⟺x1a⃗1+x2a⃗2+⋯+xka⃗k=b⃗[\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}=\vec{b}\quad\Longleftrightarrow\quad x_1\vec{a}_1+x_2\vec{a}_2+\cdots+x_k\vec{a}_k=\vec{b}
Conditions
  1. a⃗j\vec{a}_j denotes the jjth column of AA

  2. xjx_j denotes the jjth component of x⃗\vec{x}

  3. The equivalence uses the standard column interpretation of matrix-vector multiplication

Prerequisites
  1. Matrix equation setup and dimension matching
  2. a⃗1\vec{a}_1,a⃗2\vec{a}_2,…\ldots,a⃗k\vec{a}_k
  3. x1x_1,x2x_2,…\ldots,xkx_k

Interpretation of an inconsistent system via column space

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says there is no solution, then explains that this means no set of weights on the column vectors of AA can get to bb, equivalently no linear combination equals bb, and finally says bb is not in the column space of AA.

  2. Formula
    Observation

    Cyan handwritten text reads "No solution to Ax⃗=b⃗A\vec{x}=\vec{b}"; yellow handwritten text reads "b⃗\vec{b} is not in the C(A)C(A)".

Definition
Explanation

When Ax⃗=b⃗A\vec{x}=\vec{b} has no solution, the vector b⃗\vec{b} cannot be obtained as any linear combination of the columns of AA. The clip names this geometrically by saying b⃗\vec{b} is not in the column space C(A)C(A).

Formula
Ax⃗=b⃗ has no solution ⟺b⃗∉C(A)A\vec{x}=\vec{b}\text{ has no solution }\Longleftrightarrow \vec{b}\notin C(A)
Conditions
  1. C(A)C(A) is the set of all linear combinations of the columns of AA

  2. The equivalence is stated for the specific system under discussion

Prerequisites
  1. Matrix-vector product as a linear combination of columns
  2. C(A)C(A)

Inconsistent system Ax=b when b is outside C(A)C(A)

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board states 'No solution to Ax = b' and shows the equivalent column-combination equation.

  2. Diagram
    Observation

    A cyan arrow labeled b points outside the purple plane labeled C(A)C(A), with text 'b is not in the C(A)C(A)'.

Definition
Explanation

This segment begins from the case where the exact linear system Ax=b has no solution because b does not belong to the column space of A. The video expresses Ax both as a matrix product and as a linear combination of the columns of A, making clear that solvability requires b to lie in C(A)C(A).

Formula
Ax⃗=b⃗,x⃗∈Rk, b⃗∈Rn,[a⃗1 a⃗2 ⋯ a⃗k][x1x2⋮xk]=b⃗,x1a⃗1+x2a⃗2+⋯+xka⃗k=b⃗A\vec{x}=\vec{b},\quad \vec{x}\in\mathbb{R}^k,\ \vec{b}\in\mathbb{R}^n,\quad [\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}=\vec{b},\quad x_1\vec{a}_1+x_2\vec{a}_2+\cdots+x_k\vec{a}_k=\vec{b}
Conditions
  1. A is n×kn\times k

  2. x∈Rkx \in R^k

  3. b∈Rnb \in R^n

  4. b ∉ C(A)C(A) in the motivating picture

Least-squares objective as distance minimization

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The speaker writes 'minimize ||b - Ax*||'.

  2. Audio
    Observation

    He says, 'when I say close, I'm talking about length' and 'I want to minimize the length of b minus Ax*'.

  3. Diagram
    Observation

    The vector Ax* is identified with v inside C(A)C(A), while b remains outside the plane.

Definition
Explanation

The least-squares idea is introduced geometrically: choose x* so that Ax*, viewed as a point v in the column space, is as close as possible to b. Closeness is measured by vector length, so the problem becomes minimizing the norm of the residual b - Ax*.

Formula
Choose x∗ so that ∥b⃗−Ax⃗∗∥ is minimized.\text{Choose }x^*\text{ so that }\|\vec{b}-A\vec{x}^*\|\text{ is minimized.}
Conditions
  1. Ax* must lie in C(A)C(A)

  2. Distance is measured by Euclidean length

Prerequisites
  1. Inconsistent system Ax=b when b is outside C(A)C(A)

Every Ax lies in the column space C(A)C(A)

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, 'Ax is going to be a member of my column space' and repeats that any Ax is in the column space.

  2. Formula
    Observation

    He writes v = Ax* on the diagram.

  3. Diagram
    Observation

    The vector v is placed inside the purple plane labeled C(A)C(A).

Method
Explanation

The video uses the identity v = Ax* to translate the optimization into a geometric search inside C(A)C(A). Because Ax is always a linear combination of the columns of A, the candidate output vector must remain in the column space, and the task is to pick the best such vector relative to b.

Formula
v⃗=Ax⃗∗∈C(A)\vec{v}=A\vec{x}^*\in C(A)
Conditions
  1. x∗∈Rkx* \in R^k

  2. A is n×kn\times k

Prerequisites
  1. Inconsistent system Ax=b when b is outside C(A)C(A)
  2. Least-squares objective as distance minimization

Residual vector written componentwise

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board expands the norm into the vector [b1−v1b_1 - v_1, b2−v2b_2 - v_2, ..., bn−vnb_n - v_n]^T.

  2. Audio
    Observation

    The speaker says to take the difference between each of the elements.

Formula
Explanation

Once v = Ax* is introduced, the residual b - Ax* is rewritten as b - v and then displayed as the vector of componentwise differences. This makes explicit that the distance being minimized depends on all coordinate mismatches between b and the chosen column-space vector v.

Formula
b⃗−v⃗=[b1−v1b2−v2⋮bn−vn]\vec{b}-\vec{v}=\begin{bmatrix}b_1-v_1\\b_2-v_2\\\vdots\\b_n-v_n\end{bmatrix}
Conditions
  1. b,v∈Rnv \in R^n

Prerequisites
  1. Every Ax lies in the column space C(A)C(A)
  2. Least-squares objective as distance minimization

Squared Euclidean length equals sum of squared residuals

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board shows ||[b1−v1b_1-v_1, ..., bn−vnb_n-v_n]^T||^2=(b1−v1)2+(b2−v2)22 = (b_1-v_1)^2 + (b_2-v_2)^2 + ... + (bn−vnb_n-v_n)^2.

  2. Audio
    Observation

    The speaker says, 'the length squared of this is just going to be ...' and lists the squared terms.

Formula
Explanation

The lecture converts the geometric minimization of length into an algebraic minimization of a sum of squares. By squaring the residual norm, the objective becomes the explicit quadratic expression in the componentwise errors, which motivates the terminology 'least squares'.

Formula
∥[b1−v1b2−v2⋮bn−vn]∥2=(b1−v1)2+(b2−v2)2+⋯+(bn−vn)2\left\|\begin{bmatrix}b_1-v_1\\b_2-v_2\\\vdots\\b_n-v_n\end{bmatrix}\right\|^2=(b_1-v_1)^2+(b_2-v_2)^2+\cdots+(b_n-v_n)^2
Conditions
  1. Using the standard Euclidean norm on RnR^n

Prerequisites
  1. Residual vector written componentwise

Terminology: least squares estimate / solution / approximation

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, 'I want to get the least squares estimate here' and explains that this is why it is called the least squares estimate.

  2. Formula
    Observation

    The board adds the label 'least squares' pointing to x*.

Definition
Explanation

After deriving the sum-of-squared-residuals form, the video names the object x* as the least squares estimate, also calling it the least squares solution or least squares approximation for Ax=b. The terminology is tied directly to minimizing the sum of squared component errors.

Formula
x∗ is the least squares estimate for Ax=b.x^*\text{ is the least squares estimate for }Ax=b.
Conditions
  1. The minimization is over Ax* ∈ C(A)C(A) approximating b

Prerequisites
  1. Squared Euclidean length equals sum of squared residuals
  2. Least-squares objective as distance minimization

Least-squares setup when Ax=b has no solution

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board states 'No solution to Ax=b' and introduces x^* with A x^* as close to b as possible.

  2. Audio
    Observation

    The speaker says there is no solution to this, but maybe we can find some x^* where A times x^* is as close to b as possible.

Definition
Explanation

When the exact linear system Ax=b is inconsistent, the video defines a least-squares approach: choose x^* so that A x^* lies in the column space of A and is as close as possible to b.

Formula
No solution to Ax=b;choose x∗ so that Ax∗ is as close to b as possible\text{No solution to } Ax=b;\quad \text{choose } x^* \text{ so that } A x^* \text{ is as close to } b \text{ as possible}
Conditions
  1. Ax=b has no exact solution

  2. A x^* must lie in C(A)C(A)

Minimization objective for least squares

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board writes 'minimize ||b - A x^*||'.

  2. Audio
    Observation

    The speaker says they want to get this vector to be as close to b as possible.

Formula
Explanation

The least-squares problem is posed as minimizing the norm of the difference between b and A x^*.

Formula
min⁡  ∥b−Ax∗∥\min \; \|b - A x^*\|
Conditions
  1. b is fixed

  2. x^* is the variable being chosen

Prerequisites
  1. Least-squares setup when Ax=b has no solution

Closest vector in a subspace is the orthogonal projection

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says the closest vector in any subspace to a vector not in that subspace is the projection.

  2. Formula
    Observation

    The board later writes A x^* = proj_{C(A)C(A)} b.

Method
Explanation

For a vector outside a subspace, the nearest point inside that subspace is its orthogonal projection onto the subspace. In this lesson the subspace is C(A)C(A).

Formula
closest vector in C(A) to b=proj⁡C(A)b\text{closest vector in } C(A) \text{ to } b = \operatorname{proj}_{C(A)} b
Conditions
  1. C(A)C(A) is treated as a subspace

  2. b is not in C(A)C(A)

Prerequisites
  1. Least-squares setup when Ax=b has no solution
Claims and conditions · 15

Equivalence between inconsistency and non-membership in the column space

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker states three equivalent ways of describing the same situation: no weights on the columns reach bb, no linear combination equals bb, and bb is not in the column space of AA.

  2. Formula
    Observation

    The board shows the expanded linear-combination equation and the statement "b⃗\vec{b} is not in the C(A)C(A)".

Proposition
Statement

For the system Ax⃗=b⃗A\vec{x}=\vec{b}, having no solution is equivalent to saying that b⃗\vec{b} is not a linear combination of the columns of AA, equivalently b⃗∉C(A)\vec{b}\notin C(A).

Hypotheses
  1. AA is an n×kn\times k matrix

  2. x⃗∈Rk\vec{x}\in\mathbb{R}^k and b⃗∈Rn\vec{b}\in\mathbb{R}^n

  3. C(A)C(A) denotes the column space of AA

Quantifiers

For the given matrix AA and vector b⃗\vec{b}, the absence of any x⃗\vec{x} satisfying Ax⃗=b⃗A\vec{x}=\vec{b} is equivalent to b⃗∉C(A)\vec{b}\notin C(A).

If b is not in C(A)C(A), then Ax=b has no exact solution

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board explicitly states 'b is not in the C(A)C(A)'.

  2. Diagram
    Observation

    The cyan vector b is drawn outside the purple plane C(A)C(A).

  3. Audio
    Observation

    The spoken explanation continues from the premise that there is no solution to Ax=b.

Proposition
Statement

In the setup shown, because b is not in the column space C(A)C(A), the equation Ax=b has no solution.

Hypotheses
  1. A is n×kn\times k

  2. x∈Rkx \in R^k

  3. b∈Rnb \in R^n

  4. b ∉ C(A)C(A)

Quantifiers

For the displayed A, x, and b in this example setup.

Every product Ax belongs to C(A)C(A)

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, 'any Ax is going to be in your column space' and repeats the statement.

  2. Formula
    Observation

    He writes v = Ax* and places v inside C(A)C(A).

Proposition
Statement

For any x∈Rkx \in R^k, the vector Ax is a member of the column space C(A)C(A).

Hypotheses
  1. A is n×kn\times k

  2. x∈Rkx \in R^k

Quantifiers

For all x in RkR^k.

Least-squares problem as choosing the closest column-space vector to b

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The objective is written as minimize ||b - Ax*||.

  2. Audio
    Observation

    The speaker says he wants Ax* to be as close as possible to b and identifies Ax* with v in C(A)C(A).

Proposition
Statement

Choosing x* to minimize ||b - Ax*|| is equivalent to choosing the vector v = Ax* in C(A)C(A) that is closest to b.

Hypotheses
  1. b∈Rnb \in R^n

  2. A is n×kn\times k

  3. x∗∈Rkx* \in R^k

  4. Distance measured by Euclidean norm

Quantifiers

For the given A and b in the inconsistent-system setup.

Squared residual norm equals sum of squared residuals

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board displays the equality between the squared norm and the sum of squared component differences.

  2. Audio
    Observation

    The speaker verbally derives the same expression as the length squared.

Proposition
Statement

For b,v∈Rnv \in R^n, ||b-v||^2=(b1−v1)2+(b2−v2)22 = (b_1-v_1)^2 + (b_2-v_2)^2 + ... + (bn−vnb_n-v_n)^2.

Hypotheses
  1. b,v∈Rnv \in R^n

  2. Standard Euclidean norm

Quantifiers

For all component indices i=1i = 1,...,n.

Closest-point claim for a subspace

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker states that the closest vector to b that is in the subspace is the projection of b onto the column space.

  2. Diagram
    Observation

    The diagram shows b outside the plane C(A)C(A) and identifies the nearest point in the plane as the projection.

Proposition
Statement

If b is not in a subspace, then the closest vector to b inside that subspace is the orthogonal projection of b onto the subspace.

Hypotheses
  1. There is a subspace, here C(A)C(A)

  2. b is a vector not in that subspace

Quantifiers

For the given vector b and the given subspace C(A)C(A).

Least-squares characterization claim

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board writes and boxes A x^* = proj_{C(A)C(A)} b.

  2. Audio
    Observation

    The speaker says Ax needs to be equal to the projection of b on my column space.

Proposition
Statement

The least-squares solution x^* is characterized by A x^* = proj_{C(A)C(A)} b.

Hypotheses
  1. Ax=b has no exact solution

  2. x^* is chosen to minimize ||b - A x^*||

Quantifiers

For the least-squares solution x^* associated with the inconsistent system Ax=b.

Closest-point characterization of least squares

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says Ax* is as close to b as possible and ties this to projection onto the subspace.

  2. Formula
    Observation

    Left panel: where Ax* is as close to b as possible.

Proposition
Statement

For the least-squares problem, Ax* is the vector in C(A)C(A) that is closest to b.

Hypotheses
  1. x* is a least-squares solution to Ax=b.

  2. Distance is measured by the Euclidean norm.

Quantifiers

for the given matrix A and vector b

Projection residual orthogonality

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says if I take the projection of b ... minus b, I'm going to get this vector ... this vector right here is orthogonal.

  2. Formula
    Observation

    Ax⃗∗=proj⁡C(A)b⃗A\vec x^{*}=\operatorname{proj}_{C(A)}\vec b and Ax⃗∗−b⃗∈C(A)⊥A\vec x^{*}-\vec b \in C(A)^{\perp}.

Proposition
Statement

If Ax⃗∗=proj⁡C(A)b⃗A\vec x^{*}=\operatorname{proj}_{C(A)}\vec b, then Ax⃗∗−b⃗A\vec x^{*}-\vec b is orthogonal to C(A)C(A).

Hypotheses
  1. Projection is orthogonal projection onto C(A)C(A).

Quantifiers

for all vectors in C(A)C(A)

Column-space orthogonal complement identity

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    C(A)⊥=N(AT)C(A)^{\perp}=N(A^T).

  2. Audio
    Observation

    The speaker states this identity directly.

Theorem
Statement

C(A)⊥=N(AT)C(A)^{\perp}=N(A^T).

Hypotheses
  1. A is a real matrix.

Quantifiers

for the matrix A under discussion

Derivation of the normal equations

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    AT(Ax⃗∗−b⃗)=0A^T(A\vec x^{*}-\vec b)=0 and then ATAx⃗∗−ATb⃗=0A^T A\vec x^{*}-A^T\vec b=0.

  2. Audio
    Observation

    The speaker says this times A transpose has got to be equal to zero and then simplifies.

Proposition
Statement

If Ax⃗∗−b⃗∈N(AT)A\vec x^{*}-\vec b \in N(A^T), then ATAx⃗∗=ATb⃗A^T A\vec x^{*}=A^T\vec b.

Hypotheses
  1. x* is a least-squares solution.

  2. Ax⃗∗−b⃗A\vec x^{*}-\vec b belongs to the left null space of A.

Quantifiers

for the least-squares solution x*

The normal equations always have a solution

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says: 'This right here will always have a solution, and this right here is our least squares solution.'

  2. Formula
    Observation

    The referenced equation on screen is ATAxA^T A x^* = ATbA^T b.

Uncertainties
  1. The video states solvability but does not prove it in this clip.

Proposition
Statement

The equation ATAx=ATbA^T A x = A^T b always has a solution, and that solution is the least-squares solution of the original inconsistent system.

Hypotheses
  1. The original system is Ax=bA x = b.

  2. The equation under discussion is the normal equation ATAx=ATbA^T A x = A^T b.

Quantifiers

For the matrix A and vector b under discussion in this lesson segment.

Derivations and proofs · 9

Rewriting Ax⃗=b⃗A\vec{x}=\vec{b} as a column linear combination

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "Let's just expand out A," writes the columns, multiplies by the components of xx, and states that the result is the same equation.

  2. Formula
    Observation

    The board displays [a⃗1 a⃗2 ⋯ a⃗k][x1x2⋮xk]=b⃗[\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}=\vec{b} and then x1a⃗1+x2a⃗2+⋯+xka⃗k=b⃗x_1\vec{a}_1+x_2\vec{a}_2+\cdots+x_k\vec{a}_k=\vec{b}.

Proof
Steps
  1. Expression
    Ax⃗=b⃗A\vec{x}=\vec{b}
    Explanation

    Start from the original matrix equation.

    Justification

    Given setup in the video.

    Shown in the video
  2. Expression
    A=[a⃗1 a⃗2 ⋯ a⃗k]A=[\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]
    Explanation

    Write AA explicitly as its column vectors.

    Justification

    Definition of a matrix by its columns.

    Shown in the video
  3. Expression
    x⃗=[x1x2⋮xk]\vec{x}=\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}
    Explanation

    Write the unknown vector in components.

    Justification

    Standard coordinate representation of a vector in Rk\mathbb{R}^k.

    Shown in the video
  4. Expression
    [a⃗1 a⃗2 ⋯ a⃗k][x1x2⋮xk]=b⃗[\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}=\vec{b}
    Explanation

    Substitute the column form of AA and component form of x⃗\vec{x} into the equation.

    Justification

    Substitution into the original equation.

    Shown in the video
  5. Expression
    x1a⃗1+x2a⃗2+⋯+xka⃗k=b⃗x_1\vec{a}_1+x_2\vec{a}_2+\cdots+x_k\vec{a}_k=\vec{b}
    Explanation

    Evaluate the matrix-vector product as the corresponding weighted sum of columns.

    Justification

    Column interpretation of matrix-vector multiplication.

    Shown in the video
Conclusion

The system Ax⃗=b⃗A\vec{x}=\vec{b} is exactly the statement that some linear combination of the columns of AA equals b⃗\vec{b}.

From geometric closeness to the least-squares objective

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board progresses from minimize ||b - Ax*|| to v = Ax*, then to the residual vector and finally to the sum of squared terms.

  2. Audio
    Observation

    The speaker narrates each step: closeness means length, Ax* is in C(A)C(A), take componentwise differences, then square and sum them.

Intuitive argument
Steps
  1. Expression
    minimize ∥b⃗−Ax⃗∗∥\text{minimize }\|\vec{b}-A\vec{x}^*\|
    Explanation

    Start from the goal of making Ax* as close as possible to b.

    Justification

    Stated directly by the speaker as the definition of 'close' via length.

    Shown in the video
  2. Expression
    v⃗=Ax⃗∗∈C(A)\vec{v}=A\vec{x}^*\in C(A)
    Explanation

    Rename the candidate output vector as v and note that it lies in the column space.

    Justification

    The speaker says Ax is a member of the column space and writes v = Ax*.

    Shown in the video
  3. Expression
    b⃗−v⃗=[b1−v1b2−v2⋮bn−vn]\vec{b}-\vec{v}=\begin{bmatrix}b_1-v_1\\b_2-v_2\\\vdots\\b_n-v_n\end{bmatrix}
    Explanation

    Rewrite the residual as the vector of componentwise differences.

    Justification

    The speaker explicitly takes the difference between each element of b and v.

    Shown in the video
  4. Expression
    ∥[b1−v1b2−v2⋮bn−vn]∥2=(b1−v1)2+(b2−v2)2+⋯+(bn−vn)2\left\|\begin{bmatrix}b_1-v_1\\b_2-v_2\\\vdots\\b_n-v_n\end{bmatrix}\right\|^2=(b_1-v_1)^2+(b_2-v_2)^2+\cdots+(b_n-v_n)^2
    Explanation

    Square the residual norm to obtain a sum of squared errors.

    Justification

    The speaker states that the length squared is the sum of the squared component differences.

    Shown in the video
  5. Expression
    x∗ is the least squares estimate for Ax=bx^*\text{ is the least squares estimate for }Ax=b
    Explanation

    Name the minimizing choice x* as the least squares estimate.

    Justification

    The speaker says this is why the quantity is called the least squares estimate, solution, or approximation.

    Shown in the video
Conclusion

The least-squares problem is motivated as minimizing the sum of squared componentwise residuals between b and the best available vector Ax* in C(A)C(A).

Matrix product Ax as a linear combination of columns

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board shows [a1a2a_1 a_2 ... aka_k][x1x_1 ... xkx_k]^T=bT = b and x1a1x_1 a_1 + ... + xkak=bx_k a_k = b.

  2. Audio
    Observation

    The speaker explains Ax as multiplying a vector in RkR^k by the columns of A.

Visual argument
Steps
  1. Expression
    [a⃗1 a⃗2 ⋯ a⃗k][x1x2⋮xk]=b⃗[\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}=\vec{b}
    Explanation

    Write Ax using the columns of A and the entries of x.

    Justification

    Displayed on the board as the expanded matrix equation.

    Shown in the video
  2. Expression
    x1a⃗1+x2a⃗2+⋯+xka⃗k=b⃗x_1\vec{a}_1+x_2\vec{a}_2+\cdots+x_k\vec{a}_k=\vec{b}
    Explanation

    Interpret the product as a linear combination of the columns of A.

    Justification

    The speaker explains that multiplying by A combines its columns with the coefficients xix_i.

    Shown in the video
Conclusion

Solving Ax=b exactly would require expressing b as a linear combination of the columns of A, i.e. requiring b∈C(A)b \in C(A).

Derivation of the projection characterization of least squares

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board moves from 'minimize ||b - A x^*||' to 'A x^* = proj_{C(A)C(A)} b'.

  2. Audio
    Observation

    The speaker explains that the closest vector in the subspace is the projection, so Ax must equal that projection.

Intuitive argument
Steps
  1. Expression
    No solution to Ax=b\text{No solution to } Ax=b
    Explanation

    Start from an inconsistent linear system.

    Justification

    Stated directly on the board and in the audio.

    Shown in the video
  2. Expression
    min⁡  ∥b−Ax∗∥\min \; \|b - A x^*\|
    Explanation

    Replace exact solvability by the goal of making A x^* as close as possible to b.

    Justification

    Explicitly written as the least-squares objective.

    Shown in the video
  3. Expression
    Ax∗∈C(A)A x^* \in C(A)
    Explanation

    Any vector of the form A x^* lies in the column space of A.

    Justification

    Audio explanation that A times a vector is a linear combination of the column vectors.

    Shown in the video
  4. Expression
    closest vector in C(A) to b=proj⁡C(A)b\text{closest vector in } C(A) \text{ to } b = \operatorname{proj}_{C(A)} b
    Explanation

    Use the geometric fact that the nearest point in a subspace is the orthogonal projection.

    Justification

    Spoken principle referenced from earlier videos and illustrated by the diagram.

    Shown in the video
  5. Expression
    Ax∗=proj⁡C(A)bA x^* = \operatorname{proj}_{C(A)} b
    Explanation

    Therefore the least-squares solution is characterized by equality with the projection.

    Justification

    Combines the previous steps.

    Shown in the video
Conclusion

The least-squares problem is equivalent to requiring A x^* to be the projection of b onto C(A)C(A).

Rearranging the least-squares equation by subtracting b

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board writes A x^* - b = proj_{C(A)C(A)} b - b.

  2. Audio
    Observation

    The speaker says they will subtract b from both sides to see if something interesting appears.

Proof
Steps
  1. Expression
    Ax∗=proj⁡C(A)bA x^* = \operatorname{proj}_{C(A)} b
    Explanation

    Begin with the boxed characterization of the least-squares solution.

    Justification

    Previously derived and written on the board.

    Shown in the video
  2. Expression
    Ax∗−b=proj⁡C(A)b−bA x^* - b = \operatorname{proj}_{C(A)} b - b
    Explanation

    Subtract b from both sides.

    Justification

    Valid algebraic operation applied to both sides of an equality.

    Shown in the video
Conclusion

The equation is rewritten in residual form, setting up the next step of the lesson.

Geometric-to-algebraic derivation of the normal equations

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker moves from orthogonality to membership in C(A)⊥C(A)^{\perp}, then to N(AT)N(A^T), then multiplies by ATA^T and expands.

  2. Formula
    Observation

    Visible chain: Ax⃗∗=proj⁡C(A)b⃗A\vec x^{*}=\operatorname{proj}_{C(A)}\vec b; Ax⃗∗−b⃗∈C(A)⊥A\vec x^{*}-\vec b \in C(A)^{\perp}; C(A)⊥=N(AT)C(A)^{\perp}=N(A^T); Ax⃗∗−b⃗∈N(AT)A\vec x^{*}-\vec b \in N(A^T); AT(Ax⃗∗−b⃗)=0A^T(A\vec x^{*}-\vec b)=0; ATAx⃗∗−ATb⃗=0A^T A\vec x^{*}-A^T\vec b=0.

Uncertainties
  1. The clip does not prove C(A)⊥=N(AT)C(A)^{\perp}=N(A^T); it invokes it as known from earlier videos.

Proof
Steps
  1. Expression
    Ax⃗∗=proj⁡C(A)b⃗A\vec x^{*}=\operatorname{proj}_{C(A)}\vec b
    Explanation

    Start from the geometric characterization of the least-squares fit as the orthogonal projection of b onto the column space.

    Justification

    Stated explicitly on the board and in the audio.

    Shown in the video
  2. Expression
    Ax⃗∗−b⃗∈C(A)⊥A\vec x^{*}-\vec b \in C(A)^{\perp}
    Explanation

    The difference between the projected point and b is orthogonal to the subspace C(A)C(A).

    Justification

    Definition/property of orthogonal projection, as explained by the speaker.

    Shown in the video
  3. Expression
    C(A)⊥=N(AT)C(A)^{\perp}=N(A^T)
    Explanation

    Replace the orthogonal complement of the column space by the left null space.

    Justification

    Standard fundamental-subspaces identity invoked in the video.

    Shown in the video
  4. Expression
    Ax⃗∗−b⃗∈N(AT)A\vec x^{*}-\vec b \in N(A^T)
    Explanation

    Therefore the residual vector lies in the null space of ATA^T.

    Justification

    Substitution using the previous identity.

    Shown in the video
  5. Expression
    AT(Ax⃗∗−b⃗)=0⃗A^T(A\vec x^{*}-\vec b)=\vec 0
    Explanation

    Multiplying a vector in N(AT)N(A^T) by ATA^T yields the zero vector.

    Justification

    Definition of null space.

    Shown in the video
  6. Expression
    ATAx⃗∗−ATb⃗=0⃗A^T A\vec x^{*}-A^T\vec b=\vec 0
    Explanation

    Distribute ATA^T over the difference to obtain the normal equations in homogeneous form.

    Justification

    Linearity/distributivity of matrix multiplication.

    Shown in the video
  7. Expression
    ATAx⃗∗=ATb⃗A^T A\vec x^{*}=A^T\vec b
    Explanation

    Move ATbA^T b to the other side to isolate the standard normal-equation form.

    Justification

    Algebraic rearrangement.

    Derived from the video
Conclusion

The least-squares solution satisfies the normal equations ATAxA^T A x* = ATbA^T b.

Derivation of the normal equations from Ax=b

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker explains multiplying both sides of Ax=b by ATA^T.

  2. Formula
    Observation

    The board shows ATAx=ATbA^T A x = A^T b and then labels the solution as x^*.

Proof
Steps
  1. Expression
    Ax⃗=b⃗A\vec x=\vec b
    Explanation

    Begin with the original linear system that the video says has no exact solution.

    Justification

    Stated directly in the clip as the starting problem.

    Shown in the video
  2. Expression
    AT(Ax⃗)=ATb⃗A^{\mathrm T}(A\vec x)=A^{\mathrm T}\vec b
    Explanation

    Left-multiply both sides by ATA^T.

    Justification

    Explicitly performed and narrated in the clip.

    Shown in the video
  3. Expression
    ATAx⃗=ATb⃗A^{\mathrm T}A\vec x=A^{\mathrm T}\vec b
    Explanation

    Use associativity of matrix multiplication to rewrite the left side.

    Justification

    Standard matrix algebra; the simplified form is what is written on the board.

    Derived from the video
  4. Expression
    ATAx⃗∗=ATb⃗A^{\mathrm T}A\vec x^*=A^{\mathrm T}\vec b
    Explanation

    Rename the solution of the new equation as the least-squares solution x^*.

    Justification

    The speaker explicitly identifies this solution as the least-squares solution.

    Shown in the video
Conclusion

The least-squares solution can be found by solving the normal equations ATAxA^T A x^* = ATbA^T b instead of directly computing the projection.

Derivation of the normal equations from the orthogonality condition

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board displays the chain ATA^T(b - A x^*) = 0, then ATAxA^T A x^* - ATb=0A^T b = 0, then the boxed ATAxA^T A x^* = ATbA^T b.

  2. Audio
    Observation

    The speaker identifies one side as a matrix and the other as a vector while discussing the final equation.

Proof
Steps
  1. Expression
    AT(b−Ax∗)=0A^T(b - A x^*) = 0
    Explanation

    Start from the condition that the residual is annihilated by ATA^T.

    Justification

    This is the orthogonality/residual condition written directly on the board.

    Shown in the video
  2. Expression
    ATAx∗−ATb=0A^T A x^* - A^T b = 0
    Explanation

    Distribute ATA^T over the difference and move terms to one side.

    Justification

    Standard matrix algebra applied to the displayed equation.

    Shown in the video
  3. Expression
    ATAx∗=ATbA^T A x^* = A^T b
    Explanation

    Add ATbA^T b to both sides to isolate the unknown x^* term.

    Justification

    Algebraic rearrangement of the previous line, matching the boxed equation on the board.

    Shown in the video
Conclusion

The least-squares condition is equivalently encoded by the normal equations ATAxA^T A x^* = ATbA^T b.

Geometric derivation linking projection, orthogonal complement, and null space

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The zoomed-out board shows A x^* = proj_{C(A)C(A)} b, A x^* - b∈C(A)⊥b \in C(A)^\perp, C(A)⊥=N(AT)C(A)^\perp = N(A^T), and A x^* - b∈N(AT)b \in N(A^T).

  2. Diagram
    Observation

    A geometric sketch shows b, its projection onto a plane representing C(A)C(A), and the perpendicular residual direction.

Visual argument
Steps
  1. Expression
    Ax∗=proj⁡C(A)bA x^* = \operatorname{proj}_{C(A)} b
    Explanation

    The closest point to b in C(A)C(A) is its orthogonal projection onto that subspace.

    Justification

    Displayed as the defining identity for the least-squares image vector.

    Shown in the video
  2. Expression
    Ax∗−b∈C(A)⊥A x^* - b \in C(A)^\perp
    Explanation

    The residual from b to its projection is perpendicular to the subspace C(A)C(A).

    Justification

    This is the geometric meaning of orthogonal projection, also written on the board.

    Shown in the video
  3. Expression
    C(A)⊥=N(AT)C(A)^\perp = N(A^T)
    Explanation

    The orthogonal complement of the column space is the null space of A transpose.

    Justification

    Explicit identity shown on the board after the zoom-out.

    Shown in the video
  4. Expression
    Ax∗−b∈N(AT)A x^* - b \in N(A^T)
    Explanation

    Therefore the residual also lies in N(AT)N(A^T).

    Justification

    Substitution using the previous two lines.

    Derived from the video
Conclusion

The projection picture, the orthogonal-complement statement, and the null-space statement are consistent descriptions of the same least-squares geometry.

Visual events · 12

Geometric picture of b⃗∉C(A)\vec{b}\notin C(A)

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says, "let's see if we can visualize it a bit," draws the column space of AA, assumes it is a plane in Rn\mathbb{R}^n, marks the origin, and draws bb outside the plane.

  2. Diagram
    Observation

    A purple parallelogram-like plane is drawn and labeled C(A)C(A); a cyan arrow from the origin points upward outside the plane and is labeled b⃗\vec{b}.

Uncertainties
  1. The drawing is schematic; the ambient dimension nn is not fixed beyond being at least large enough to contain a plane.

Objects
  1. Purple plane representing C(A)C(A)

  2. Origin point

  3. Cyan vector arrow labeled b⃗\vec{b}

  4. Labels C(A)C(A) and b⃗\vec{b}

Changes
  1. First the plane for C(A)C(A) is drawn.

  2. Then the origin is marked.

  3. Finally the vector b⃗\vec{b} is drawn starting at the origin and extending outside the plane.

Invariants
  1. The plane is meant to represent all linear combinations of the columns of AA.

  2. The vector b⃗\vec{b} remains outside that plane throughout the sketch.

Interpretation

The visual encodes the algebraic claim that if b⃗\vec{b} is not in the column space, then no linear combination of the columns of AA can equal b⃗\vec{b}.

Introduction of a candidate approximate solution x⃗∗\vec{x}^*

Approximate timing
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker asks what if we can find a solution that gets us close to the target and says, "I want to find some x star for now where A times x star is..."

  2. Formula
    Observation

    Yellow handwritten x⃗∗\vec{x}^* and the partial phrase "where Ax⃗∗A\vec{x}^*" appear at the lower left.

Uncertainties
  1. The defining condition for x⃗∗\vec{x}^* is not completed within the supplied clip.

Objects
  1. Symbol x⃗∗\vec{x}^*

  2. Partial expression Ax⃗∗A\vec{x}^*

Changes
  1. A new symbol x⃗∗\vec{x}^* is added after the discussion of the inconsistent system.

  2. The speaker begins to define a property of Ax⃗∗A\vec{x}^* but the clip ends before completion.

Invariants
  1. The earlier equation Ax⃗=b⃗A\vec{x}=\vec{b} remains on screen as the inconsistent target problem.

  2. The column-space sketch remains visible while x⃗∗\vec{x}^* is introduced.

Interpretation

This visual transition signals a move from exact solvability to an approximation problem, but the precise criterion for x⃗∗\vec{x}^* is not yet shown in this segment.

Geometric picture of b outside C(A)C(A) and v inside C(A)C(A)

Clear evidence
Shown in the video
Evidence
  1. Diagram
    Observation

    A purple parallelogram labeled C(A)C(A) is drawn on the right, with a cyan vector b starting from the same origin and ending outside the plane.

  2. Formula
    Observation

    Later a white vector labeled v = Ax* is drawn inside the purple plane.

Objects
  1. Purple plane labeled C(A)C(A)

  2. Cyan vector labeled b

  3. White vector labeled v = Ax*

  4. Origin point common to the drawn vectors

Changes
  1. At first only b and the plane C(A)C(A) are shown.

  2. Around 52-80 seconds, the speaker introduces v = Ax* and draws it inside C(A)C(A).

  3. The residual discussion then relates b and v by componentwise subtraction.

Invariants
  1. The plane C(A)C(A) stays fixed as the set of attainable Ax vectors.

  2. b remains outside C(A)C(A) throughout the motivating picture.

Interpretation

The diagram encodes the core idea that exact equality Ax=b is impossible here, but one can still choose a vector v in C(A)C(A) that is nearest to b.

Writing the minimization objective

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The phrase 'minimize ||b - Ax*||' is written progressively on the board.

  2. Audio
    Observation

    The speaker says he wants to minimize the length of b minus Ax*.

Objects
  1. Text 'minimize'

  2. Expression ||b - Ax*||

Changes
  1. The word 'minimize' is written first.

  2. Then the norm bars and the residual expression b - Ax* are added.

Invariants
  1. The objective remains a distance-minimization statement.

Interpretation

This visual step converts the verbal idea of 'as close as possible' into a precise optimization problem.

Expanding the residual into coordinates and squares

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board expands the norm into a column vector of differences and then into a sum of squared terms.

Objects
  1. Column vector [b1−v1b_1-v_1, ..., bn−vnb_n-v_n]^T

  2. Sum (b1−v1b_1-v_1)^2 + ... + (bn−vnb_n-v_n)^2

Changes
  1. First the residual vector is written componentwise.

  2. Then its squared norm is rewritten as a sum of squares.

Invariants
  1. The underlying quantity being minimized is still the distance between b and v = Ax*.

Interpretation

This visual expansion is what motivates the name 'least squares'.

Two-part whiteboard layout linking algebra and geometry

Clear evidence
Shown in the video
Evidence
  1. Diagram
    Observation

    Left side contains the algebraic setup for Ax=b and the minimization expression; right side contains a geometric sketch of b outside the plane C(A)C(A).

Objects
  1. Matrix equation Ax=b

  2. Minimization expression ||b - A x^*||

  3. Purple plane labeled C(A)C(A)

  4. Cyan vector labeled b

  5. Yellow vector labeled A x^*

Changes
  1. The right-side diagram gains the label proj_{C(A)C(A)} b.

  2. A new boxed equation A x^* = proj_{C(A)C(A)} b is written below the diagram.

  3. The boxed equation is later rewritten as A x^* - b = proj_{C(A)C(A)} b - b.

Invariants
  1. The left-side setup remains visible throughout the clip.

  2. The plane C(A)C(A) remains the subspace of interest.

  3. b remains outside C(A)C(A) in the geometric picture.

Interpretation

The visual organization ties the algebraic least-squares objective to the geometric idea of projecting b onto the column space.

Geometric picture of projection onto C(A)C(A)

Clear evidence
Shown in the video
Evidence
  1. Diagram
    Observation

    A cyan vector b points outside a purple plane labeled C(A)C(A); a yellow vector A x^* lies in the plane; the projection label is added near the plane.

  2. Audio
    Observation

    The speaker says the closest vector in the subspace is the projection of b onto the column space.

Uncertainties
  1. The exact arrowhead placement of the projection annotation is approximate from the sampled frames.

Objects
  1. Vector b

  2. Plane C(A)C(A)

  3. Vector A x^*

  4. Projection annotation proj_{C(A)C(A)} b

Changes
  1. The diagram shifts from showing only b and A x^* to explicitly identifying the nearest point in C(A)C(A) as proj_{C(A)C(A)} b.

Invariants
  1. b stays outside the plane.

  2. A x^* stays inside the plane.

Interpretation

The figure illustrates that minimizing ||b - A x^*|| means choosing the point in C(A)C(A) closest to b, namely the orthogonal projection.

Three-region whiteboard organization

Clear evidence
Shown in the video
Evidence
  1. Diagram
    Observation

    The screen is split into a left algebraic setup, a central geometric sketch, and a right derivation column.

  2. Animation
    Observation

    New formulas are added sequentially on the right as the explanation proceeds.

Objects
  1. left panel with Ax=b and least-squares objective

  2. central purple plane C(A)C(A) with vectors b, proj_{C(A)C(A)}b, and Ax*

  3. right panel accumulating implication chain to normal equations

Changes
  1. Right-side formulas are written step by step from projection identity to ATAx∗−ATb=0A^T A x* - A^T b = 0.

Invariants
  1. The left least-squares setup and central projection diagram remain visible while the derivation is built.

Interpretation

The visual layout links the optimization statement, the geometric projection picture, and the algebraic normal-equation derivation.

Orthogonal projection picture for least squares

Clear evidence
Shown in the video
Evidence
  1. Diagram
    Observation

    A purple plane labeled C(A)C(A) contains the yellow vector Ax⃗∗A\vec x^{*}; b is drawn outside the plane; an orange/yellow arrow connects b down to the plane.

  2. Audio
    Observation

    The speaker points to the vector going straight down onto the plane and calls it orthogonal to the column space.

Uncertainties
  1. Exact color naming varies slightly between orange and yellow in the residual arrow across frames.

Objects
  1. plane C(A)C(A)

  2. vector b outside the plane

  3. vector Ax⃗∗A\vec x^{*} inside the plane

  4. residual arrow from b to Ax⃗∗A\vec x^{*}

Changes
  1. The residual is emphasized as the perpendicular connector between b and its projection.

Invariants
  1. Ax⃗∗A\vec x^{*} stays in C(A)C(A) while b stays outside it in the sketch.

Interpretation

The diagram encodes that the best approximation to b from C(A)C(A) is obtained by dropping perpendicularly to the plane.

Color-coded summary board of least squares theory

Clear evidence
Shown in the video
Evidence
  1. Diagram
    Observation

    The blackboard contains multiple color-coded formulas: the minimization statement, the projection identity, the orthogonality relation, and the final normal equation.

Objects
  1. minimize ||b - A x^*||

  2. expanded squared-sum expression

  3. A x^* = proj_{C(A)C(A)} b

  4. ATA^T(Ax^*-b)=0

  5. ATAxA^T A x^* = ATbA^T b

  6. Ax=b with 'no solution' annotation

Changes
  1. The lower-left area is rewritten during the clip to restate the problem setup.

  2. The final normal equation is emphasized as the practical solution method.

Invariants
  1. The overall topic remains least-squares approximation for an inconsistent linear system.

  2. The projection-based geometric interpretation stays visible while the algebraic method is introduced.

Interpretation

The visual layout connects three levels of the same idea: optimization definition, geometric projection characterization, and algebraic normal-equation computation.

Close-up of the least-squares formula board

Clear evidence
Shown in the video
Evidence
  1. Diagram
    Observation

    Close-up of a black digital board with multicolored handwritten formulas about least squares, including minimize ||b - A x^*||, A x^* = proj_{C(A)C(A)} b, and ATAxA^T A x^* = ATbA^T b.

  2. Animation
    Observation

    A cursor moves among the formulas while the speaker talks.

Objects
  1. handwritten formulas

  2. cursor

  3. boxed normal equations

  4. projection identity

  5. norm expression

Changes
  1. Cursor points to different parts of the board as the speaker discusses matrix-versus-vector structure and minimization.

Invariants
  1. The board content remains the same during the close-up.

  2. The central identities A x^* = proj_{C(A)C(A)} b and ATAxA^T A x^* = ATbA^T b stay visible.

Interpretation

The close-up visually organizes the least-squares argument around three linked ideas: minimize residual norm, project onto C(A)C(A), and solve the normal equations.

Zoom-out revealing the full conceptual board

Clear evidence
Shown in the video
Evidence
  1. Animation
    Observation

    The camera zooms out to reveal a much larger board.

  2. Diagram
    Observation

    Newly visible regions include the column-combination equations at upper left and a geometric sketch with b, proj_{C(A)C(A)} b, and a plane labeled C(A)C(A).

Objects
  1. full whiteboard

  2. column-combination equations

  3. geometric sketch of b and its projection

  4. plane labeled C(A)C(A)

  5. subspace identities on the right

Changes
  1. Additional formulas appear: [a1a2a_1 a_2 ... aka_k][x1x_1 ... xkx_k]^T=bT = b and x1a1+x2a2x_1 a_1 + x_2 a_2 + ... + xkak=bx_k a_k = b.

  2. A geometric diagram becomes visible showing b, its projection onto C(A)C(A), and the perpendicular residual.

  3. Right-side subspace relations C(A)⊥=N(AT)C(A)^\perp = N(A^T) and A x^* - b∈N(AT)b \in N(A^T) become visible.

Invariants
  1. The previously seen least-squares formulas remain part of the same overall argument.

  2. The projection identity A x^* = proj_{C(A)C(A)} b continues to anchor the geometry.

Interpretation

The zoom-out connects the algebraic normal equations to the geometric picture of projecting b onto the column space and interpreting the residual as orthogonal to that space.

Misconceptions · 11

Stopping at “no solution” from row reduction

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says that up until now they would make an augmented matrix, put it in reduced row echelon form, get a line saying zero equals one, and conclude there is no solution and nothing more can be done.

Misconception

Treating a row-reduction contradiction such as 0=10=1 as the end of the analysis, with no further useful objective available.

Clarification

The video explicitly contrasts that old endpoint with a new goal: even if no exact solution exists, one may still seek a solution that gets close to the target.

Do not confuse approximate solution with exact solution

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker contrasts 'No solution to Ax=b' with making Ax* 'as close to b as possible'.

  2. Diagram
    Observation

    b is shown outside C(A)C(A), so exact equality is visually impossible.

Misconception

One might think that because we are minimizing ||b - Ax*||, the resulting x* solves Ax=b exactly.

Clarification

In this setup Ax=b has no exact solution because b ∉ C(A)C(A); x* only gives the best approximation within C(A)C(A).

Do not treat Ax* as an arbitrary vector in RnR^n

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker repeatedly says any Ax is in the column space.

  2. Formula
    Observation

    He writes v = Ax* inside C(A)C(A).

Misconception

One might forget that the candidate vector Ax* is constrained to lie in C(A)C(A).

Clarification

Because Ax* is a linear combination of the columns of A, the optimization searches only among vectors in the column space.

Minimizing length and minimizing squared length are being linked carefully

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board first shows ||b - Ax*|| and then rewrites the squared version as a sum of squares.

  2. Audio
    Observation

    The speaker says, 'Let me take the length squared actually.'

Misconception

One might think the lecture abruptly changed the problem when moving from ||b - Ax*|| to the sum of squared residuals.

Clarification

The video uses the squared norm because it has the explicit sum-of-squares form that motivates the terminology 'least squares'.

Inconsistency does not prevent approximation

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker contrasts 'no solution to this' with finding x^* so that A x^* is as close to b as possible.

  2. Formula
    Observation

    The board keeps 'No solution to Ax=b' while introducing the least-squares objective.

Misconception

If Ax=b has no solution, one might think nothing useful can be said about x.

Clarification

The video explicitly replaces exact solvability with a least-squares approximation problem: choose x^* so that A x^* is as close as possible to b.

Knowing the projection formula does not make computation easy

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says it is still pretty hard to find because taking A times the inverse of A transpose A times A transpose is hard.

Uncertainties
  1. This is a caution stated by the speaker rather than a fully worked counterexample.

Misconception

One might assume that once the projection formula is known, the least-squares solution is straightforward to compute directly.

Clarification

The video states that the explicit projection-matrix route is hard to use here, motivating an easier derivation.

Least squares is not claiming Ax=b is solvable

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    Top-left text reads No solution to Ax=b.

  2. Audio
    Observation

    The whole clip reframes the problem as finding Ax* as close to b as possible.

Misconception

One might think the appearance of x* means there is an exact solution to Ax=b.

Clarification

The video explicitly starts from the case of no solution and replaces exact equality by closest approximation in the column space.

Sign convention for the residual

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    Both Ax⃗∗−b⃗A\vec x^{*}-\vec b and proj⁡C(A)b⃗−b⃗\operatorname{proj}_{C(A)}\vec b-\vec b appear.

  2. Audio
    Observation

    The speaker says you could say b plus this vector is equal to my projection of b onto my subspace.

Misconception

Students may confuse the direction of the perpendicular error vector.

Clarification

Here the residual is taken as Ax*-b, equivalently proj_{C(A)C(A)}b-b, and the key property is that this vector is orthogonal to C(A)C(A); reversing the sign would still be orthogonal but changes the written expression.

Do not confuse the normal-equation solution with an exact solution of Ax=b

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says the solution to the multiplied equation will not be the same as the solution to the original equation.

Misconception

One might think that solving ATAx=ATbA^T A x = A^T b gives an exact solution of the original inconsistent system Ax=b.

Clarification

The video explicitly distinguishes them: the original system has no solution, while the normal equations have a solution that is the least-squares solution.

Confusing inconsistency with hopelessness

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board explicitly writes Ax=b with "no solution" nearby, then separately writes the least-squares minimization and normal equations.

  2. Audio
    Observation

    The speaker says this gives our best shot at finding a solution to Ax=b after noting the system may not be solvable exactly.

Misconception

If Ax=b has no exact solution, one might think nothing useful can be said about x.

Clarification

The board distinguishes the inconsistent system Ax=b from the least-squares problem, which still produces a meaningful x^* by minimizing ||b - A x^*||.

Overlooking dimensional structure in the normal equations

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    Speaker: "Notice, this is some matrix, and then this right here is some vector."

Misconception

One might treat ATAxA^T A x^* = ATbA^T b as just another symbolic equation without noticing the types of the objects involved.

Clarification

The speaker explicitly points out that one side is a matrix acting on a vector and the other side is a vector, emphasizing the equation's linear-algebraic structure.

Concept relations · 24

Matrix equation setup and dimension matching → Matrix-vector product as a linear combination of columns

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    After stating x⃗∈Rk\vec{x}\in\mathbb{R}^k and b⃗∈Rn\vec{b}\in\mathbb{R}^n, the speaker expands AA into columns and rewrites the equation as a linear combination.

  2. Formula
    Observation

    The board moves from Ax⃗=b⃗A\vec{x}=\vec{b} to [a⃗1 a⃗2 ⋯ a⃗k][x1x2⋮xk]=b⃗[\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}=\vec{b} and then to x1a⃗1+⋯+xka⃗k=b⃗x_1\vec{a}_1+\cdots+x_k\vec{a}_k=\vec{b}.

Application
Explanation

The dimension setup makes the column expansion meaningful: because AA has kk columns, x⃗\vec{x} supplies exactly kk weights for those columns.

Matrix-vector product as a linear combination of columns → Interpretation of an inconsistent system via column space

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker uses the linear-combination form to explain what “no solution” means and then names the column space.

  2. Formula
    Observation

    The linear-combination equation remains on screen while the statement "b⃗\vec{b} is not in the C(A)C(A)" is added.

Proof dependency
Explanation

The interpretation of inconsistency depends on first rewriting Ax⃗=b⃗A\vec{x}=\vec{b} as a statement about linear combinations of the columns of AA.

Interpretation of an inconsistent system via column space → Geometric picture of b⃗∉C(A)\vec{b}\notin C(A)

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    After saying bb is not in the column space of AA, the speaker says, "let's see if we can visualize it a bit" and draws the column space and bb outside it.

  2. Diagram
    Observation

    The symbolic statement "b⃗\vec{b} is not in the C(A)C(A)" is paired with a purple plane labeled C(A)C(A) and a cyan vector b⃗\vec{b} outside the plane.

Application
Explanation

The geometric sketch is a direct visualization of the algebraic claim that b⃗∉C(A)\vec{b}\notin C(A).

Interpretation of an inconsistent system via column space → Introduction of a candidate approximate solution x⃗∗\vec{x}^*

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker contrasts the old conclusion of “no solution” with the question of finding something that gets us close, then introduces x⃗∗\vec{x}^*.

  2. Formula
    Observation

    The board retains the inconsistent-system statements and adds x⃗∗\vec{x}^* with the partial condition "where Ax⃗∗A\vec{x}^*".

Uncertainties
  1. The exact approximation criterion for x⃗∗\vec{x}^* is not completed in this clip.

Generalizes
Explanation

Once exact solvability fails, the lesson pivots toward a broader objective: find an approximate solution rather than stop at inconsistency.

Inconsistent system Ax=b when b is outside C(A)C(A) → Least-squares objective as distance minimization

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker moves from 'No solution to Ax=b' to defining x* by minimizing ||b - Ax*||.

  2. Formula
    Observation

    Both statements appear on the same board.

Prerequisite
Explanation

The least-squares objective is introduced precisely because the exact system Ax=b is inconsistent in the pictured case.

Every Ax lies in the column space C(A)C(A) → Least-squares objective as distance minimization

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says Ax is in the column space and then explains that we want this vector to be as close as possible to b.

  2. Formula
    Observation

    v = Ax* is written inside C(A)C(A).

Proof dependency
Explanation

The geometric meaning of minimizing ||b - Ax*|| depends on recognizing that Ax* ranges over C(A)C(A).

Residual vector written componentwise → Least-squares objective as distance minimization

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The residual vector is written directly under the minimization objective.

Application
Explanation

The componentwise residual formula is the coordinate expression of the abstract norm-minimization objective.

Squared Euclidean length equals sum of squared residuals → Terminology: least squares estimate / solution / approximation

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says the sum-of-squares form is why the method is called least squares.

  2. Formula
    Observation

    The board shows the sum of squared residuals and then labels x* as least squares.

Proof dependency
Explanation

The terminology 'least squares' is justified by the explicit sum-of-squared-residuals expression.

Least-squares objective as distance minimization → Squared Euclidean length equals sum of squared residuals

Clear evidence
Shown in the video
Evidence
  1. Diagram
    Observation

    The plane C(A)C(A), vector b, and vector v = Ax* are drawn geometrically.

  2. Formula
    Observation

    The same idea is rewritten algebraically as a sum of squared component errors.

Equivalent
Explanation

The geometric minimization of distance and the algebraic minimization of the sum of squared residuals describe the same least-squares problem.

Minimization objective for least squares → Closest vector in a subspace is the orthogonal projection

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker derives the least-squares condition by invoking the closest-vector/projection principle.

  2. Formula
    Observation

    The board connects ||b - A x^*|| with A x^* = proj_{C(A)C(A)} b.

Application
Explanation

The minimization objective is solved geometrically by applying the fact that the closest vector in a subspace is the orthogonal projection.

Closest vector in a subspace is the orthogonal projection → Characterizing equation for the least-squares solution

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board writes A x^* = proj_{C(A)C(A)} b after discussing projection.

  2. Audio
    Observation

    The speaker says Ax needs to be equal to the projection of b on the column space.

Application
Explanation

The projection principle yields the defining equation for the least-squares solution.

Direct projection formula is computationally inconvenient here → Algebraic rearrangement of the least-squares equation

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says the direct projection formula is hard, so let's see if we can find an easier way.

  2. Formula
    Observation

    The board then subtracts b from both sides of A x^* = proj_{C(A)C(A)} b.

Contrast
Explanation

The clip contrasts the cumbersome direct projection-matrix method with an algebraic rearrangement intended to lead to an easier derivation.

Find an answer · 25

Why does Ax⃗=b⃗A\vec{x}=\vec{b} force x⃗∈Rk\vec{x}\in\mathbb{R}^k and b⃗∈Rn\vec{b}\in\mathbb{R}^n when AA is n×kn\times k?

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker explains that xx must be in Rk\mathbb{R}^k because there are kk columns, and bb is in Rn\mathbb{R}^n.

  2. Formula
    Observation

    x⃗∈Rk\vec{x}\in\mathbb{R}^k and b⃗∈Rn\vec{b}\in\mathbb{R}^n are written on the board.

Knowledge points
  1. Matrix equation setup and dimension matching
  2. A
  3. x⃗\vec{x}
  4. b⃗\vec{b}
  5. n, k

How is Ax⃗=b⃗A\vec{x}=\vec{b} rewritten as a linear combination of the columns of AA?

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker expands AA into columns and says the product is the same as a weighted sum of those columns.

  2. Formula
    Observation

    The board shows [a⃗1 a⃗2 ⋯ a⃗k][x1x2⋮xk]=b⃗[\vec{a}_1\ \vec{a}_2\ \cdots\ \vec{a}_k]\begin{bmatrix}x_1\\x_2\\\vdots\\x_k\end{bmatrix}=\vec{b} and x1a⃗1+⋯+xka⃗k=b⃗x_1\vec{a}_1+\cdots+x_k\vec{a}_k=\vec{b}.

Knowledge points
  1. Matrix-vector product as a linear combination of columns
  2. Rewriting Ax⃗=b⃗A\vec{x}=\vec{b} as a column linear combination
  3. a⃗1\vec{a}_1,a⃗2\vec{a}_2,…\ldots,a⃗k\vec{a}_k
  4. x1x_1,x2x_2,…\ldots,xkx_k

What does it mean algebraically and geometrically when Ax⃗=b⃗A\vec{x}=\vec{b} has no solution?

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says no set of weights on the columns reaches bb, equivalently no linear combination equals bb, and then says bb is not in the column space of AA.

  2. Formula
    Observation

    The board states "No solution to Ax⃗=b⃗A\vec{x}=\vec{b}" and "b⃗\vec{b} is not in the C(A)C(A)".

Knowledge points
  1. Interpretation of an inconsistent system via column space
  2. Equivalence between inconsistency and non-membership in the column space
  3. C(A)C(A)

Why does Ax=b have no solution when b is not in C(A)C(A)?

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    Board text says 'No solution to Ax=b' and 'b is not in the C(A)C(A)'.

Knowledge points
  1. Inconsistent system Ax=b when b is outside C(A)C(A)

What does 'as close as possible' mean in the least-squares setup?

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says 'when I say close, I'm talking about length'.

  2. Formula
    Observation

    The objective is written as minimize ||b - Ax*||.

Knowledge points
  1. Least-squares objective as distance minimization

Why is Ax* restricted to the column space C(A)C(A)?

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says any Ax is in the column space.

  2. Formula
    Observation

    v = Ax* is written inside C(A)C(A).

Knowledge points
  1. Every Ax lies in the column space C(A)C(A)

How does minimizing ||b - Ax*|| become minimizing a sum of squared residuals?

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    The board expands ||b-v|| into component differences and then squares them.

Knowledge points
  1. Residual vector written componentwise
  2. Squared Euclidean length equals sum of squared residuals

Why is x* called the least squares estimate or least squares solution?

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    The speaker says this is why it is called the least squares estimate.

  2. Formula
    Observation

    The label 'least squares' is added next to x*.

Knowledge points
  1. Terminology: least squares estimate / solution / approximation

Why introduce x^* and minimize ||b - A x^*|| when Ax=b has no solution?

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    Board shows 'No solution to Ax=b' and 'minimize ||b - A x^*||'.

  2. Audio
    Observation

    Speaker explains finding x^* as close as possible when no exact solution exists.

Knowledge points
  1. Least-squares setup when Ax=b has no solution
  2. Minimization objective for least squares

What geometric fact makes the closest vector in C(A)C(A) to b equal to proj_{C(A)C(A)} b?

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    Speaker says the closest vector in the subspace is the projection.

  2. Diagram
    Observation

    Diagram shows b outside C(A)C(A) and the nearest point in the plane.

Knowledge points
  1. Closest vector in a subspace is the orthogonal projection
  2. Closest-point claim for a subspace

How is the least-squares solution x^* characterized in terms of projection onto the column space?

Clear evidence
Shown in the video
Evidence
  1. Formula
    Observation

    Boxed equation A x^* = proj_{C(A)C(A)} b.

  2. Audio
    Observation

    Speaker says Ax needs to be equal to the projection of b on the column space.

Knowledge points
  1. Characterizing equation for the least-squares solution
  2. Least-squares characterization claim

Why subtract b from both sides of A x^* = proj_{C(A)C(A)} b instead of using the projection matrix directly?

Clear evidence
Shown in the video
Evidence
  1. Audio
    Observation

    Speaker says the direct projection formula is hard and proposes an easier way.

  2. Formula
    Observation

    Board rewrites to A x^* - b = proj_{C(A)C(A)} b - b.

Knowledge points
  1. Direct projection formula is computationally inconvenient here
  2. Algebraic rearrangement of the least-squares equation
Coverage and review notes

Covered · Audio and formulas introduce the n×kn\times k matrix equation and the required vector dimensions.

Covered · The speaker states that the system has no solution and writes this condition on the board.

Covered · The matrix is expanded into columns and the equation is rewritten as a linear combination.

Covered · The speaker interprets inconsistency as nonexistence of suitable weights and as b⃗∉C(A)\vec{b}\notin C(A).

Covered · A schematic plane for C(A)C(A) and an outside vector b⃗\vec{b} are drawn to visualize the algebraic statement.

Covered · The speaker contrasts the previous row-reduction endpoint with the new question of getting close to a solution.

Covered · A new candidate x⃗∗\vec{x}^* is introduced for an approximate solution, but the defining condition is cut off before completion.

Covered · Opening board state and spoken setup: Ax=b has no solution because b is outside C(A)C(A); matrix equation is rewritten as a column combination.

Covered · Speaker defines closeness as length and writes the minimization objective ||b - Ax*||.

Covered · Ax* is renamed v and placed in C(A)C(A); the geometric meaning of the optimization is explained.

Covered · The residual b - v is expanded componentwise into a column vector.

Covered · The squared norm is rewritten as a sum of squared component differences.

Covered · Speaker names x* as the least squares estimate / solution / approximation and ties the terminology to the sum of squares.

Covered · Introduces inconsistency of Ax=b and the least-squares minimization objective.

Covered · States the geometric principle that the closest vector in a subspace is the orthogonal projection.

Covered · Derives and boxes A x^* = proj_{C(A)C(A)} b as the characterization of the least-squares solution.

Covered · Explains that the direct projection-matrix formula is hard to use and motivates an easier route.

Covered · Subtracts b from both sides to obtain A x^* - b = proj_{C(A)C(A)} b - b.

Covered · Continuous whiteboard lecture segment: the clip begins with the least-squares setup and projection diagram, develops the residual orthogonality and left-null-space identity, and ends with the normal equations ATAx∗−ATb=0A^T A x* - A^T b = 0.

Covered · Opening board already displays the least-squares minimization setup, projection identity, orthogonality relation, and the normal equation result while the speaker finishes an algebraic simplification.

Covered · The speaker restates the original inconsistent system Ax=b and defines x^* as the minimizer of ||b - A x^*||, explaining the name least squares.

Covered · The geometric characterization is given: the closest vector in C(A)C(A) to b is proj_{C(A)C(A)} b, so A x^* equals that projection.

Covered · The speaker introduces the simpler algebraic method, multiplies Ax=b on the left by ATA^T, writes ATAx=ATbA^T A x = A^T b, and identifies its solution as the least-squares solution.

Covered · Close-up board discussion of minimizing ||b-Ax^*||, projection onto C(A)C(A), residual orthogonality, and the normal equations.

Covered · Zoom-out reveals the full board, including column-combination form, geometric projection sketch, subspace identity, and the speaker's remark that the concept is currently abstract but will be useful later.

Explore the knowledge in this video

Open video knowledge graph →

  • Least squares ExplanationAt 3:25
    Why this connection?

    Reviewed current material from 205 seconds defines least squares as minimizing the residual norm for an inconsistent system, identifies the closest attainable vector as an orthogonal projection onto the column space, and derives the normal equations ATAx=ATbA^T A x=A^T b.

Questions this video answers

Understand why

↗
Understand why

↗
Understand why

↗
Understand why

↗
Find a method

↗
Find a method

↗
Find a method

↗