Dimension setup for
For an matrix , the equation forces and . This is the starting framework used in the clip before discussing solvability.
Khan Academy · YouTube · 15:32
This 180-second whiteboard segment sets up the motivation for least squares by studying when has no solution. It first fixes dimensions for an matrix, then rewrites the equation as a linear combination of the columns of . From there it identifies inconsistency with the statement and illustrates that geometrically with a plane for the column space and a vector outside it. The final seconds pivot toward approximation by introducing a candidate , but the exact criterion is not completed within the supplied clip. This 180-second whiteboard segment motivates the least-squares approximation for an inconsistent linear system Ax=b. It starts from the geometric fact that b is not in the column space , so no exact solution exists. The lecturer then defines x* by minimizing the distance ||b - Ax*||, observes that Ax* = v must lie in , expands the residual b - v into components, and rewrites the squared length as a sum of squared errors. This algebraic form directly explains the terminology 'least squares estimate', 'least squares solution', or 'least squares approximation'. This 180-second whiteboard segment develops the geometric meaning of a least-squares solution when Ax=b is inconsistent. It begins by replacing exact solvability with the objective of minimizing ||b - A x^*||, then uses the fact that the closest vector in a subspace is the orthogonal projection to characterize the solution by A x^* = proj_{} b. The speaker notes that computing the projection matrix directly is cumbersome, so the clip ends by subtracting b from both sides to set up an easier derivation. This 180-second whiteboard segment explains least squares geometrically and algebraically. Starting from the unsolvable system Ax=b, it defines x* so that Ax* is the closest point to b in the column space , identifies Ax* with proj_{} b, and uses the perpendicular residual to show lies in . It then invokes , multiplies by , and derives the normal equations * = . This 180-second whiteboard segment reviews the least-squares problem for an inconsistent linear system Ax=b. It first restates the definition of a least-squares solution x^* as the vector minimizing ||b - A x^*||, then connects that minimization to the geometry of projecting b onto the column space . After noting that directly computing the projection is cumbersome, the clip introduces the standard algebraic shortcut: multiply the original equation on the left by to obtain the normal equations . The speaker emphasizes that this new equation always has a solution and that its solution is the desired least-squares solution, not an exact solution of the original inconsistent system. This 32-second clip is a whiteboard-style summary of the least-squares method for an inconsistent linear system Ax=b. The speaker explains that if one can solve the displayed matrix-vector equation, then one has made the best possible attempt at solving Ax=b by minimizing the error between b and A x^*. The board ties together three central viewpoints: minimizing ||b - A x^*||, identifying A x^* with proj_{} b, and deriving the normal equations ^* = from the orthogonality condition (b - A x^*) = 0. In the last seconds, the camera zooms out to show the larger conceptual map, including the column-combination form of Ax=b, the geometric picture of projecting b onto the plane , and the subspace identity . The speaker closes by saying the material is still somewhat abstract but will become clearly useful in the next video.
Use the learning inspector for key ideas and moments, or open the reading tabs for the complete notes.
Generated from the video's visuals and explanation; not verbatim speech.
The lesson opens by writing a general matrix equation on the board: , with specified as an matrix. The speaker immediately uses dimension compatibility to place the unknowns correctly: because has columns, must lie in , and because the product has rows, must lie in . This establishes the setting for the whole discussion.
Next, the speaker adds a hypothesis that matters: suppose this system has no solution. That statement is written explicitly as "No solution to ". The purpose is not yet to solve the system, but to understand what failure of solvability means structurally.
To expose that structure, the matrix is rewritten in terms of its columns, , and multiplied by the component vector . The board then shows the equivalent linear-combination form . This step translates the abstract matrix equation into a concrete statement about weights on the columns of .
With that reformulation in hand, the speaker interprets the assumption of no solution: there is no choice of weights on the columns of that produces . Equivalently, no linear combination of the columns of equals . The clip then names this geometrically by writing that is not in the column space .
The algebraic claim is then visualized. A purple plane-like region is drawn and labeled , standing for the collection of all linear combinations of the columns of . An origin is marked, and a cyan arrow labeled is drawn starting from the origin but pointing outside that plane. The picture makes the previous sentence intuitive: if lies outside the column space, then it cannot be reached by any combination of the columns.
After establishing that exact solvability fails, the speaker contrasts this with an older habit of stopping there: form an augmented matrix, row-reduce, encounter a contradiction such as , and conclude there is nothing more to do. The lesson instead asks whether one can do better by finding something close to a solution.
That question leads directly to a new symbol, , written at the lower left together with the partial phrase "where ". Within this supplied segment, the intended role of is clear at the level of motivation: it is being introduced as a candidate associated with an approximate solution. However, the exact condition it is supposed to satisfy is cut off before the clip ends.
The board opens on an inconsistent linear system. We have with and , and the lecturer has already written that there is no solution to . On the left, the same equation is expanded in two equivalent ways: as a matrix times a column vector, [ ... ][ ... ]^, and as a linear combination of columns, + ... + . On the right, a purple plane labeled represents the column space, while a cyan arrow labeled b points outside that plane. The key point is that exact solvability would require b to be expressible as a combination of the columns of A, but the picture explicitly says b is not in .
From that failed exact solution, the lecture pivots to approximation. The speaker says that when he talks about getting 'close,' he means close in length. He therefore introduces x* by the requirement that A x* be as close as possible to b, and writes the objective as minimize ||b - A x*||. Mathematically, this changes the question from 'solve exactly' to 'choose x* so that the residual vector b - A x* has smallest possible Euclidean norm.'
To make the geometry clearer, the lecturer renames the candidate output vector. He says A x is always a member of the column space, because multiplying any vector in by A forms a linear combination of the columns of A. He writes * and places v inside the plane . Thus the optimization is no longer abstract: among all vectors v that lie in , we want the one nearest to the external point b. The invariant constraint is ; the quantity being reduced is the distance from b to v.
Next the residual is written in coordinates. Since *, the vector b - A x* becomes b - v. The lecturer expands this as the column vector [, , ..., ]^T. This step shows that the single geometric distance ||b - v|| is built from all coordinate mismatches between the target vector b and the chosen column-space vector v.
The lecture then squares the residual norm. He states that the length squared of [, ..., ]^T is ()^ + ... + ()^2. This is the crucial algebraic reformulation: minimizing the Euclidean distance is represented by minimizing an explicit sum of squared errors. The reason for squaring here is not to change the geometric target but to expose the additive quadratic structure that will name the method.
Finally, the terminology is attached to the construction. Because the quantity being minimized is a sum of squares of the residuals, the lecturer calls x* the least squares estimate for , also referring to it as the least squares solution or least squares approximation. The segment closes with the conceptual chain fully visible: b outside prevents an exact solution; x* is chosen so that A x* = v in is nearest to b; and the nearest-distance condition is expressed algebraically as minimizing the sum of squared componentwise errors.
The board starts from an inconsistent system: there is no solution to Ax=b. Instead of stopping there, the lesson introduces a least-squares goal: find x^* so that A x^* is as close as possible to b.
That closeness is expressed by minimizing ||b - A x^*||. Geometrically, A x^* must lie in the column space , because multiplying A by any vector produces a linear combination of the columns of A.
The speaker then recalls a general geometric fact: among all vectors in a subspace, the one closest to an external vector is the orthogonal projection of that external vector onto the subspace.
Applying that fact with subspace and external vector b gives the key characterization of the least-squares solution: A x^* = proj_{} b. This equation is written and boxed on the board.
However, the speaker points out that using the explicit projection formula is still hard computationally, so the lesson looks for an easier path.
To begin that easier derivation, b is subtracted from both sides of the boxed equation, producing A x^* - b = proj_{} b - b, which reframes the problem in terms of residuals and prepares the next argument.
The board opens with the approximate-system context: a matrix equation [ ... ][ ... ]^ is shown, together with the note that the goal is least squares. The speaker frames x* as the choice for which Ax* is as close as possible to b, and the objective is written as minimizing ||b-Ax*||, expanded into a sum of squared component differences.
At the center, a geometric picture supports the algebra. A purple plane labeled represents the column space. The vector b is drawn outside that plane, while the yellow vector Ax* lies inside it. The speaker identifies Ax* with the orthogonal projection of b onto , making explicit that the least-squares fit is the nearest point in the column space.
The explanation then focuses on the error vector between b and its projection. Using the right side of the board, the speaker writes and notes that this is the same as proj_{} b - b once the projection identity is used. The key claim is that this difference is perpendicular to the plane , so it belongs to .
Next, the video translates the geometric orthogonality condition into a standard matrix identity. The board states , and the speaker calls the left null space of A. From this, the residual condition is rewritten as .
Once the residual is known to lie in the left null space, the speaker applies the defining property of that null space: multiplying it by gives the zero vector. This produces .
Finally, the expression is expanded by distributivity to , which is the normal-equation form derived in this clip. The segment closes on this algebraic condition, linking the original projection geometry to the practical linear system used to compute least-squares solutions.
The clip opens on a dense blackboard summary of the least-squares topic. Visible relations include the minimization statement for ||b - A x^*||, the projection identity A x^* = proj_{} b, an orthogonality relation involving (Ax^*-b), and the normal equation ^* = . The speaker is finishing an algebraic simplification that leaves the normal-equation form on the board.
The explanation then steps back to the origin of the problem. The system is inconsistent, so there is no exact solution. Instead of solving that impossible equation, the goal is to choose a vector x^* that makes A x^* as close as possible to b. This is expressed by minimizing the residual norm ||b - A x^*||. The term 'least squares' is justified because minimizing the length of the residual is equivalent to minimizing the sum of squared component differences.
Next, the video gives the geometric interpretation. Among all vectors in the column space , the one closest to b is the orthogonal projection of b onto . Therefore the optimal image A x^* must satisfy A x^* = proj_{} b. If some x^* in produces that projected vector, then that x^* is the least-squares solution.
The speaker then notes that computing the projection directly is conceptually clear but algebraically cumbersome. A simpler route is to transform the original inconsistent equation. Starting from , multiply both sides on the left by . This gives , the normal equations.
Finally, the clip stresses the logical status of this new equation. Its solution is not an exact solution of the original system . Rather, the normal equations are guaranteed to have a solution, and that solution is precisely the least-squares solution x^* sought at the start.
The clip opens on a dense handwritten board centered on the least-squares problem for a linear system Ax=b. The speaker begins by pointing out the structural difference between the objects in the equation being discussed: one side is a matrix and the other is a vector. This sets up the practical claim that if we can solve this displayed equation, then we have done the best possible job of approximating a solution to the original system.
The board makes that “best possible job” precise through the minimization statement minimize ||b - A x^*||. In words, x^* is chosen so that A x^* is as close as possible to b. The speaker then states the consequence directly: we have minimized the error, we obtain A x^*, and the difference between A x^* and b is as small as possible. This is exactly what the board labels as the least-square solution.
Alongside the minimization formulation, the board presents an equivalent geometric identity: A x^* = proj_{} b. That is, the image of the least-squares solution is not arbitrary; it is the orthogonal projection of b onto the column space of A. The board also records the residual condition in several compatible forms, including A x^* - and (b - A x^*) = 0. From this orthogonality condition, the displayed algebra rewrites the problem into the boxed normal equations ^* = .
In the final seconds, the camera zooms out to reveal the wider conceptual layout. On the left, Ax=b is rewritten as a column-combination problem, , clarifying that solving the system means asking whether b lies in the span of the columns of A. In the middle, a geometric sketch shows b, its projection onto the plane representing , and the perpendicular residual direction. On the right, the board states , linking the orthogonal-complement description of the residual to the null space of A transpose. Over this full-board view, the speaker remarks that the topic still feels abstract in this clip but promises that the next video will show why the idea is very useful.
For an matrix , the equation forces and . This is the starting framework used in the clip before discussing solvability.
Writing by columns turns into . Thus solving the system means asking whether can be formed from the columns of .
The clip states that if has no solution, then no linear combination of the columns of equals . In column-space language, this is exactly the assertion that is not in .
A plane-like region labeled represents all attainable linear combinations of the columns of . Drawing outside that region visualizes why the equation cannot be solved exactly.
Instead of ending at “no solution,” the lesson pivots toward finding a vector that gets close to the target. The symbol is introduced, but the precise defining condition is not completed within this segment.
The segment begins with the equation , where A is , , and . The board states that there is no solution because b is not in the column space . The same equation is rewritten as + ... + , showing that an exact solution would require b to be a linear combination of the columns of A.
Since exact equality is impossible, the lecturer defines x* by making A x* as close as possible to b. 'Close' is made precise using vector length, so the problem becomes minimizing the norm of the residual b - A x*.
The speaker introduces * and notes that any product A x is a linear combination of the columns of A, hence belongs to . Geometrically, the least-squares problem is to choose the vector v in the plane that is nearest to the external vector b.
Once * is substituted, the residual b - A x* is rewritten as b - v and expanded componentwise. This displays the error as the vector of coordinate differences between b and the chosen column-space vector v.
The lecturer then squares the residual norm. The squared Euclidean length of the residual vector equals the sum of the squared componentwise errors. This algebraic form is what motivates the name of the method.
Because the objective is a sum of squared residuals, x* is called the least squares estimate for , also the least squares solution or least squares approximation. The terminology refers to minimizing the total squared error between b and the best attainable vector A x* in .
When Ax=b has no exact solution, the video defines a least-squares approach by seeking x^* such that A x^* is as close as possible to b. This is formalized as minimizing the residual norm ||b - A x^*||.
For a vector outside a subspace, the nearest point inside the subspace is its orthogonal projection. In this lesson the relevant subspace is the column space .
Combining the minimization objective with the projection principle yields the central equation of the clip: the least-squares solution satisfies A x^* = proj_{} b.
The speaker notes that computing the projection via is difficult in practice, motivating a different derivation route.
To search for an easier method, the boxed equation is rewritten by subtracting b from both sides, giving a residual form that sets up the next step of the lesson.
The clip defines the least-squares solution x* by requiring Ax* to be as close as possible to b. On the board this is expressed through the minimization of ||b-Ax*|| and visually through the column-space plane . The important point is that x* is not presented as an exact solution to Ax=b, but as the coefficient vector producing the best approximation inside the range of A.
The geometric core of the lesson is the boxed identity A x* = proj_{} b. This says that the least-squares image of x* is exactly the orthogonal projection of b onto the column space. The diagram reinforces this by placing Ax* in the plane and b outside it.
Because Ax* is the projection of b onto , the difference vector from b to Ax* is perpendicular to the entire column space. The board records this as . This is the bridge from the geometric picture to the later algebraic derivation.
The speaker then uses the fundamental-subspaces identity , naming the left null space of A. This converts the orthogonality statement into a null-space membership statement for the residual.
From , multiplying by gives zero. Expanding the product yields the normal equations in the form shown at the end of the clip. This is the main algebraic payoff of the geometric argument.
For an inconsistent system , the least-squares solution x^* is defined as the vector that minimizes the residual size ||b - A x^*||. The name comes from equivalently minimizing the sum of squared component errors.
The best approximant to b inside the column space is the orthogonal projection of b onto . Hence the least-squares condition can be written as A x^* = proj_{} b.
A practical way to compute the least-squares solution is to left-multiply the original equation by , producing . Solving this system gives the least-squares solution x^*.
Explore conditions, steps and evidence. Supplementary explanations are labeled separately from content shown in the video.
The speaker says, "Let's say I have some matrix A."
Yellow handwritten is written at the upper left.
A
An matrix whose columns are used to form linear combinations.
The speaker says, "I have the equation Ax is equal to b" and later "x would have to be a member of Rk."
Yellow handwritten appears in , with written nearby.
Unknown vector in the matrix equation .
The speaker says, "Ax is equal to b" and "b is a member of Rn."
Yellow handwritten appears in , with written nearby.
Right-hand-side vector in the equation .
The speaker says, "it's an n by k matrix" and explains that has entries because there are columns.
Yellow handwritten is placed under .
n, k
Dimensions of : rows and columns.
Positive integers
The speaker says, "If I write A like this, A1, A2 ... all the way through Ak" and refers to them as column vectors.
Pink handwritten bracketed expression is written below the initial equation.
,,,
Column vectors of , indexed from to .
Each is a column of ; hence
The speaker says, "multiply it times x1, x2 all the way through xk."
Pink handwritten column vector is written next to the column expansion of .
,,,
Scalar components of the unknown vector , used as weights on the columns of .
for
The speaker says, "b is not in the column space of A."
Yellow handwritten text states " is not in the ".
Column space of , i.e. the set of all linear combinations of the columns of .
Subspace of
The speaker says, "I want to find some x star for now where A times x star is..."
Yellow handwritten appears at the lower left, followed by "where ".
The sentence and displayed condition after are cut off before completion.
A candidate vector introduced for a better approximate solution when has no exact solution.
Intended to lie in by analogy with , but the clip does not explicitly restate the domain
Matrix A appears in and is described as .
The speaker refers to multiplying a vector by the matrix A and getting a member of its column space.
A
An matrix whose columns span .
matrix
is written beside .
The speaker says any vector in times the matrix A gives a member of the column space.
x
Unknown vector in .
is written beside .
The speaker repeatedly refers to making Ax* as close as possible to b.
b
Target vector in that generally lies outside in this setup.
x* is introduced in the line 'x* where Ax* is as close to b as possible'.
The speaker defines x* through minimizing ||b - Ax*||.
x^*
Vector chosen so that Ax* is the closest point in to b.
The speaker introduces a matrix that is by and the equation , then states the dimensions of and .
Yellow handwritten , , , , and appear across the top of the board.
The clip begins with a general linear system written as , where is an matrix. Dimension compatibility forces to have entries and to have entries, so and .
has rows and columns
The product is defined only when has entries
The resulting vector has entries, matching
The speaker expands into its column vectors and multiplies by the components of , saying this is the same equation rewritten.
Pink handwritten is written, followed by .
The equation is rewritten by expressing in terms of its columns. Multiplying the column-expanded matrix by the component vector gives the weighted sum of the columns of , showing that solving is equivalent to finding weights that express as a linear combination of the columns of .
denotes the th column of
denotes the th component of
The equivalence uses the standard column interpretation of matrix-vector multiplication
The speaker says there is no solution, then explains that this means no set of weights on the column vectors of can get to , equivalently no linear combination equals , and finally says is not in the column space of .
Cyan handwritten text reads "No solution to "; yellow handwritten text reads " is not in the ".
When has no solution, the vector cannot be obtained as any linear combination of the columns of . The clip names this geometrically by saying is not in the column space .
is the set of all linear combinations of the columns of
The equivalence is stated for the specific system under discussion
The board states 'No solution to Ax = b' and shows the equivalent column-combination equation.
A cyan arrow labeled b points outside the purple plane labeled , with text 'b is not in the '.
This segment begins from the case where the exact linear system Ax=b has no solution because b does not belong to the column space of A. The video expresses Ax both as a matrix product and as a linear combination of the columns of A, making clear that solvability requires b to lie in .
A is
b ∉ in the motivating picture
The speaker writes 'minimize ||b - Ax*||'.
He says, 'when I say close, I'm talking about length' and 'I want to minimize the length of b minus Ax*'.
The vector Ax* is identified with v inside , while b remains outside the plane.
The least-squares idea is introduced geometrically: choose x* so that Ax*, viewed as a point v in the column space, is as close as possible to b. Closeness is measured by vector length, so the problem becomes minimizing the norm of the residual b - Ax*.
Ax* must lie in
Distance is measured by Euclidean length
The speaker says, 'Ax is going to be a member of my column space' and repeats that any Ax is in the column space.
He writes v = Ax* on the diagram.
The vector v is placed inside the purple plane labeled .
The video uses the identity v = Ax* to translate the optimization into a geometric search inside . Because Ax is always a linear combination of the columns of A, the candidate output vector must remain in the column space, and the task is to pick the best such vector relative to b.
A is
The board expands the norm into the vector [, , ..., ]^T.
The speaker says to take the difference between each of the elements.
Once v = Ax* is introduced, the residual b - Ax* is rewritten as b - v and then displayed as the vector of componentwise differences. This makes explicit that the distance being minimized depends on all coordinate mismatches between b and the chosen column-space vector v.
b,
The board shows ||[, ..., ]^T||^ + ... + ()^2.
The speaker says, 'the length squared of this is just going to be ...' and lists the squared terms.
The lecture converts the geometric minimization of length into an algebraic minimization of a sum of squares. By squaring the residual norm, the objective becomes the explicit quadratic expression in the componentwise errors, which motivates the terminology 'least squares'.
Using the standard Euclidean norm on
The speaker says, 'I want to get the least squares estimate here' and explains that this is why it is called the least squares estimate.
The board adds the label 'least squares' pointing to x*.
After deriving the sum-of-squared-residuals form, the video names the object x* as the least squares estimate, also calling it the least squares solution or least squares approximation for Ax=b. The terminology is tied directly to minimizing the sum of squared component errors.
The minimization is over Ax* ∈ approximating b
The board states 'No solution to Ax=b' and introduces x^* with A x^* as close to b as possible.
The speaker says there is no solution to this, but maybe we can find some x^* where A times x^* is as close to b as possible.
When the exact linear system Ax=b is inconsistent, the video defines a least-squares approach: choose x^* so that A x^* lies in the column space of A and is as close as possible to b.
Ax=b has no exact solution
A x^* must lie in
The board writes 'minimize ||b - A x^*||'.
The speaker says they want to get this vector to be as close to b as possible.
The least-squares problem is posed as minimizing the norm of the difference between b and A x^*.
b is fixed
x^* is the variable being chosen
The speaker says the closest vector in any subspace to a vector not in that subspace is the projection.
The board later writes A x^* = proj_{} b.
For a vector outside a subspace, the nearest point inside that subspace is its orthogonal projection onto the subspace. In this lesson the subspace is .
is treated as a subspace
b is not in
The speaker states three equivalent ways of describing the same situation: no weights on the columns reach , no linear combination equals , and is not in the column space of .
The board shows the expanded linear-combination equation and the statement " is not in the ".
For the system , having no solution is equivalent to saying that is not a linear combination of the columns of , equivalently .
is an matrix
and
denotes the column space of
For the given matrix and vector , the absence of any satisfying is equivalent to .
The board explicitly states 'b is not in the '.
The cyan vector b is drawn outside the purple plane .
The spoken explanation continues from the premise that there is no solution to Ax=b.
In the setup shown, because b is not in the column space , the equation Ax=b has no solution.
A is
b ∉
For the displayed A, x, and b in this example setup.
The speaker says, 'any Ax is going to be in your column space' and repeats the statement.
He writes v = Ax* and places v inside .
For any , the vector Ax is a member of the column space .
A is
For all x in .
The objective is written as minimize ||b - Ax*||.
The speaker says he wants Ax* to be as close as possible to b and identifies Ax* with v in .
Choosing x* to minimize ||b - Ax*|| is equivalent to choosing the vector v = Ax* in that is closest to b.
A is
Distance measured by Euclidean norm
For the given A and b in the inconsistent-system setup.
The board displays the equality between the squared norm and the sum of squared component differences.
The speaker verbally derives the same expression as the length squared.
For b,, ||b-v||^ + ... + ()^2.
b,
Standard Euclidean norm
For all component indices ,...,n.
The speaker states that the closest vector to b that is in the subspace is the projection of b onto the column space.
The diagram shows b outside the plane and identifies the nearest point in the plane as the projection.
If b is not in a subspace, then the closest vector to b inside that subspace is the orthogonal projection of b onto the subspace.
There is a subspace, here
b is a vector not in that subspace
For the given vector b and the given subspace .
The board writes and boxes A x^* = proj_{} b.
The speaker says Ax needs to be equal to the projection of b on my column space.
The least-squares solution x^* is characterized by A x^* = proj_{} b.
Ax=b has no exact solution
x^* is chosen to minimize ||b - A x^*||
For the least-squares solution x^* associated with the inconsistent system Ax=b.
The speaker says Ax* is as close to b as possible and ties this to projection onto the subspace.
Left panel: where Ax* is as close to b as possible.
For the least-squares problem, Ax* is the vector in that is closest to b.
x* is a least-squares solution to Ax=b.
Distance is measured by the Euclidean norm.
for the given matrix A and vector b
The speaker says if I take the projection of b ... minus b, I'm going to get this vector ... this vector right here is orthogonal.
and .
If , then is orthogonal to .
Projection is orthogonal projection onto .
for all vectors in
.
The speaker states this identity directly.
.
A is a real matrix.
for the matrix A under discussion
and then .
The speaker says this times A transpose has got to be equal to zero and then simplifies.
If , then .
x* is a least-squares solution.
belongs to the left null space of A.
for the least-squares solution x*
The speaker says: 'This right here will always have a solution, and this right here is our least squares solution.'
The referenced equation on screen is ^* = .
The video states solvability but does not prove it in this clip.
The equation always has a solution, and that solution is the least-squares solution of the original inconsistent system.
The original system is .
The equation under discussion is the normal equation .
For the matrix A and vector b under discussion in this lesson segment.
The speaker says, "Let's just expand out A," writes the columns, multiplies by the components of , and states that the result is the same equation.
The board displays and then .
Start from the original matrix equation.
Given setup in the video.
Write explicitly as its column vectors.
Definition of a matrix by its columns.
Write the unknown vector in components.
Standard coordinate representation of a vector in .
Substitute the column form of and component form of into the equation.
Substitution into the original equation.
Evaluate the matrix-vector product as the corresponding weighted sum of columns.
Column interpretation of matrix-vector multiplication.
The system is exactly the statement that some linear combination of the columns of equals .
The board progresses from minimize ||b - Ax*|| to v = Ax*, then to the residual vector and finally to the sum of squared terms.
The speaker narrates each step: closeness means length, Ax* is in , take componentwise differences, then square and sum them.
Start from the goal of making Ax* as close as possible to b.
Stated directly by the speaker as the definition of 'close' via length.
Rename the candidate output vector as v and note that it lies in the column space.
The speaker says Ax is a member of the column space and writes v = Ax*.
Rewrite the residual as the vector of componentwise differences.
The speaker explicitly takes the difference between each element of b and v.
Square the residual norm to obtain a sum of squared errors.
The speaker states that the length squared is the sum of the squared component differences.
Name the minimizing choice x* as the least squares estimate.
The speaker says this is why the quantity is called the least squares estimate, solution, or approximation.
The least-squares problem is motivated as minimizing the sum of squared componentwise residuals between b and the best available vector Ax* in .
The board shows [ ... ][ ... ]^ and + ... + .
The speaker explains Ax as multiplying a vector in by the columns of A.
Write Ax using the columns of A and the entries of x.
Displayed on the board as the expanded matrix equation.
Interpret the product as a linear combination of the columns of A.
The speaker explains that multiplying by A combines its columns with the coefficients .
Solving Ax=b exactly would require expressing b as a linear combination of the columns of A, i.e. requiring .
The board moves from 'minimize ||b - A x^*||' to 'A x^* = proj_{} b'.
The speaker explains that the closest vector in the subspace is the projection, so Ax must equal that projection.
Start from an inconsistent linear system.
Stated directly on the board and in the audio.
Replace exact solvability by the goal of making A x^* as close as possible to b.
Explicitly written as the least-squares objective.
Any vector of the form A x^* lies in the column space of A.
Audio explanation that A times a vector is a linear combination of the column vectors.
Use the geometric fact that the nearest point in a subspace is the orthogonal projection.
Spoken principle referenced from earlier videos and illustrated by the diagram.
Therefore the least-squares solution is characterized by equality with the projection.
Combines the previous steps.
The least-squares problem is equivalent to requiring A x^* to be the projection of b onto .
The board writes A x^* - b = proj_{} b - b.
The speaker says they will subtract b from both sides to see if something interesting appears.
Begin with the boxed characterization of the least-squares solution.
Previously derived and written on the board.
Subtract b from both sides.
Valid algebraic operation applied to both sides of an equality.
The equation is rewritten in residual form, setting up the next step of the lesson.
The speaker moves from orthogonality to membership in , then to , then multiplies by and expands.
Visible chain: ; ; ; ; ; .
The clip does not prove ; it invokes it as known from earlier videos.
Start from the geometric characterization of the least-squares fit as the orthogonal projection of b onto the column space.
Stated explicitly on the board and in the audio.
The difference between the projected point and b is orthogonal to the subspace .
Definition/property of orthogonal projection, as explained by the speaker.
Replace the orthogonal complement of the column space by the left null space.
Standard fundamental-subspaces identity invoked in the video.
Therefore the residual vector lies in the null space of .
Substitution using the previous identity.
Multiplying a vector in by yields the zero vector.
Definition of null space.
Distribute over the difference to obtain the normal equations in homogeneous form.
Linearity/distributivity of matrix multiplication.
Move to the other side to isolate the standard normal-equation form.
Algebraic rearrangement.
The least-squares solution satisfies the normal equations * = .
The speaker explains multiplying both sides of Ax=b by .
The board shows and then labels the solution as x^*.
Begin with the original linear system that the video says has no exact solution.
Stated directly in the clip as the starting problem.
Left-multiply both sides by .
Explicitly performed and narrated in the clip.
Use associativity of matrix multiplication to rewrite the left side.
Standard matrix algebra; the simplified form is what is written on the board.
Rename the solution of the new equation as the least-squares solution x^*.
The speaker explicitly identifies this solution as the least-squares solution.
The least-squares solution can be found by solving the normal equations ^* = instead of directly computing the projection.
The board displays the chain (b - A x^*) = 0, then ^* - , then the boxed ^* = .
The speaker identifies one side as a matrix and the other as a vector while discussing the final equation.
Start from the condition that the residual is annihilated by .
This is the orthogonality/residual condition written directly on the board.
Distribute over the difference and move terms to one side.
Standard matrix algebra applied to the displayed equation.
Add to both sides to isolate the unknown x^* term.
Algebraic rearrangement of the previous line, matching the boxed equation on the board.
The least-squares condition is equivalently encoded by the normal equations ^* = .
The zoomed-out board shows A x^* = proj_{} b, A x^* - , , and A x^* - .
A geometric sketch shows b, its projection onto a plane representing , and the perpendicular residual direction.
The closest point to b in is its orthogonal projection onto that subspace.
Displayed as the defining identity for the least-squares image vector.
The residual from b to its projection is perpendicular to the subspace .
This is the geometric meaning of orthogonal projection, also written on the board.
The orthogonal complement of the column space is the null space of A transpose.
Explicit identity shown on the board after the zoom-out.
Therefore the residual also lies in .
Substitution using the previous two lines.
The projection picture, the orthogonal-complement statement, and the null-space statement are consistent descriptions of the same least-squares geometry.
The speaker says, "let's see if we can visualize it a bit," draws the column space of , assumes it is a plane in , marks the origin, and draws outside the plane.
A purple parallelogram-like plane is drawn and labeled ; a cyan arrow from the origin points upward outside the plane and is labeled .
The drawing is schematic; the ambient dimension is not fixed beyond being at least large enough to contain a plane.
Purple plane representing
Origin point
Cyan vector arrow labeled
Labels and
First the plane for is drawn.
Then the origin is marked.
Finally the vector is drawn starting at the origin and extending outside the plane.
The plane is meant to represent all linear combinations of the columns of .
The vector remains outside that plane throughout the sketch.
The visual encodes the algebraic claim that if is not in the column space, then no linear combination of the columns of can equal .
The speaker asks what if we can find a solution that gets us close to the target and says, "I want to find some x star for now where A times x star is..."
Yellow handwritten and the partial phrase "where " appear at the lower left.
The defining condition for is not completed within the supplied clip.
Symbol
Partial expression
A new symbol is added after the discussion of the inconsistent system.
The speaker begins to define a property of but the clip ends before completion.
The earlier equation remains on screen as the inconsistent target problem.
The column-space sketch remains visible while is introduced.
This visual transition signals a move from exact solvability to an approximation problem, but the precise criterion for is not yet shown in this segment.
A purple parallelogram labeled is drawn on the right, with a cyan vector b starting from the same origin and ending outside the plane.
Later a white vector labeled v = Ax* is drawn inside the purple plane.
Purple plane labeled
Cyan vector labeled b
White vector labeled v = Ax*
Origin point common to the drawn vectors
At first only b and the plane are shown.
Around 52-80 seconds, the speaker introduces v = Ax* and draws it inside .
The residual discussion then relates b and v by componentwise subtraction.
The plane stays fixed as the set of attainable Ax vectors.
b remains outside throughout the motivating picture.
The diagram encodes the core idea that exact equality Ax=b is impossible here, but one can still choose a vector v in that is nearest to b.
The phrase 'minimize ||b - Ax*||' is written progressively on the board.
The speaker says he wants to minimize the length of b minus Ax*.
Text 'minimize'
Expression ||b - Ax*||
The word 'minimize' is written first.
Then the norm bars and the residual expression b - Ax* are added.
The objective remains a distance-minimization statement.
This visual step converts the verbal idea of 'as close as possible' into a precise optimization problem.
The board expands the norm into a column vector of differences and then into a sum of squared terms.
Column vector [, ..., ]^T
Sum ()^2 + ... + ()^2
First the residual vector is written componentwise.
Then its squared norm is rewritten as a sum of squares.
The underlying quantity being minimized is still the distance between b and v = Ax*.
This visual expansion is what motivates the name 'least squares'.
Left side contains the algebraic setup for Ax=b and the minimization expression; right side contains a geometric sketch of b outside the plane .
Matrix equation Ax=b
Minimization expression ||b - A x^*||
Purple plane labeled
Cyan vector labeled b
Yellow vector labeled A x^*
The right-side diagram gains the label proj_{} b.
A new boxed equation A x^* = proj_{} b is written below the diagram.
The boxed equation is later rewritten as A x^* - b = proj_{} b - b.
The left-side setup remains visible throughout the clip.
The plane remains the subspace of interest.
b remains outside in the geometric picture.
The visual organization ties the algebraic least-squares objective to the geometric idea of projecting b onto the column space.
A cyan vector b points outside a purple plane labeled ; a yellow vector A x^* lies in the plane; the projection label is added near the plane.
The speaker says the closest vector in the subspace is the projection of b onto the column space.
The exact arrowhead placement of the projection annotation is approximate from the sampled frames.
Vector b
Plane
Vector A x^*
Projection annotation proj_{} b
The diagram shifts from showing only b and A x^* to explicitly identifying the nearest point in as proj_{} b.
b stays outside the plane.
A x^* stays inside the plane.
The figure illustrates that minimizing ||b - A x^*|| means choosing the point in closest to b, namely the orthogonal projection.
The screen is split into a left algebraic setup, a central geometric sketch, and a right derivation column.
New formulas are added sequentially on the right as the explanation proceeds.
left panel with Ax=b and least-squares objective
central purple plane with vectors b, proj_{}b, and Ax*
right panel accumulating implication chain to normal equations
Right-side formulas are written step by step from projection identity to .
The left least-squares setup and central projection diagram remain visible while the derivation is built.
The visual layout links the optimization statement, the geometric projection picture, and the algebraic normal-equation derivation.
A purple plane labeled contains the yellow vector ; b is drawn outside the plane; an orange/yellow arrow connects b down to the plane.
The speaker points to the vector going straight down onto the plane and calls it orthogonal to the column space.
Exact color naming varies slightly between orange and yellow in the residual arrow across frames.
plane
vector b outside the plane
vector inside the plane
residual arrow from b to
The residual is emphasized as the perpendicular connector between b and its projection.
stays in while b stays outside it in the sketch.
The diagram encodes that the best approximation to b from is obtained by dropping perpendicularly to the plane.
The blackboard contains multiple color-coded formulas: the minimization statement, the projection identity, the orthogonality relation, and the final normal equation.
minimize ||b - A x^*||
expanded squared-sum expression
A x^* = proj_{} b
(Ax^*-b)=0
^* =
Ax=b with 'no solution' annotation
The lower-left area is rewritten during the clip to restate the problem setup.
The final normal equation is emphasized as the practical solution method.
The overall topic remains least-squares approximation for an inconsistent linear system.
The projection-based geometric interpretation stays visible while the algebraic method is introduced.
The visual layout connects three levels of the same idea: optimization definition, geometric projection characterization, and algebraic normal-equation computation.
Close-up of a black digital board with multicolored handwritten formulas about least squares, including minimize ||b - A x^*||, A x^* = proj_{} b, and ^* = .
A cursor moves among the formulas while the speaker talks.
handwritten formulas
cursor
boxed normal equations
projection identity
norm expression
Cursor points to different parts of the board as the speaker discusses matrix-versus-vector structure and minimization.
The board content remains the same during the close-up.
The central identities A x^* = proj_{} b and ^* = stay visible.
The close-up visually organizes the least-squares argument around three linked ideas: minimize residual norm, project onto , and solve the normal equations.
The camera zooms out to reveal a much larger board.
Newly visible regions include the column-combination equations at upper left and a geometric sketch with b, proj_{} b, and a plane labeled .
full whiteboard
column-combination equations
geometric sketch of b and its projection
plane labeled
subspace identities on the right
Additional formulas appear: [ ... ][ ... ]^ and + ... + .
A geometric diagram becomes visible showing b, its projection onto , and the perpendicular residual.
Right-side subspace relations and A x^* - become visible.
The previously seen least-squares formulas remain part of the same overall argument.
The projection identity A x^* = proj_{} b continues to anchor the geometry.
The zoom-out connects the algebraic normal equations to the geometric picture of projecting b onto the column space and interpreting the residual as orthogonal to that space.
The speaker says that up until now they would make an augmented matrix, put it in reduced row echelon form, get a line saying zero equals one, and conclude there is no solution and nothing more can be done.
Treating a row-reduction contradiction such as as the end of the analysis, with no further useful objective available.
The video explicitly contrasts that old endpoint with a new goal: even if no exact solution exists, one may still seek a solution that gets close to the target.
The speaker contrasts 'No solution to Ax=b' with making Ax* 'as close to b as possible'.
b is shown outside , so exact equality is visually impossible.
One might think that because we are minimizing ||b - Ax*||, the resulting x* solves Ax=b exactly.
In this setup Ax=b has no exact solution because b ∉ ; x* only gives the best approximation within .
The speaker repeatedly says any Ax is in the column space.
He writes v = Ax* inside .
One might forget that the candidate vector Ax* is constrained to lie in .
Because Ax* is a linear combination of the columns of A, the optimization searches only among vectors in the column space.
The board first shows ||b - Ax*|| and then rewrites the squared version as a sum of squares.
The speaker says, 'Let me take the length squared actually.'
One might think the lecture abruptly changed the problem when moving from ||b - Ax*|| to the sum of squared residuals.
The video uses the squared norm because it has the explicit sum-of-squares form that motivates the terminology 'least squares'.
The speaker contrasts 'no solution to this' with finding x^* so that A x^* is as close to b as possible.
The board keeps 'No solution to Ax=b' while introducing the least-squares objective.
If Ax=b has no solution, one might think nothing useful can be said about x.
The video explicitly replaces exact solvability with a least-squares approximation problem: choose x^* so that A x^* is as close as possible to b.
The speaker says it is still pretty hard to find because taking A times the inverse of A transpose A times A transpose is hard.
This is a caution stated by the speaker rather than a fully worked counterexample.
One might assume that once the projection formula is known, the least-squares solution is straightforward to compute directly.
The video states that the explicit projection-matrix route is hard to use here, motivating an easier derivation.
Top-left text reads No solution to Ax=b.
The whole clip reframes the problem as finding Ax* as close to b as possible.
One might think the appearance of x* means there is an exact solution to Ax=b.
The video explicitly starts from the case of no solution and replaces exact equality by closest approximation in the column space.
Both and appear.
The speaker says you could say b plus this vector is equal to my projection of b onto my subspace.
Students may confuse the direction of the perpendicular error vector.
Here the residual is taken as Ax*-b, equivalently proj_{}b-b, and the key property is that this vector is orthogonal to ; reversing the sign would still be orthogonal but changes the written expression.
The speaker says the solution to the multiplied equation will not be the same as the solution to the original equation.
One might think that solving gives an exact solution of the original inconsistent system Ax=b.
The video explicitly distinguishes them: the original system has no solution, while the normal equations have a solution that is the least-squares solution.
The board explicitly writes Ax=b with "no solution" nearby, then separately writes the least-squares minimization and normal equations.
The speaker says this gives our best shot at finding a solution to Ax=b after noting the system may not be solvable exactly.
If Ax=b has no exact solution, one might think nothing useful can be said about x.
The board distinguishes the inconsistent system Ax=b from the least-squares problem, which still produces a meaningful x^* by minimizing ||b - A x^*||.
Speaker: "Notice, this is some matrix, and then this right here is some vector."
One might treat ^* = as just another symbolic equation without noticing the types of the objects involved.
The speaker explicitly points out that one side is a matrix acting on a vector and the other side is a vector, emphasizing the equation's linear-algebraic structure.
After stating and , the speaker expands into columns and rewrites the equation as a linear combination.
The board moves from to and then to .
The dimension setup makes the column expansion meaningful: because has columns, supplies exactly weights for those columns.
The speaker uses the linear-combination form to explain what “no solution” means and then names the column space.
The linear-combination equation remains on screen while the statement " is not in the " is added.
The interpretation of inconsistency depends on first rewriting as a statement about linear combinations of the columns of .
After saying is not in the column space of , the speaker says, "let's see if we can visualize it a bit" and draws the column space and outside it.
The symbolic statement " is not in the " is paired with a purple plane labeled and a cyan vector outside the plane.
The geometric sketch is a direct visualization of the algebraic claim that .
The speaker contrasts the old conclusion of “no solution” with the question of finding something that gets us close, then introduces .
The board retains the inconsistent-system statements and adds with the partial condition "where ".
The exact approximation criterion for is not completed in this clip.
Once exact solvability fails, the lesson pivots toward a broader objective: find an approximate solution rather than stop at inconsistency.
The speaker moves from 'No solution to Ax=b' to defining x* by minimizing ||b - Ax*||.
Both statements appear on the same board.
The least-squares objective is introduced precisely because the exact system Ax=b is inconsistent in the pictured case.
The speaker says Ax is in the column space and then explains that we want this vector to be as close as possible to b.
v = Ax* is written inside .
The geometric meaning of minimizing ||b - Ax*|| depends on recognizing that Ax* ranges over .
The residual vector is written directly under the minimization objective.
The componentwise residual formula is the coordinate expression of the abstract norm-minimization objective.
The speaker says the sum-of-squares form is why the method is called least squares.
The board shows the sum of squared residuals and then labels x* as least squares.
The terminology 'least squares' is justified by the explicit sum-of-squared-residuals expression.
The plane , vector b, and vector v = Ax* are drawn geometrically.
The same idea is rewritten algebraically as a sum of squared component errors.
The geometric minimization of distance and the algebraic minimization of the sum of squared residuals describe the same least-squares problem.
The speaker derives the least-squares condition by invoking the closest-vector/projection principle.
The board connects ||b - A x^*|| with A x^* = proj_{} b.
The minimization objective is solved geometrically by applying the fact that the closest vector in a subspace is the orthogonal projection.
The board writes A x^* = proj_{} b after discussing projection.
The speaker says Ax needs to be equal to the projection of b on the column space.
The projection principle yields the defining equation for the least-squares solution.
The speaker says the direct projection formula is hard, so let's see if we can find an easier way.
The board then subtracts b from both sides of A x^* = proj_{} b.
The clip contrasts the cumbersome direct projection-matrix method with an algebraic rearrangement intended to lead to an easier derivation.
The speaker explains that must be in because there are columns, and is in .
and are written on the board.
The speaker expands into columns and says the product is the same as a weighted sum of those columns.
The board shows and .
The speaker says no set of weights on the columns reaches , equivalently no linear combination equals , and then says is not in the column space of .
The board states "No solution to " and " is not in the ".
Board text says 'No solution to Ax=b' and 'b is not in the '.
The speaker says 'when I say close, I'm talking about length'.
The objective is written as minimize ||b - Ax*||.
The speaker says any Ax is in the column space.
v = Ax* is written inside .
The board expands ||b-v|| into component differences and then squares them.
The speaker says this is why it is called the least squares estimate.
The label 'least squares' is added next to x*.
Board shows 'No solution to Ax=b' and 'minimize ||b - A x^*||'.
Speaker explains finding x^* as close as possible when no exact solution exists.
Speaker says the closest vector in the subspace is the projection.
Diagram shows b outside and the nearest point in the plane.
Boxed equation A x^* = proj_{} b.
Speaker says Ax needs to be equal to the projection of b on the column space.
Speaker says the direct projection formula is hard and proposes an easier way.
Board rewrites to A x^* - b = proj_{} b - b.
Covered · Audio and formulas introduce the matrix equation and the required vector dimensions.
Covered · The speaker states that the system has no solution and writes this condition on the board.
Covered · The matrix is expanded into columns and the equation is rewritten as a linear combination.
Covered · The speaker interprets inconsistency as nonexistence of suitable weights and as .
Covered · A schematic plane for and an outside vector are drawn to visualize the algebraic statement.
Covered · The speaker contrasts the previous row-reduction endpoint with the new question of getting close to a solution.
Covered · A new candidate is introduced for an approximate solution, but the defining condition is cut off before completion.
Covered · Opening board state and spoken setup: Ax=b has no solution because b is outside ; matrix equation is rewritten as a column combination.
Covered · Speaker defines closeness as length and writes the minimization objective ||b - Ax*||.
Covered · Ax* is renamed v and placed in ; the geometric meaning of the optimization is explained.
Covered · The residual b - v is expanded componentwise into a column vector.
Covered · The squared norm is rewritten as a sum of squared component differences.
Covered · Speaker names x* as the least squares estimate / solution / approximation and ties the terminology to the sum of squares.
Covered · Introduces inconsistency of Ax=b and the least-squares minimization objective.
Covered · States the geometric principle that the closest vector in a subspace is the orthogonal projection.
Covered · Derives and boxes A x^* = proj_{} b as the characterization of the least-squares solution.
Covered · Explains that the direct projection-matrix formula is hard to use and motivates an easier route.
Covered · Subtracts b from both sides to obtain A x^* - b = proj_{} b - b.
Covered · Continuous whiteboard lecture segment: the clip begins with the least-squares setup and projection diagram, develops the residual orthogonality and left-null-space identity, and ends with the normal equations .
Covered · Opening board already displays the least-squares minimization setup, projection identity, orthogonality relation, and the normal equation result while the speaker finishes an algebraic simplification.
Covered · The speaker restates the original inconsistent system Ax=b and defines x^* as the minimizer of ||b - A x^*||, explaining the name least squares.
Covered · The geometric characterization is given: the closest vector in to b is proj_{} b, so A x^* equals that projection.
Covered · The speaker introduces the simpler algebraic method, multiplies Ax=b on the left by , writes , and identifies its solution as the least-squares solution.
Covered · Close-up board discussion of minimizing ||b-Ax^*||, projection onto , residual orthogonality, and the normal equations.
Covered · Zoom-out reveals the full board, including column-combination form, geometric projection sketch, subspace identity, and the speaker's remark that the concept is currently abstract but will be useful later.
Reviewed current material from 205 seconds defines least squares as minimizing the residual norm for an inconsistent system, identifies the closest attainable vector as an orthogonal projection onto the column space, and derives the normal equations .
The least-squares fit is identified with the orthogonal projection of onto because the goal of least squares is to find the vector in the subspace that is closest to . A fundamental geometric property of subspaces states that the unique closest point in a subspace to an external vector is its orthogonal projection.
Conditions: is a subspace of ; ; Distance is measured by the Euclidean norm
The solution to the normal equations is considered the least-squares solution because it satisfies the necessary and sufficient condition for minimizing the residual norm . The derivation shows that minimizing this norm is equivalent to requiring the residual to be orthogonal to the column space .
Conditions: The original system may be inconsistent; has a solution
Every product is a member of the column space because matrix-vector multiplication is defined as a linear combination of the columns of . Specifically, if and , then .
Conditions: is an matrix; ; is the span of the columns of
The equation has no solution because solving it is equivalent to finding weights such that . The column space is defined as the set of all possible linear combinations of the columns of .
Conditions: is an matrix; and ; denotes the column space of
Minimizing the Euclidean norm is equivalent to minimizing its square, . The squared norm expands algebraically into the sum of the squared differences of corresponding components: , where .
Conditions: ; Standard Euclidean norm is used;
When has no exact solution, the least-squares solution is defined as the vector that minimizes the Euclidean norm of the residual, . Geometrically, this means choosing such that is the closest possible vector to within the column space .
Conditions: The system is inconsistent (no exact solution exists); Distance is measured by the standard Euclidean norm
The residual vector is orthogonal to the column space because is the orthogonal projection of onto . Orthogonality to means is in the orthogonal complement .
Conditions: is the orthogonal projection of onto ; (Fundamental Theorem of Linear Algebra); Matrix multiplication distributes over subtraction