Artificial Intelligence from First Principles A Pedagogical Mathematical Journey from Linear Regression to Large Language Models Global Notation Registry Throughout this text, we rigorously maintain consistent mathematical notation. When a new mathematical object is constructed, it will be added to the concepts below. Symbol Meaning x An input vector or scalar. y A target value or label (ground truth). ˆ y A model’s predicted output. W A weight matrix containing trainable parameters. b A bias vector or scalar. θ The set of all trainable parameters in a system. L A loss function mapping predictions and targets to a scalar error. η The learning rate (a scalar governing optimization step size). E An embedding matrix containing vector representations of tokens. Q, K, V Query, Key, and Value matrices used in attention mechanisms. R d A d -dimensional vector space over the real numbers. E The expectation operator. L A likelihood function; plain L remains reserved for loss. H Entropy. D KL ( p ∥ q ) Kullback–Leibler divergence from q to p J or J f A Jacobian matrix, usually the Jacobian of function f H f The Hessian matrix of a scalar-valued function f z A pre-activation, logit, or intermediate scalar/vector, as specified locally. a An activation. g An upstream gradient or adjoint when that role is stated. ̄ v The adjoint ∂L/∂v of an intermediate quantity v δ ( ℓ ) The loss gradient with respect to a layer- ℓ pre-activation. λ A nonnegative regularization-strength parameter. p keep The probability that a unit is retained under dropout. 1 Contents Global Notation Registry 1 I Mathematical Preliminaries 10 1 What Is AI? 11 1.1 Intelligence as Prediction and Adaptation . . . . . . . . . . . . . . . . . . . . . . 11 1.2 The Taxonomy of Artificial Intelligence . . . . . . . . . . . . . . . . . . . . . . . 11 1.3 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 1.4 Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 1.5 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2 Scalars, Vectors, Matrices, and Tensors 13 2.1 The Insufficiency of a Single Number . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.2 Vectors: Lists of Numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.3 Matrices: Grids of Numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 2.4 Tensors: The Generalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 2.5 Vector and Matrix Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 2.5.1 The Dot Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 2.5.2 Matrix-Vector Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.6 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.7 Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.8 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 3 Functions and Linear Transformations 17 3.1 The Mathematical Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 3.2 Parameterized Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 3.3 Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 3.4 Affine Transformations (Adding Bias) . . . . . . . . . . . . . . . . . . . . . . . . 17 3.5 Function Composition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 3.6 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 3.7 Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 3.8 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 4 Linear Regression 20 4.1 A Prediction Problem in Ordinary Language . . . . . . . . . . . . . . . . . . . . 20 4.2 From a Pattern to a Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 4.3 Measuring Error . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 4.4 A Complete Numerical Forward Pass . . . . . . . . . . . . . . . . . . . . . . . . . 22 4.5 Derivatives of the Loss . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 4.6 Gradient Descent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 4.7 Why This Optimization Problem Is Especially Friendly . . . . . . . . . . . . . . 26 2 CONTENTS 3 4.8 From One Feature to Many . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 4.9 Batch, Stochastic, and Minibatch Training . . . . . . . . . . . . . . . . . . . . . . 28 4.10 A Probabilistic Interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 4.11 Fitting Is Not the Final Objective . . . . . . . . . . . . . . . . . . . . . . . . . . 30 4.12 Regularization and Weight Decay . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 4.13 Linear Regression as a Learning System . . . . . . . . . . . . . . . . . . . . . . . 31 4.14 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 4.15 Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 4.16 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 II Mathematical Language 33 5 Probability and Random Variables 34 5.1 Outcomes, Sample Spaces, and Events . . . . . . . . . . . . . . . . . . . . . . . . 34 5.2 Probability Rules . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 5.3 Empirical Frequency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 5.4 Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 5.5 Independence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 5.6 Bayes’ Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 5.7 Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 5.8 Discrete Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 5.9 Continuous Distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 5.10 The Gaussian Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 5.11 Expectation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 5.12 Variance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 5.13 Joint and Conditional Distributions . . . . . . . . . . . . . . . . . . . . . . . . . 38 5.14 Covariance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 5.15 A Complete Housing Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39 5.16 Uncertainty and Sampling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39 5.17 Why AI Needs Probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39 5.18 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40 5.19 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40 6 Expectation, Likelihood, and Information 41 6.1 Expectation as a Probability-Weighted Calculation . . . . . . . . . . . . . . . . . 41 6.1.1 A Decision Under Uncertainty . . . . . . . . . . . . . . . . . . . . . . . . 41 6.2 Linearity of Expectation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 6.3 Empirical Averages as Estimates . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 6.4 Parameters Turn Distributions into Families . . . . . . . . . . . . . . . . . . . . . 43 6.5 Likelihood: Data Fixed, Parameters Variable . . . . . . . . . . . . . . . . . . . . 43 6.6 A Bernoulli Likelihood Worked by Hand . . . . . . . . . . . . . . . . . . . . . . . 43 6.7 Why Log-Likelihood Is Used . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 6.8 Gaussian Likelihood and Squared Error . . . . . . . . . . . . . . . . . . . . . . . 44 6.9 Information as Surprise . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 6.10 Entropy: Expected Surprise . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 6.11 Cross-Entropy: Predicting One Distribution with Another . . . . . . . . . . . . . 46 6.12 Kullback–Leibler Divergence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46 6.13 From Information to Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 6.14 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 6.15 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 CONTENTS 4 7 Geometry of High-Dimensional Data 48 7.1 Observations as Points in Feature Space . . . . . . . . . . . . . . . . . . . . . . . 48 7.2 Distance and Feature Scaling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48 7.3 Angles and Cosine Similarity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49 7.4 Orthogonality and Independent Directions . . . . . . . . . . . . . . . . . . . . . . 49 7.5 Subspaces and Span . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 7.6 Projection onto a Direction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 7.7 Least Squares as Projection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51 7.8 A Numerical Projection Example . . . . . . . . . . . . . . . . . . . . . . . . . . . 51 7.9 High-Dimensional Concentration . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 7.10 The Curse of Dimensionality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 7.11 Intrinsic and Ambient Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 7.12 Principal Directions and Compression . . . . . . . . . . . . . . . . . . . . . . . . 53 7.13 Representation Geometry in AI . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 7.14 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 7.15 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54 III Calculus, Computational Graphs, and Autodiff 55 8 Derivatives and the Geometry of Change 56 8.1 Change Across an Interval . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 8.2 The Difference Quotient . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 8.3 Deriving the Derivative of a Square . . . . . . . . . . . . . . . . . . . . . . . . . . 57 8.4 Tangent Lines and Local Linearization . . . . . . . . . . . . . . . . . . . . . . . . 57 8.5 Derivative Notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 8.6 Basic Differentiation Rules . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 8.6.1 Constants, Sums, and Multiples . . . . . . . . . . . . . . . . . . . . . . . . 58 8.6.2 The Power Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 8.6.3 The Product Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 8.6.4 The Quotient Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 8.7 Exponential and Logarithmic Change . . . . . . . . . . . . . . . . . . . . . . . . 59 8.8 Derivatives of a Learning Objective . . . . . . . . . . . . . . . . . . . . . . . . . . 59 8.9 Stationary Points . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 8.10 Second Derivatives and Curvature . . . . . . . . . . . . . . . . . . . . . . . . . . 60 8.11 Nondifferentiable Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 8.12 Numerical Differentiation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 8.13 Why AI Needs Derivatives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 8.14 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62 9 Partial Derivatives and Multivariable Calculus 63 9.1 A Surface Instead of a Curve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63 9.2 Partial Derivatives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63 9.3 The Gradient . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64 9.4 Multivariable Local Linearization . . . . . . . . . . . . . . . . . . . . . . . . . . . 64 9.5 Directional Derivatives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 9.6 Contours and the Gradient . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 9.7 Gradient Descent in Vector Form . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 9.8 A Two-Parameter Numerical Update . . . . . . . . . . . . . . . . . . . . . . . . . 66 9.9 Matrix Derivatives for Linear Regression . . . . . . . . . . . . . . . . . . . . . . . 67 9.10 Vector-Valued Functions and Jacobians . . . . . . . . . . . . . . . . . . . . . . . 67 9.11 Second Derivatives and the Hessian . . . . . . . . . . . . . . . . . . . . . . . . . . 68 CONTENTS 5 9.12 Saddles, Plateaus, and Valleys . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68 9.13 Why AI Needs Multivariable Calculus . . . . . . . . . . . . . . . . . . . . . . . . 68 9.14 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69 10 The Chain Rule 70 10.1 A Nested Calculation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70 10.2 The Scalar Chain Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71 10.3 Local Derivatives Multiply Along a Path . . . . . . . . . . . . . . . . . . . . . . . 71 10.4 A Complete Scalar Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72 10.5 Why Gradient Contributions Accumulate . . . . . . . . . . . . . . . . . . . . . . 73 10.6 Chain Rule for Several Inputs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73 10.7 Gradient Form of the Chain Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 10.8 A Vector Chain-Rule Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 10.9 Chain Rule Through an Affine Map . . . . . . . . . . . . . . . . . . . . . . . . . 75 10.10Vanishing and Exploding Products . . . . . . . . . . . . . . . . . . . . . . . . . . 75 10.11The Chain Rule as a Computational Procedure . . . . . . . . . . . . . . . . . . . 76 10.12What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 10.13Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 11 Computational Graphs 77 11.1 Why Expressions Need Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . 77 11.2 Nodes and Directed Edges . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78 11.3 Roles of Nodes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78 11.4 Forward Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 11.5 Local Operations and Local Derivatives . . . . . . . . . . . . . . . . . . . . . . . 79 11.6 Backward Traversal of the Graph . . . . . . . . . . . . . . . . . . . . . . . . . . . 80 11.7 Why Contributions Add at a Branch . . . . . . . . . . . . . . . . . . . . . . . . . 80 11.8 Shared Intermediate Values . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81 11.9 Topological Order . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81 11.10A Prediction-and-Loss Graph . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 82 11.11Graphs and Equivalent Expressions . . . . . . . . . . . . . . . . . . . . . . . . . . 82 11.12What a Computational Graph Does Not Do . . . . . . . . . . . . . . . . . . . . . 83 11.13What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 11.14Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 11.15Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 12 Reverse-Mode Automatic Differentiation 84 12.1 Three Ways to Obtain Derivatives . . . . . . . . . . . . . . . . . . . . . . . . . . 84 12.1.1 Symbolic Differentiation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84 12.1.2 Numerical Differentiation . . . . . . . . . . . . . . . . . . . . . . . . . . . 84 12.1.3 Automatic Differentiation . . . . . . . . . . . . . . . . . . . . . . . . . . . 85 12.2 The Reverse-Mode Question . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85 12.3 The Local Reverse Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85 12.4 A Complete Reverse-Mode Example . . . . . . . . . . . . . . . . . . . . . . . . . 86 12.4.1 Forward Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86 12.4.2 Local Derivatives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86 12.4.3 Reverse Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86 12.4.4 Direct Check . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87 12.5 Why One Backward Pass Produces Many Derivatives . . . . . . . . . . . . . . . . 87 12.6 The Jacobian Perspective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88 12.7 Forward Mode . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88 12.8 Forward Mode Versus Reverse Mode . . . . . . . . . . . . . . . . . . . . . . . . . 88 CONTENTS 6 12.9 Reverse Rules for Common Scalar Operations . . . . . . . . . . . . . . . . . . . . 89 12.10Reverse Mode Through an Affine Map . . . . . . . . . . . . . . . . . . . . . . . . 89 12.11Storage and Recalculation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90 12.12Control, Nondifferentiability, and Numerical Reality . . . . . . . . . . . . . . . . 90 12.13What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90 12.14Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90 12.15Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90 13 A Micrograd-Level Scalar Autograd System 92 13.1 The Mathematical Record for One Scalar Node . . . . . . . . . . . . . . . . . . . 92 13.2 Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92 13.3 Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 13.4 Powers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 13.5 Composition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94 13.6 The Need for Accumulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94 13.7 A Rich Scalar Graph . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 13.7.1 Forward Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 13.7.2 Local Derivatives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 13.7.3 Reverse Traversal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96 13.8 The Same Graph at a Different Point . . . . . . . . . . . . . . . . . . . . . . . . . 96 13.9 A Scalar Neuron . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97 13.10Connection to the Affine Backward Formula . . . . . . . . . . . . . . . . . . . . . 98 13.11Reverse Order and Node Readiness . . . . . . . . . . . . . . . . . . . . . . . . . . 98 13.12Repeated Backward Evaluations . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 13.13What the Scalar System Captures . . . . . . . . . . . . . . . . . . . . . . . . . . 99 13.14Limitations of a Scalar View . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 13.15What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 13.16Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100 13.17Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100 14 Backpropagation from First Principles 101 14.1 From a Scalar Graph to a Layered Model . . . . . . . . . . . . . . . . . . . . . . 101 14.2 A Hand-Computable Network . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102 14.3 Complete Forward Propagation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102 14.4 Starting the Backward Pass . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103 14.5 Gradients of the Output Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103 14.6 Backward Through the Activation . . . . . . . . . . . . . . . . . . . . . . . . . . 104 14.7 Gradients of the First Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 104 14.8 All Parameter Gradients . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 14.9 The Same Derivation in Vector Form . . . . . . . . . . . . . . . . . . . . . . . . . 105 14.10Why the Order Is Reversed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 14.11Why Backpropagation Scales . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 14.12Backpropagation and Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 14.13What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 14.14Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 14.15Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 IV Neural Networks 108 15 The Perceptron and the Artificial Neuron 109 15.1 A Weighted Evidence Score . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109 CONTENTS 7 15.2 The Threshold Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109 15.3 The Decision Boundary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110 15.4 Logic Gates as Perceptrons . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110 15.5 A Perceptron as an Artificial Neuron . . . . . . . . . . . . . . . . . . . . . . . . . 111 15.6 Learning from Mistakes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111 15.7 Why the Hard Threshold Blocks Backpropagation . . . . . . . . . . . . . . . . . 111 15.8 From Decisions to Smooth Activations . . . . . . . . . . . . . . . . . . . . . . . . 111 15.9 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112 15.10Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112 15.11Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112 16 Activation Functions and Nonlinearity 113 16.1 Why Nonlinearity Changes the Function Class . . . . . . . . . . . . . . . . . . . 113 16.2 Sigmoid . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114 16.2.1 Derivative . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114 16.2.2 Saturation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114 16.3 Hyperbolic Tangent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114 16.4 Rectified Linear Unit . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115 16.4.1 Why ReLU Became Important . . . . . . . . . . . . . . . . . . . . . . . . 115 16.4.2 Dead Units . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115 16.5 GELU . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 116 16.6 Comparing Gradient Flow . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 116 16.7 Nonlinearity and Geometric Folding . . . . . . . . . . . . . . . . . . . . . . . . . 116 16.8 Choosing an Activation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 16.9 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 16.10Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 16.11Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 17 Multilayer Perceptrons 118 17.1 From One Unit to a Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118 17.2 A One-Hidden-Layer MLP . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118 17.3 Layer Terminology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119 17.4 A Complete Numerical MLP . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119 17.5 A Hidden Representation Is Learned . . . . . . . . . . . . . . . . . . . . . . . . . 119 17.6 Representing Exclusive-Or . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120 17.7 Adding More Hidden Layers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120 17.8 Width and Depth . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120 17.9 Parameter Count . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 17.10Geometric Interpretation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 17.11What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 17.12Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 17.13Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 18 Forward Propagation 123 18.1 Layer-by-Layer Notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123 18.2 A Three-Layer Numerical Network . . . . . . . . . . . . . . . . . . . . . . . . . . 123 18.3 The Shape Ledger . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124 18.4 Bias Broadcasting for a Batch . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 18.5 Forward Propagation as Composition . . . . . . . . . . . . . . . . . . . . . . . . . 125 18.6 Intermediate Values Must Be Preserved . . . . . . . . . . . . . . . . . . . . . . . 125 18.7 Output Construction Depends on the Task . . . . . . . . . . . . . . . . . . . . . 125 18.8 Forward Propagation Is Not Learning . . . . . . . . . . . . . . . . . . . . . . . . 126 CONTENTS 8 18.9 What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126 18.10Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126 18.11Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126 19 Training a Neural Network with Backpropagation 127 19.1 The Training Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 19.2 Step 1: Forward Propagation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 19.3 Step 2: Backward Propagation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128 19.3.1 Output Parameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128 19.3.2 Through ReLU . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128 19.3.3 Hidden Parameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129 19.4 Step 3: Parameter Update . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129 19.5 Step 4: A New Forward Pass . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129 19.6 Reading the Gradient Signs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130 19.7 Repeated Iterations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130 19.8 From One Example to a Minibatch . . . . . . . . . . . . . . . . . . . . . . . . . . 131 19.9 Training Error and Generalization . . . . . . . . . . . . . . . . . . . . . . . . . . 131 19.10Why Neural Training Retains the Linear-Regression Skeleton . . . . . . . . . . . 131 19.11What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132 19.12Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132 19.13Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132 20 Initialization, Numerical Stability, and Gradient Flow 133 20.1 Why Zero Initialization Fails in Hidden Layers . . . . . . . . . . . . . . . . . . . 133 20.2 Scale Matters as Much as Symmetry . . . . . . . . . . . . . . . . . . . . . . . . . 134 20.3 A Numerical Scale Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134 20.4 Forward and Backward Scale Must Both Be Considered . . . . . . . . . . . . . . 135 20.5 Xavier Initialization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 20.6 He Initialization for ReLU . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 20.7 Gradient Flow Through Depth . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136 20.8 Saturation and Vanishing Gradients . . . . . . . . . . . . . . . . . . . . . . . . . 136 20.9 Exploding Gradients . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136 20.10Finite-Precision Arithmetic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136 20.11Stable Exponential Normalization . . . . . . . . . . . . . . . . . . . . . . . . . . . 137 20.12Stable Logarithmic Expressions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137 20.13Input and Parameter Scaling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138 20.14Diagnosing Flow Layer by Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . 138 20.15What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138 20.16Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138 20.17Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139 21 Dropout, Weight Decay, and Deep-Network Regularization 140 21.1 Training Fit Is Not the Final Objective . . . . . . . . . . . . . . . . . . . . . . . 140 21.2 Weight Magnitude and Sensitivity . . . . . . . . . . . . . . . . . . . . . . . . . . 140 21.3 L 2 Regularization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 21.4 A Weight-Decay Calculation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 21.5 Geometric Interpretation of the Penalty . . . . . . . . . . . . . . . . . . . . . . . 142 21.6 Weight Decay and Scale Across Layers . . . . . . . . . . . . . . . . . . . . . . . . 142 21.7 Dropout as Random Representation Perturbation . . . . . . . . . . . . . . . . . . 142 21.8 A Numerical Dropout Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 21.9 Backward Propagation Through Dropout . . . . . . . . . . . . . . . . . . . . . . 143 21.10Why Dropout Can Regularize . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 CONTENTS 9 21.11Training and Evaluation Modes . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144 21.12Dropout Strength . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144 21.13Combining Weight Decay and Dropout . . . . . . . . . . . . . . . . . . . . . . . . 144 21.14Early Stopping and Validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144 21.15Regularization Does Not Repair Dataset Problems . . . . . . . . . . . . . . . . . 144 21.16What We Have Constructed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145 21.17Why This Was Necessary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145 21.18Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145 Part I Mathematical Preliminaries 10 Chapter 1 What Is AI? The term “Artificial Intelligence” is ubiquitous, yet its technical boundaries are frequently mis- understood. Before we construct the mathematical machinery required to build an AI system, we must establish precisely what we mean by intelligence, learning, and representation. 1.1 Intelligence as Prediction and Adaptation In the context of this book, we do not treat intelligence as a biological mystery or a philosophical phenomenon of consciousness. Instead, we define intelligence operationally as the capacity to observe data, form internal representations, and use those representations to make accurate predictions or optimal decisions in novel situations. When a human catches a thrown ball, their brain is subconsciously predicting a parabolic trajectory based on visual data. When a language model completes a sentence, it is predicting the most statistically likely subsequent word based on the context. At its mathematical core, much of artificial intelligence is the science of prediction under uncertainty. 1.2 The Taxonomy of Artificial Intelligence Artificial Intelligence (AI) is a broad academic discipline. Within it are nested subfields that rep- resent increasingly specific mathematical mechanisms. It is crucial to distinguish these bound- aries. Definition 1.1 (Machine Learning) Machine Learning (ML) is a subset of AI where systems are not explicitly programmed with fixed rules. Instead, they are equipped with a mathematical model whose internal parameters are adjusted automatically based on data, enabling the system to learn to map inputs to outputs. Definition 1.2 (Deep Learning) Deep Learning (DL) is a subset of Machine Learning utilizing Artificial Neural Networks with multiple layers. These layers act as sequential mathematical transformations, allowing the system to learn increasingly complex hierarchical representations of the data (e.g., from raw pixels, to edges, to shapes, to objects). Definition 1.3 (Transformers and LLMs) Transformers are a specific, highly successful math- ematical architecture within Deep Learning designed to process sequential data using a mech- anism called attention Large Language Models (LLMs) are massive Transformer networks trained on vast amounts of text data to predict subsequent sequences of characters or words. Our journey in this book will mirror this hierarchy. We will begin with basic Machine Learning (Linear Regression), construct the foundations of Deep Learning (Neural Networks and Backpropagation), and eventually build the complex architectures of modern AI (Transformers and LLMs). 11 CHAPTER 1. WHAT IS AI? 12 Artificial Intelligence (AI) Machine Learning (ML) Deep Learning (DL) Transformers & LLMs Figure 1.1: The nested hierarchy of Artificial Intelligence. 1.3 What We Have Constructed We have established a clear, non-anthropomorphic definition of intelligence centered on predic- tion and adaptation. We have mapped the conceptual boundaries separating AI, ML, DL, and LLMs. 1.4 Why This Was Necessary By stripping away the science-fiction connotations of AI, we prepare ourselves to treat the subject purely as applied mathematics. Knowing that ML is fundamentally about adjusting parameters based on data gives us our mandate: we must now learn the mathematical language required to represent data (Part II) and the calculus required to adjust parameters (Parts III and IV). 1.5 Exercises Exercise 1.1 (Conceptual) Explain why a standard pocket calculator, despite solving complex arithmetic flawlessly, is not considered a Machine Learning system. Exercise 1.2 (Conceptual) Based on Figure 1.1, is it possible for a system to be considered Artificial Intelligence but not Machine Learning? If so, what might such a system look like conceptually? We now develop the numerical language used to express observations and transformations. Chapter 2 Scalars, Vectors, Matrices, and Tensors To teach a machine to recognize patterns, we must first translate the physical world—images, sound waves, or text—into a language a computer understands: numbers. This chapter builds the mathematical structures used to store and manipulate these numbers. 2.1 The Insufficiency of a Single Number The Problem: Suppose we want to predict the price of a house. If the only information we have is the house’s area (e.g., 2000 square feet), we can represent this input as a single number: x = 2000. Definition 2.1 (Scalar) A scalar is a single numerical value, representing a magnitude. We denote that x is a real-number scalar by writing x ∈ R However, what if we also know the house’s age (15 years) and the number of bedrooms (3)? A single scalar is insufficient to represent multiple distinct features simultaneously. We need an ordered list. 2.2 Vectors: Lists of Numbers To group multiple features belonging to a single entity, we use a vector. Definition 2.2 (Vector) A vector is an ordered, one-dimensional array of numbers. If a vector contains d real numbers, we say it belongs to the d -dimensional space R d For our house, we define the feature vector x : x = 2000 15 3 (2.1) Because x contains 3 elements, we write x ∈ R 3 We refer to individual elements using sub- scripts. Here, x 1 = 2000, x 2 = 15, and x 3 = 3. By convention in machine learning, vectors are written as column vectors (arranged vertically). Geometric Interpretation: A vector is not just a list; it is a coordinate in space. A vector v ∈ R 2 , such as v = [ 3 2 ] , can be drawn as an arrow starting at the origin (0 , 0) and ending at the point (3 , 2). Thus, a machine learning dataset containing thousands of inputs is geometrically a ”cloud” of thousands of points in a high-dimensional space. 13 CHAPTER 2. SCALARS, VECTORS, MATRICES, AND TENSORS 14 2.3 Matrices: Grids of Numbers The Problem: A vector describes a single house. What if we have a dataset of 4 different houses? We need a structure to hold multiple vectors. Definition 2.3 (Matrix) A matrix is a two-dimensional rectangular grid of numbers. A matrix with m rows and n columns is denoted as W ∈ R m × n Let us construct a matrix X (by convention, matrices are capitalized) where each row represents one house, and each column represents a feature (Area, Age, Bedrooms). If we have 4 houses and 3 features, X ∈ R 4 × 3 : X = 2000 15 3 1500 5 2 3000 20 4 1200 2 1 (2.2) We can index into a matrix using two subscripts: X i,j refers to the entry in the i -th row and j -th column. Here, X 2 , 1 = 1500 (the area of the second house). 2.4 Tensors: The Generalization Scalars, vectors, and matrices are all specific instances of a more general mathematical object. Definition 2.4 (Tensor) A tensor is an n -dimensional array of numbers. The number of dimensions is called the order or rank of the tensor. Scalar (Order 0) x ∈ R Vector (Order 1) x ∈ R d Matrix (Order 2) X ∈ R m × n 3D Tensor (Order 3) T ∈ R c × h × w Figure 2.1: The hierarchy of mathematical representations in machine learning. Modern AI heavily utilizes 3D and 4D tensors. For example, a color image is a 3D tensor: it has height, width, and 3 color channels (Red, Green, Blue). An image of size 256 × 256 pixels is mathematically a tensor T ∈ R 3 × 256 × 256 2.5 Vector and Matrix Operations To build models, we must compute with these objects. 2.5.1 The Dot Product The most fundamental operation between two vectors of the same size is the dot product. It takes two vectors and returns a single scalar Definition 2.5 (Dot Product) Given two vectors u, v ∈ R d , their dot product is the sum of the products of their corresponding entries: u · v = u ⊤ v = d ∑ i =1 u i v i = u 1 v 1 + u 2 v 2 + · · · + u d v d (2.3) (Note: u ⊤ denotes the transpose of u , turning a column vector into a row vector, allowing standard matrix multiplication rules to apply). CHAPTER 2. SCALARS, VECTORS, MATRICES, AND TENSORS 15 Example 2.6 (Calculating a Dot Product) Let u = 1 3 − 2 and v = 4 0 5 . Both are in R 3 u · v = (1)(4) + (3)(0) + ( − 2)(5) = 4 + 0 − 10 = − 6 The result is − 6, a scalar. Geometrically, the dot product measures similarity and alignment between two vectors in space. 2.5.2 Matrix-Vector Multiplication The Problem: Suppose we have an input vector x ∈ R n , and we want to transform it into a new vector y ∈ R m . We accomplish this by multiplying x by a matrix W ∈ R m × n The rule for matrix-vector multiplication is