概率与统计:不确定性科学
一门全面的大学初级课程,介绍概率与统计的数学基础。课程要求具备一年微积分知识,涵盖概率模型、随机变量、期望、抽样分布、似然与贝叶斯推断,以及变量之间的关系。
课程概述
📚 内容概要
一门全面的大学本科水平课程,系统介绍概率与统计的数学基础。课程要求具备一年微积分知识,涵盖概率模型、随机变量、期望、抽样分布、似然与贝叶斯推断,以及变量之间的关系。
通过基于微积分的概率与统计推断,掌握不确定性这一严谨的数学科学。
作者: Michael J. Evans 与 Jeffrey S. Rosenthal
致谢: 作者感谢多所院校(包括多伦多大学、麦克马斯特大学和普渡大学)的评审人及同事们的贡献。同时,也特别提及多伦多大学提供的资金与基础设施支持。
🎯 学习目标
- 使用样本空间、事件和概率测度,定义一个正式的概率模型。
- 应用组合原理(排列、子集、二项系数)解决均匀概率问题。
- 利用全概率公式和贝叶斯定理分析多阶段系统,并根据新信息更新信念。
- 定义并区分离散型与绝对连续型随机变量,及其各自的概率/密度函数。
- 识别并应用关键概率分布(如伯努利、二项、泊松、正态等)来建模现实世界现象。
- 计算多元分布的边缘密度、条件分布,并评估独立性。
- 计算离散、连续及混合型随机变量的期望值、方差和协方差。
- 应用“无意识统计学家定律”(LOTUS)和线性性质,计算变换后变量的期望。
- 利用概率生成函数(PGF)和矩生成函数(MGF)推导各阶矩。
- 定义并推导独立同分布序列函数的抽样分布。
课程 共 11 课时 · 预计 33.0h
课程
Lesson
This lesson introduces formal probability models as a rigorous framework to replace subjective intuition, highlighting the relative frequency interpretation and the Law of Large Numbers. Students learn to apply these mathematical structures to quantify uncertainty and manage risk in complex, real-world scenarios where human cognitive biases often fail.
This lesson introduces random variables as deterministic functions that map sample space outcomes to real numbers, providing a quantitative framework for probability. Students learn to utilize indicator functions, understand probability distributions, and apply the continuity of probability to analyze complex events.
This lesson introduces mathematical expectation as the long-run average of a random variable and explores its core properties, including linearity, monotonicity, and independence. Students learn to apply the Law of the Unconscious Statistician (LOTUS) to efficiently calculate the expected values of transformed variables without needing to derive their specific probability distributions.
This lesson introduces sampling distributions as the probability laws governing statistics, which are functions of independent and identically distributed (i.i.d.) random variables. Students learn to derive exact distributions for small sample sizes and explore how these concepts bridge the gap between raw data and statistical inference.
This lesson introduces statistical inference as the formal process of using sample data to estimate the underlying probability distributions and mechanics of a system. It emphasizes that inference is necessary to distinguish between inherent random variation and structural uncertainty, allowing researchers to make robust predictions beyond simple descriptive summaries.
This lesson explores likelihood-based inference, focusing on how the likelihood function quantifies the support for different parameter values given observed data. Students learn to use log-likelihoods for computational efficiency, apply the Central Limit Theorem for asymptotic inference, and utilize Fisher Information to measure the precision of statistical estimates.
This lesson introduces the Bayesian paradigm, which treats unknown parameters as random variables rather than fixed constants to allow for direct probability statements about them. Students learn to construct a complete Bayesian model by combining a sampling model with a prior distribution to update beliefs through the joint distribution.
This lesson explores the mathematical foundations of optimal statistical inference by defining the Mean Squared Error (MSE) as the sum of an estimator's variance and squared bias. Students learn how to minimize this error to identify the best estimators and understand the role of sufficiency and posterior means in decision theory.
This lesson introduces model checking as a critical validation step that ensures statistical inferences are grounded in reality rather than mathematical fiction. Students will learn to distinguish between parameter estimation and model validation, emphasizing that even the most precise calculations are meaningless if the underlying model assumptions do not accurately reflect the data-generating process.
This lesson defines a statistical relationship as any change in the conditional distribution of a response variable $Y$ when a predictor $X$ varies, moving beyond simple correlation to include shifts in mean, variance, or shape. It also emphasizes that establishing causality requires rigorous experimental design, such as blinding and blocking, to account for confounding variables and eliminate bias.
This lesson introduces stochastic processes as systems that evolve over time through probabilistic rather than deterministic rules, with a primary focus on the Simple Random Walk. Students learn to calculate path probabilities and expected values while exploring key concepts like the parity rule, Martingale fairness, and the foundational mechanics of the Gambler’s Ruin model.