# AI — Stale course audit

- URL: https://getstale.tech/run/web_20260428T132556Z
- Target role: Ml / Data Engineer
- Course focus: Artificial Intelligence — Broad Introductory Survey
- Audit date: 2026-04-28T13:25:56Z
- Verified findings: 1
- Structured data: https://getstale.tech/run/web_20260428T132556Z.json

## Findings

### 1. Incorrect (medium severity)

**Location:** Week 4-Informed Search Algorithms.pdf, Page 37 (Properties of A* Search)

**What the slide says:**

> Optimal? YES if used Admissible Heuristic

**Primary source:** https://inst.eecs.berkeley.edu/~cs188/fa22/assets/notes/cs188-fa22-note02.pdf (verified)

> An additional caveat of graph search is that it tends to ruin the optimality of A*, even under admissible heuristics. ... Hence, to maintain optimality under A* graph search, we need an even stronger property than admissibility, consistency.

**What to learn instead:** State the optimality claim conditional on the search variant: A* tree search is optimal with an admissible heuristic; A* graph search (which the course teaches in Week 3 with a closed/reached set to avoid revisiting states) is optimal only with a consistent (monotonic) heuristic h(n) <= c(n,n') + h(n'), because an admissible-but-inconsistent heuristic can cause a state to be locked into the reached set via a suboptimal path before the optimal path is discovered.

## Market fit

This curriculum delivers a faithful Russell & Norvig classical/symbolic-AI tour with an introductory ML chapter, but the ml_data job market it lists on its own job-description slides has moved on — most demanded skills (Python, pandas, scikit-learn, PyTorch, hugging face) sit outside this survey's depth bound, while the three legitimate, in-bound extensions are deepening statistics, and bridging the existing KR/IR material to RAG and vector-database concepts.

### Statistics — extending the partial DS treatment beyond named-only foundations and a handful of formulas to a conceptual tour of distributions, sampling, and inferential statistics (still survey-depth, no proofs) (high)

The course names statistics as a DS foundation and operationally uses a few statistical objects — Z-score normalization (Week 12 p.12), MAE/MSE means (Week 12 p.59), median imputation (Week 12 p.12), and the assertion 'Data Scientists have a strong background in statistics' (Week 12 p.65) — but never introduces distributions, sampling, hypothesis testing, or inference even at a conceptual level. Market demand is 75% (9/12 ml_data postings). Extending the existing DS-foundations slide with a conceptual overview of descriptive vs inferential statistics fits the survey depth bound (no formal proofs, no Python tutorials).

### Retrieval-Augmented Generation (RAG) — connecting the existing knowledge-representation and information-retrieval material to modern retrieval-augmented LLM pipelines at a conceptual level (high)

The course teaches knowledge representation deeply (Week 7 FOL KB; Week 8 ontologies — 'An ontology = concepts+properties+axioms+values' p.39; WordNet/ConceptNet/CYC reusable KBs p.29) and separately names information retrieval and question-answering as NLP applications (Week 10 p.23: 'Index and search large texts ... Question answering; Knowledge acquisition'). These two partially-covered threads are the natural conceptual bridge to RAG and make it a fair in-scope extension. 5/12 ml_data postings demand RAG. Survey-depth conceptual coverage only — no LangChain code, no embedding-math derivations.

### Vector databases / embedding-based knowledge representation — extending the symbolic-KR unit with a conceptual contrast between ontology/FOL representations and dense-vector representations used for semantic search (medium)

Week 8 teaches one specific paradigm of knowledge representation — ontologies and OWL ('Ontology is an explicit specification of conceptualization', p.3) — and Week 10 names information retrieval as an NLP application (p.23). The course's KR taxonomy (Week 7 p.3 contrasts Propositional / First-Order / Temporal / Probability / Fuzzy logics by 'what they commit to as primitives') is the obvious place to add a conceptual sibling: vector / embedding representations and how they enable semantic similarity search. 5/12 ml_data postings demand vector databases. This is a conceptual extension of two partially-covered topics, not a Pinecone tutorial — it stays inside the depth bound.

## Recommended topics

The three in-bound prescriptions all extend topics the course already names but stops short on: (1) lifting the existing 'statistics is foundational' slide into a conceptual descriptive-vs-inferential tour (75% of ml_data postings), (2) bridging Week 7/8 knowledge bases to Week 10's already-named IR/QA via a conceptual RAG slide (42%), and (3) adding vector/embedding representations as a sibling row to Week 7's KR-paradigm taxonomy (42%). Together they touch all 12 ml_data sample postings through the statistics+RAG+vector union.

### #1 Descriptive vs inferential statistics — conceptual tour of distributions, sampling, and hypothesis testing (~4h to learn)

The course already names statistics as the spine of data science — extend that one slide into a conceptual descriptive-vs-inferential tour so students leave with a vocabulary the 75% of ml_data postings demanding 'statistics' actually mean.

Keywords: statistics, distributions, sampling, hypothesis testing, inference, descriptive statistics

Where it fits: Artificial Intelligence (Week 12 — Intro to DS and ML) · Week 12 already names statistics as a DS foundation (p.7 'Principles can be statistical, computational, algorithmic, visual'; p.65 'Data Scientists have a strong background in statistics') and uses a few statistical objects operationally (Z-score normalization p.12, MAE/MSE p.59, median imputation p.12). The gap is that the slide naming statistics never opens into descriptive vs inferential framing. A conceptual one-section extension on the existing DS-foundations slide (what a distribution is, what sampling is, what inference is for) extends partial coverage without breaking survey depth.

### #2 Retrieval-Augmented Generation (RAG) as the modern bridge between knowledge representation and NLP question answering (~3h to learn)

Week 10 already lists 'index large texts' and 'question answering' on the same slide and Weeks 7–8 already teach knowledge bases — a single conceptual slide on RAG ties those threads to what 5 of 12 ml_data postings now ask for.

Keywords: rag, retrieval-augmented generation, information retrieval, question answering, llms, knowledge bases

Where it fits: Artificial Intelligence (Week 10 — Intro to NLP) · Week 10 already names the exact ingredients of RAG without connecting them: p.23 lists 'Index and search large texts ... Information extraction ... Question answering; Knowledge acquisition' as NLP applications. Pair that with Week 7/8's deep treatment of knowledge bases (FOL KB, OWL ontologies, WordNet/ConceptNet/CYC reusable KBs at Week 8 p.29) and the conceptual story 'retrieve from a knowledge source, then have a generator answer' is a one-slide extension of an already-named application. Stays at survey depth — no LangChain code, no embedding math.

### #3 Vector / embedding-based knowledge representation as a sibling paradigm to symbolic KR (~3h to learn)

Week 7's KR taxonomy already invites the question 'what other ways can knowledge be represented?' — adding embeddings/vector stores as a conceptual sibling closes the gap to the 5 ml_data postings demanding vector databases without leaving the survey lane.

Keywords: vector databases, embeddings, semantic search, dense representations, knowledge representation

Where it fits: Artificial Intelligence (Week 8 — Knowledge Representation: Ontology) · Week 7 p.3 already presents a taxonomy of representation paradigms ('what they commit to as primitives' — Propositional / First-Order / Temporal / Probability / Fuzzy) and Week 8 deepens one specific branch ('Ontology is an explicit specification of conceptualization', p.3). Adding a conceptual sibling row — dense-vector / embedding representations and the semantic-similarity-search retrieval they enable — is the natural extension of an existing taxonomy slide. It also gives the Week 10 IR/QA application slide a second concrete substrate beyond keyword indexing. No Pinecone tutorial, no math — just the conceptual contrast 'symbolic vs distributed'.

---
Produced by Stale (https://getstale.tech). Request a course audit: https://getstale.tech/request-audit
