The Observation Theory EncyclopediaFrom TSKAboutBy kindBy chapterBy Lean fileLedgerProvenance

latent semantic analysis instrument

DefinitionThe truncated singular value decomposition of the term-document matrix, ordinarily uncentred, which coincides with principal components on documents only after centring. Chapter 12. Also latent semantic.
ExampleA term-document matrix of 20000 terms by 5000 documents reduced to 100 singular components gives each document 100 coordinates.
BookData Mining as Observation, draft 0.2, commit f3914f0; entry id latent-semantic-analysis, kind instrument.
Statusno ledger row names this entry. Corrections: none recorded.
Defining equationnone
Assumptions and scope
  • The truncated singular value decomposition of the term-document matrix, ordinarily uncentred, with the top components kept. It coincides with principal components on documents only after the matrix is centred, or under an explicitly uncentred convention, and the fraction of squared singular values kept grows with the components kept and reaches one at full rank.
  • It is the identity reader on term-document variance, and the components it keeps are the directions of largest variance whether or not the consumer reads them. The non-oracle test fit a frozen embedding of 100 components on the training split only and recovered the reader blind.
Prior artnone recorded
Evidencegeometric-observation/chapters/ch10_the_blind_probe.md:100-118, geometric-observation/claims/LEDGER.md, lean/DataMiningAsObservation/PCA.lean, lean/DataMiningAsObservation/ExplainedVariance.lean, lean/DataMiningAsObservation/TFIDF.lean
Reviewedsemantic review 2026-09-06; generated 2026-09-10 from records at the commits on the provenance page.
directioneigenvaluekept | dropped
The top singular components of the term-document matrix.

Equation

none

Conditions

Conditions are curated in entries.toml rather than read from a record.

Ledger

none

First stated

Deerwester, Dumais, Furnas, Landauer, and Harshman, indexing by latent semantic analysis, 1990, as chapter 12 section 12.1 of Data Mining as Observation reads it, with the program’s frozen embedding in geometric-observation/chapters/ch10_the_blind_probe.md:100-118.

Measurements

Where the book states it Numbers, as the book’s sources table records them Source
chapter 12 section 12.3 041 non-oracle, frozen LSA TF-IDF to SVD 100 train-only, AUROC 0.975 vs 0.910, flip tied, magnitude overshot, partial geometric-observation/chapters/ch10_the_blind_probe.md:100-118; geometric-observation/claims/LEDGER.md row GO-B-blind 041

Failures and corrections

none

Invariance envelope

none declared

Machine checked

lean/DataMiningAsObservation/PCA.lean, theorems varAlong_basis, varAlong_le, varAlong_ge, dropped_eq, dropped_nonneg, at observation-data-mining f3914f0; what the check covers is stated in the book’s appendix C.

lean/DataMiningAsObservation/ExplainedVariance.lean, theorems explained_mem_unit, explained_mono, explained_full, retained_identity, retained_example, at observation-data-mining f3914f0; what the check covers is stated in the book’s appendix C.

lean/DataMiningAsObservation/TFIDF.lean, theorems weight_everywhere, weight_nonneg, weight_antitone, tf_mem_unit, at observation-data-mining f3914f0; what the check covers is stated in the book’s appendix C.

Used in

Data Mining as Observation primer L, chapters 0, 12.

Related

TF-IDF; principal component analysis; explained variance; bag of words; blind probe.

See also

Book equations stated beside the entry’s terms, not defining it: 0.36, 0.5, 4.2.

Ledger rows that cite the entry’s records without naming it: GO-B-legal (035→036).

Status

Generated 2026-09-10 by encyclopedia/generate.py; book at observation-data-mining f3914f0; the commit of every record is listed in the encyclopedia’s provenance.

← Laplacianlaw of large numbers →