Make money doing the work you believe in
This is, indeed, a great lecture, but it leaves out the other half of the story. AKA how our brains work similarly.
Both the human brain and autoregressive deep language models do continuous next-word prediction before word onset, both match their pre-onset predictions to the incoming word to calculate post-onset surprise, and both rely on contextual embeddings to represent words in natural contexts.
Both systems construct highly parallel internal structural maps to organize language. Models that have received enough training to achieve sufficiently high next-word prediction performance also acquire representations of sentences that are predictive of human fMRI responses.
The human brain both represents probability distributions and performs probabilistic inference.
Models possess brain-like hierarchies, and achieve functional and anatomical correspondence to human brains at high semantic abstraction levels.
Cross-entropy loss is a mathematical tool used to literally model how the brain minimizes surprise, processes information, and performs classification tasks.
And more…
Citations:



