Notes on AI Interpretability and Building

Jakub Dvořák
Jakub Dvořák's avatar
Engineer & founder moving into AI-safety research — how neural networks represent what they know (mechanistic interpretability, superposition). Prague · MFF UK.