
Principa
—
Research scientist
Links
Focus
Interpretability
Nischal is a research scientist at Principia interested in theoretical understanding of AI systems and using that understanding for interpretability, i.e., theory-first interpretability. He is currently interested in solvable models of learning dynamics and their applications, and mean-field theories of learning and representation formation, and would be interested in either pushing the theory frontier and developing basic understanding of safety-relevant phenomena (e.g., silent alignment, representation structure in mean-field networks, solvable models of superposition, etc.) or applying various ideas coming from theory for practical purposes (early detection of sudden changes in NNs during learning, new weight space/representation space interpretability tools, etc.).