English

NeuroX: A Toolkit for Analyzing Individual Neurons in Neural Networks

Computation and Language 2018-12-27 v1

Abstract

We present a toolkit to facilitate the interpretation and understanding of neural network models. The toolkit provides several methods to identify salient neurons with respect to the model itself or an external task. A user can visualize selected neurons, ablate them to measure their effect on the model accuracy, and manipulate them to control the behavior of the model at the test time. Such an analysis has a potential to serve as a springboard in various research directions, such as understanding the model, better architectural choices, model distillation and controlling data biases.

Keywords

Cite

@article{arxiv.1812.09359,
  title  = {NeuroX: A Toolkit for Analyzing Individual Neurons in Neural Networks},
  author = {Fahim Dalvi and Avery Nortonsmith and D. Anthony Bau and Yonatan Belinkov and Hassan Sajjad and Nadir Durrani and James Glass},
  journal= {arXiv preprint arXiv:1812.09359},
  year   = {2018}
}

Comments

AAAI Conference on Artificial Intelligence (AAAI 2019), Demonstration track, pages 2