English

Alignment midtraining for animals

Computation and Language 2026-05-05 v3 Artificial Intelligence

Abstract

We investigate the robustness of value alignment via midtraining with synthetic documents, using animal compassion as a value that is both important in its own right and orthogonal to existing alignment efforts. To evaluate compassionate reasoning, we develop and publicly release Animal Norms In Moral Assessment (ANIMA), a 26-question evaluation spanning 13 ethical dimensions, publicly available as a dataset and Inspect evaluation. On ANIMA, training with 3000 documents achieves 77% compared to 40% for instruction-tuning approaches, with generalization to human compassion and no degradation in standard safety benchmarks or capabilities. However, subsequent unrelated instruction-tuning degrades the intervention, with the advantage disappearing after 5000 samples. Our exploratory results suggest document-based value interventions may require explicit preservation strategies to remain effective through typical training pipelines.

Cite

@article{arxiv.2604.13076,
  title  = {Alignment midtraining for animals},
  author = {Jasmine Brazilek and Miles Tidmarsh},
  journal= {arXiv preprint arXiv:2604.13076},
  year   = {2026}
}

Comments

34 pages

R2 v1 2026-07-01T12:09:25.181Z