Training on probes: Research ideas
Alignment Forum1w4 min read
Recap Sequel to Previous Post . This post might not make sense without it. Last post, I told some stories about how training on probes might let our judgments on easy domains generalize to harder domains by leveraging an AI's model of the world. Training using the gradient of the probe teaches the model to fool the probe, but it turns out doing RL against the probe doesn't teach the model to fool it! Following the brain-like story The brain-like story about training on probes analogizes them to instincts. I have hardwired instincts for eating sugar, and learned knowledge about candy stores, an
