Martian argues that interpretability research needs to move past static mechanistic analysis toward understanding agentic, long-horizon task behavior in models.
Need help?
Contact usMartian argues that interpretability research needs to move past static mechanistic analysis toward understanding agentic, long-horizon task behavior in models.