Labelbox research examines Natural Language Autoencoders (NLAs) and their use in interpreting LLM behavior. The researchers investigated whether activation patterns at specific neural network layers could identify branchpoints where models decide between legitimate solutions and shortcut behaviors, using Gemma 3 27B and Terminal Wrench tasks as their test case.