What if large language models had eyes, ears, and other sensors that let them perceive the real world? What kind of internal representation would they build? Could such systems become fully autonomous, and what scientific and technical breakthroughs would it take to get there?
![]()