001 · ANONYMOUS ACL SUBMISSION, UNDER REVIEW
GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation
What changes when you stop showing a vision-language model a photograph and let it walk around Street View instead. Six models, a developed/developing gap that image quality does not explain, and an action budget where more turned out to be worse.
1,200locations
100cities
6models
5 kmbest median error
56×cost of no limit