Findings of EMNLP 2026

What happens when you let a model look around?

Every geolocation benchmark hands an AI one frozen photo and asks where it was taken. But nobody plays GeoGuessr that way. GeoAgent drops vision-language models into Street View and lets them walk.

Arka Mukherjee · Soham Roy · Kartikeya Trivedi · Shreya Ghosh Published 21 August 2026 · Findings of EMNLP 2026 · CC BY 4.0
10 km 100 km 1,000 km Gemini 3.0 Flash 5 km GPT-5 Mini 182 km Gemma 3 27B 402 km Geo-R1 7B 576 km Llama 4 Scout 580 km Claude Haiku 4.5 688 km Human players 996 km Random guess 10,210 km
Median error · log scale · agentic navigation
closer to the centre is better
1,200Street View
locations
100cities, six
continents
6models
evaluated
9navigation
actions
8actions per
round, max
5experimental
conditions
TL;DR

We introduce GeoAgent, a collection of geolocalization data and an embodied navigation environment, showing significant limitations and insights into LLM-based agentic navigation.

Keywords
embodied agentsautonomous agents agent evaluationcorpus creation benchmarking
Abstract

Modern Vision-Language Models (VLMs) perform well above the human baseline in geolocalization tasks. However, most efforts to study AI behavior on the task remain limited to static image-based retrieval, classification, and predictions. We argue that faithful recreation of the task should involve embodied navigation, where a multimodal agent autonomously explores its surroundings to gather observations before submitting a prediction. To this end, we introduce GeoAgent, an agentic environment-based benchmark that requires agents to navigate Street View environments to refine their geolocalization through reasoning sequentially. Our analysis shows that modern VLMs struggle to discern regional patterns while succeeding at country- and continent-level predictions. When compared to static image-based baselines, agentic navigation significantly improves accuracy across established metrics. We also note severe bias in a developed/developing region context across frontier model architectures and poor self-improvement capabilities given incorrect priors. Overall, our work establishes the challenges of embodied navigation and geospatial reasoning.

The setup

A photograph can’t tell you what’s behind you.

Static benchmarks measure recognition. The thing that actually makes geolocation hard is deciding where to look next.

When a person plays GeoGuessr, the opening frame is rarely the useful one. They spin around hunting for a street sign, walk to the junction to read a shopfront, glance down to check which side of the road the traffic keeps. The identifying detail is somewhere in the scene — it just isn’t in the direction the camera happens to be pointing.

Static benchmarks freeze that decision away. A model gets a fixed set of views and must infer from whatever they happen to contain; if the decisive cue lies around the corner, it is unreachable forever. Even video-based geolocalization keeps the model a passive observer riding somebody else’s camera path.

GeoAgent restores the missing half. The model is dropped into a fully navigable Street View panorama with a toolbox, and after every move it sees a fresh screenshot and revises its guess.

Passive inference tests recognition, not reasoning under uncertainty.
The paper’s central argument, paraphrased
The static baseline · from the paper
A 4-view composite: four Street View screenshots at 0, 90, 180 and 270 degrees stitched into one strip.
The 4-view composite. Four cardinal headings from the spawn point, stitched into a single image. One inference, no recourse — if the decisive cue is round the corner, it does not exist.
An 8-view composite: the four spawn-point views on the top row, and four more captured after one MOVE_FORWARD action on the bottom row.
The 8-view composite adds a second panorama taken after one MOVE_FORWARD. More pixels, still no agency — and as the explorer below shows, the gain over 4-view is small. It isn’t the quantity of imagery that matters.
The alternative · embodied

Nine actions. Eight turns.

Rotate, walk, tilt, or return to the start — then commit. A running hypothesis, revised after every observation, with the agent choosing what to look at next.

ROTATE_LEFTMOVE_FORWARD LOOK_DOWNGUESS
The environment

Nine actions, eight turns, one commitment.

The model never touches a mouse. It emits a tool call, the environment moves the camera, and a fresh screenshot comes back.

Figure 1 · the framework
The GeoAgent framework: a Street View screenshot feeds a VLM, which produces an Observation and Reasoning, then issues a toolcall from a toolbox of actions back to the environment.
The ReAct loop. A screenshot goes to the model; it writes an observation and its reasoning, then picks an action from the toolbox; the environment moves the camera and returns a new screenshot. Repeat, up to eight times.
Figure 2 · the live interface
The GeoAgent web interface: a Street View panorama of a residential street in Inverness, with a round counter, score, navigation arrow, world mini-map and a MAKE GUESS button.
What the agent actually sees. A real navigable panorama — this one from Inverness, Scotland — with the round counter, the movement arrow, and the guess control. Every sample carries full metadata: continent, country, city, tier, development status, capture date, and both centre and camera coordinates.
ROTATE:<degrees>Rotate to a specific heading, 0–360°
ROTATE_LEFTRotate 90° left from current view
ROTATE_RIGHTRotate 90° right from current view
ROTATE_BEHINDRotate 180° to look behind
LOOK_UPTilt the camera up by 20°
LOOK_DOWNTilt the camera down by 20°
MOVE_FORWARDMove forward along the road or path
RETURN_STARTReturn to the starting position
GUESSSubmit the final location, end the round

Every turn the model returns the same five JSON fields — what it observes, its reasoning, a confidence, the action it wants, and its current guess — carrying up to five prior observations as context. Look, think, move.

One design choice matters more than it looks: the model never emits coordinates. It names a place, and the environment resolves that name through the Geocoding API. This kills an entire failure mode where a model identifies the right city then hallucinates wrong numbers for it.

Step through a round to see the loop tighten →

Step 1/6 START confidence medium
39°heading
observations
reasoning
guess
An example trajectory, in the JSON format the environment enforces. Illustrative reconstruction — the paper publishes only fragments of a worked example, so intermediate steps demonstrate the loop rather than transcribe a logged run.
The dataset

Built to defeat memorisation.

A benchmark of famous landmarks measures how much of the internet a model swallowed. This one is weighted deliberately the other way.

100 cities, stratified into five recognisability tiers by annual visitor numbers — a proxy for how heavily a city features in training data. Then the distribution is inverted: Tier 5 supplies 40% of all samples, Tier 1 just 2.85%.

For each city, 100 coordinates were sampled within a two-mile radius; points without navigable Street View were dropped, leaving 1,200 locations. Cities were also tagged developed or developing per UNCTAD — 46 and 54 — which sets up the study’s most uncomfortable finding.

Read the numbers carefully

Because the hard tier is deliberately over-weighted, absolute accuracies here are not real-world performance estimates. They are a stress test. The paper asks that results always be reported stratified by tier and development level.

Figure 5 · from the paper
A heatmap of average accuracy for six models across five recognisability tiers, in four navigation modes. Accuracy falls from Tier 1 to Tier 5 for nearly every model.
Accuracy decays down the tier ladder for nearly every model in every mode. GPT-5 Mini holds up best; Gemma 3 27B and Claude Haiku 4.5 fall hardest. The values are accuracy averaged across continent, country and city.
Tier 5 — the biggest slice. 480 of the 1,200 samples come from the least-visited cities in the set: Inuvik, Siem Reap, Butare, Imphal. A model cannot coast on landmark recall here; it has to read vegetation, road markings, utility poles. Accuracy falls steadily from Tier 1 to Tier 5 for nearly every model tested.
Interactive · the main result

Switch the condition. Watch what embodiment changes.

Same models, same 1,200 locations, five different ways of looking. Pick a condition and a metric — bars are on a shared scale, so they stay comparable as you switch.

Condition
Metric
City-level accuracy (%) higher is better ↑
Show the full results table (all 5 conditions × 6 models)
ModelMean dist.MedianReas. len.Cont.CountryCity

Exploring helps — most of the time.

Gemini 3.0 Flash is the clearest case: city accuracy climbs from 39.5% on the static composite to 50.1% when allowed to navigate, a 26.8% relative gain, with median error collapsing from 118.5 km to 5 km.

But two of six models do not benefit. Llama 4 Scout gets worse with the controls. And Geo-R1 7B — the model purpose-built for geolocation — is the only one that degrades against both static baselines. Static skill does not transfer.

The control condition

Is it the reasoning, or just more pixels?

To find out, the team ran a random walk: the same number of moves, chosen at random. Every general-purpose model beat its own random-walk twin — GPT-5 Mini improved from 1,581 km to 900 km mean error.

Try it above: switch between Random walk and Agentic. The exploring is doing real work.

5 km Gemini 3.0 Flash’s median error, once it could move

Down from 118.5 km on a static composite of the same location. That is the difference between naming the right region and standing on the right street.

Finding one

Exploration refines a hypothesis. It rarely rescues one.

The sharpest result in the study only becomes visible once a model can act over time.

Split every round by how it ended. For rounds the model eventually got right, extra actions steadily pull the guess toward the truth. For rounds it got wrong, extra actions do almost nothing — the curve flattens, and for some models it reverses.

What separates them is where the model started. A logistic regression predicting final accuracy from initial distance gives β = −0.43, p < 0.001.

The failure mode has a name

Confirmation bias in sequential reasoning

Models struggle to abandon an initial misconception even when shown contradictory visual evidence. Gemini 3.0 Flash is the starkest instance — on rounds it ultimately gets wrong, its guesses actively collapse beyond about five actions.

01,250 2,5003,7505,000 Distance of the model’s FIRST guess (km) Rounds that ended CORRECT median 847 km Rounds that ended WRONG median 2,341 km
The round was already decided on action one. Boxes span the interquartile range; the heavy line is the median. Rounds that would go on to fail were already almost three times further out at the very first guess — before any exploring had happened.
Figure 3 · from the paper
Two line charts of improvement from the initial guess across eight actions. Left panel, correct city guesses: every model climbs steadily toward 100 percent. Right panel, incorrect guesses: all models stay flat near zero, and Gemini 3.0 Flash collapses to about minus 190 percent after action six.
The asymmetry, in one picture. Left: on rounds that end correct, every model converges steadily toward the target with each extra action. Right: on rounds that end wrong, the same models flatline — extra exploration buys almost nothing. Gemini 3.0 Flash is the dramatic outlier, its guesses collapsing past action six rather than recovering. Exploration compounds a good start and entrenches a bad one.
Finding two

Every model was twice as lost in the developing world.

Not slightly worse. Not only on the hardest tier. Systematically, across architectures, open and closed alike.

Developed regions Developing regions
01,000 2,0003,0004,000 Mean error (km) 646 1,317 GPT-5 Mini 2.0× worse 1,308 2,537 Gemma 3 27B 1.9× worse 1,622 2,472 Llama 4 Scout 1.5× worse 1,783 3,713 Claude Haiku 4.5 2.1× worse
Mean Haversine error by development status, agentic navigation. Accuracy deltas are messier across metrics — Gemma 3 27B actually scores higher continent and country accuracy in developing regions — but on raw distance the direction is unanimous.

The obvious objection is that this is Street View’s fault, not the models’. Coverage is famously uneven; surely developing-region panoramas are older and blurrier?

The authors went and checked, measuring image age, Google’s internal quality score, and empirical sharpness across all 1,200 panoramas. The result contradicts the assumption.

The data does not explain the gap
2.51 vs 3.09Image age, yrs
developing newer · p=.034
0.982 vs 0.979Quality score
developing higher · p=.030
1,540 vs 1,499Sharpness
no difference · p=.70

Developing-region panoramas are statistically newer and marginally higher quality. The gap is in what the models learned to see, not in the pixels.

Figure 4 · from the paper
A heatmap of developed minus developing performance deltas for six models across four navigation modes, covering continent, country and city accuracy and mean distance.
Developed–developing deltas across every model and navigation mode. The mean-distance column is the one to read: it is negative almost everywhere, meaning lower error in developed regions. Gemini 3.0 Flash carries the largest country-level gaps (over 20 points in every mode), which suggests its strong overall numbers are disproportionately driven by developed-region familiarity. Note the paper’s own accuracy sign convention here is inconsistent with its caption, so the distance column — where every source in the paper agrees — is the safest read.
Finding three

One model spun in circles. Another just walked.

Identical tools, identical budget — and strikingly different habits. The habits predicted the scores.

Rotation and movement are not equivalent. Spinning in place resamples a panorama the model has largely already seen. Walking forward reaches genuinely new information — a different junction, another shopfront, a new sign.

The correlations agree: move actions associate positively with every accuracy measure, more strongly than raw action count. But all effects are modest (|ρ| < 0.25), which the paper reads as action quality mattering more than action quantity.

A quieter pattern sits in the same data: models explore developed-region locations slightly more (5.95 vs 5.65 actions), and when they do fully explore a developing-region scene they rotate rather than walk, 2.1 : 1 against 1.4 : 1.

Forward moves Rotations
01.5 34.567.5 Mean actions per round (of 8) Llama 4 Scout 0.41 moves6.14 rotations Gemma 3 27B 4.20 moves0.76 Gemini 3.0 Flash 3.57 moves0.39 Claude Haiku 4.5 2.79 moves1.59 GPT-5 Mini 2.42 moves1.51
Llama 4 Scout burns 7.55 of its 8 actions almost entirely on rotation, pivoting on the spot without ever walking down the street. Gemini 3.0 Flash — the best city-level guesser — does the opposite.
56× More expensive per round — for worse answers

Lifting the 8-action cap tripled Llama 4 Scout’s actions to 23 per round and worsened its error by 76%. GPT-5 Mini doubled its actions and got 35% worse. Only one of three models improved, and it did so while taking fewer actions. The cap is not a compromise — it is the finding.

Finding four

They know the continent. They rarely know the town.

Country accuracy runs 52–81%. City accuracy runs 14–50%. That cliff is the most consistent shape in the study — and humans fall off it too.

Macro cues are easy and abundant: vegetation, climate, driving side, general building stock. What fails is municipal infrastructure, local business signage, regional architectural variants.

The TF-IDF analysis over the models’ own reasoning text supplies the mechanism, and it is uncomfortable reading. Models frequently produce exactly the right visual descriptors — “tropical vegetation”, “weathered infrastructure”, “motorcycle traffic” — and then map them to the wrong conclusion.

Thai locations are labelled Vietnamese 23% of the time; Vietnamese ones Thai 18%. And the countries with the worst misclassification rates — Mozambique at 67%, Thailand at 54% — have no country-specific vocabulary at all in the models’ reasoning. When distinctive cues are missing, models fall back on generic regional description, and generic description cannot disambiguate.

Real cues do surface

The models aren’t hallucinating — they’re under-resolving.

The vocabulary analysis picks up genuinely diagnostic detail: gingko for Japan, tram tracks and shutters for Switzerland, the Chao Phraya for Thailand.

Vocabulary overlaps only 36.6–52.7% between agentic and static conditions, so these are model habits — not artefacts of the environment.

Claude Haiku 4.5 · Egypt
Its top descriptors for Egyptian locations are Georgian — a systematic Egypt/Georgia confusion completely invisible in aggregate accuracy.
orthodoxcaucasustbilisi
Gemma 3 27B · China
A Japan-prior leak: Chinese street scenes described in Japanese vocabulary, suggesting East Asian imagery collapses onto one dominant prior.
japaneselanternsjapanese writing
Llama 4 Scout · China
Characteristic terms include the Google logo — the model is reading Street View’s own interface furniture rather than the scene inside it.
google logofeatures signs
GPT-5 Mini · anywhere
The strongest country-level model anchors at sub-city granularity — naming districts and street furniture, not just nations. Granularity tracks accuracy.
haussmanncentral parissoi
The human baseline

Every model beat the people.

Six undergraduate volunteers — beginner to intermediate GeoGuessr players, required to think aloud into a speech-to-text tool so their reasoning chains could be recorded — played the same rounds. Their median error was 996 km; their city accuracy 7.1%. Every model, in nearly every condition, beat them.

Two caveats keep this honest. No expert players were recruited, so the supported claim is that frontier models exceed non-expert humans — skilled GeoGuessr players are formidable, and that comparison stays open. And between-annotator variance was large, from 0.20 to 0.71 continent accuracy.

Notably, the country-to-city collapse shows up in the humans too: 46.8% → 4.3% for the strongest annotator. That asymmetry looks like a property of the task itself, not a deficiency of models.

996 kmHuman median
error
5 kmBest model
median error
7.1%Human city
accuracy
50.1%Best model city
accuracy
The annotation portal
The human annotation interface: a navigable Street View panorama with an OBSERVATIONS panel asking what clues the player sees, a world mini-map with a placed marker, and a MAKE GUESS button.
Humans played in the same environment, with one addition: a panel for recording the clues they actually used before placing a guess. Annotators were recruited across the skill spectrum, assessed on a geographic test beforehand, and each scored a stratified subset of the same 1,200-sample pool. They were paid in lunch coupons.
Caveats & ethics

What this study does not show.

01

It is not a map of the world

1,200 locations across 100 hand-chosen cities is an evaluation harness, not geographic coverage. The tier weighting over-represents obscure cities on purpose.

02

The humans were not experts

Six undergraduates, beginner to intermediate. The supported claim is about non-expert humans, and the authors say so themselves.

03

The action space is simplified

Nine actions mirror Street View’s capture topology but exclude zoom, multi-step planning, and cross-modal lookup — all things a human uses freely.

04

On misuse

Better agentic geolocalization lowers the barrier to surveillance and unauthorised location inference. The authors recommend restricted API access and geofencing downstream.

Dataset restrictions, stated by the authors

GeoAgent runs only on public Street View imagery, which Google already blurs for faces and licence plates, and processes no user-generated content. The dataset should not be used to train or fine-tune models, to power production geolocation services, or in any application that could identify individuals incidentally captured in Street View imagery. Metadata is released; raw imagery is not redistributed.

In summary

Five things to take away.

01

Embodiment changes what you can measure

Letting a model act exposes confirmation bias, exploration style, and self-correction failure — none of which a static benchmark can see.

02

Refinement, not recovery

Gains concentrate where the opening hypothesis was already close. Bad first guesses are rarely recovered — and often reinforced.

03

The gap is behavioural, not photographic

Roughly 2× worse error in developing regions, with an image-quality audit ruling out the obvious data explanation.

04

More budget is not more intelligence

Unlimited actions cost up to 56× more per round and made two of three models less accurate.

05

Static skill does not transfer

The model purpose-built for geolocation was the only one to get worse when finally allowed to navigate.

Cite the paper

@inproceedings{mukherjee2026geoagent, title = {GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation}, author = {Mukherjee, Arka and Roy, Soham and Trivedi, Kartikeya and Ghosh, Shreya}, booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026}, year = {2026}, url = {https://anonymous.4open.science/r/geoagent-1162} }