research.kartikeya.me

Papers, explained properly.

Interactive walkthroughs of research worth understanding — the argument, the method, and the numbers you can actually poke at.


001 · ANONYMOUS ACL SUBMISSION, UNDER REVIEW

GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation

Authors redacted for double-blind review

What changes when you stop showing a vision-language model a photograph and let it walk around Street View instead. Six models, a developed/developing gap that image quality does not explain, and an action budget where more turned out to be worse.

1,200locations
100cities
6models
5 kmbest median error
56×cost of no limit
Vision-language modelsAgentsBenchmark
Read the walkthrough →