360CityArena benchmark shows Gemini 2.5 Flash scoring 17.1% against 77.3% for humans

A new benchmark called 360CityArena tests how well embodied AI agents can explore and reason about real city environments, and the results show a wide gap between the best-performing model and humans. The benchmark is built from a photorealistic reconstruction of the Akihabara district in Tokyo, assembled from 602 360-degree video segments covering 85 streets. On top of this reconstruction sit 175 human-crafted tasks split into three categories: environment understanding, path reasoning, and spatial reasoning, together covering skills such as localization, landmark search, path planning, and reasoning about spatial relationships between locations. The stated motivation is that existing outdoor benchmarks either lack photorealism or lack complexity, leaving a considerable gap between what those benchmarks measure and how real urban environments actually behave. When the authors evaluated state-of-the-art large multimodal model (LMM) based agents on 360CityArena, the strongest one tested, Gemini 2.5 Flash, scored 17.1%, far below the 77.3% humans achieved on the same tasks. The authors present this gap as evidence of substantial unresolved challenges in city-scale embodied navigation and reasoning, and offer 360CityArena as a testbed for future work on photorealistic urban-district navigation and spatial reasoning.
Key facts
- 360CityArena is built from 602 360-degree video segments covering 85 streets of Tokyo's Akihabara district
- The benchmark includes 175 human-crafted tasks across three categories: environment understanding, path reasoning, and spatial reasoning
- Gemini 2.5 Flash, the strongest LMM-based agent tested, scored 17.1% on the benchmark
- Human performance on the same tasks was measured at 77.3%
- The authors say existing outdoor benchmarks lack either the photorealism or the complexity of real urban environments
Why it matters
Most embodied-agent benchmarks trade off realism for scale or vice versa: simplified simulators are easy to build but do not resemble real cities, while real-world testing is hard to standardize. 360CityArena tries to close that gap by reconstructing an actual, dense urban district (Akihabara) photorealistically from 360-degree video, then layering structured, human-crafted tasks on top. The resulting 60-point gap between human performance (77.3%) and the best tested agent (17.1%) is a concrete data point on how far current LMM-based agents are from operating reliably in real, unstructured city environments.
Who it affects
Researchers building or evaluating embodied navigation agents gain a new, more realistic testbed than prior outdoor benchmarks offered. Teams working on robotics, autonomous delivery, or any agent meant to operate in real urban spaces can use the results as a baseline for how far current models are from human-level urban competence. The source does not name the authors or their institution.
How to use it
360CityArena is presented as a benchmark and testbed rather than a product; the source gives no pricing, license, or access details, and does not name any models beyond Gemini 2.5 Flash as having been evaluated.
How solid is it
The benchmark's scale is stated precisely: 602 video segments across 85 streets and 175 tasks split into three defined categories, which gives it more structure than a loosely scraped test set. The headline comparison, 77.3% human versus 17.1% for the best model, is a specific, reproducible metric rather than a qualitative claim. The source does not describe how human performance was measured, such as the number of testers or the exact protocol, which limits how that baseline can be independently checked.
Risks and caveats
The source names only one evaluated model, Gemini 2.5 Flash, and calls it the strongest without listing what else was compared against it, so the breadth of the evaluation is unclear. No publication date, venue, or author affiliation is given in the available text, which limits the ability to assess provenance or peer review status. As with any single-district benchmark, performance on Akihabara may not generalize evenly to other urban layouts or cities.