GS-QA Leaderboard

Benchmarking geospatial reasoning.

Transparent evaluation across 3,300 questions and 53 vector, raster, and raster-vector task templates.

Published benchmark

Cite GS-QA2

Zhuocheng Shang, Shahd Elmahallawy, Zabir Al Nazi, Vagelis Hristidis, and Ahmed Eldawy. 2026. GS-QA2: A Benchmark for Question Answering over Raster–Vector Data. In Proceedings of the 34th ACM International Conference on Advances in Geographic Information Systems (SIGSPATIAL '26). Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3841645.3843431

@inproceedings{shang2026gsqa2,
  author    = {Zhuocheng Shang and Shahd Elmahallawy and Zabir Al Nazi and
               Vagelis Hristidis and Ahmed Eldawy},
  title     = {{GS-QA2}: A Benchmark for Question Answering over Raster--Vector Data},
  booktitle = {Proceedings of the 34th ACM International Conference on Advances
               in Geographic Information Systems},
  series    = {SIGSPATIAL '26},
  year      = {2026},
  month     = nov,
  publisher = {Association for Computing Machinery},
  address   = {New York, NY, USA},
  doi       = {10.1145/3841645.3843431},
  isbn      = {979-8-4007-2950-8}
}

Data and artifacts

Use the benchmark behind the rankings

Questions live in a versioned Hugging Face dataset. The leaderboard's paper-reported result rows are available here as JSON and CSV.

Benchmark dataset

GS-QA2 questions

3,300 questions: 2,800 vector-only and 500 raster-only or raster–vector. Pin the dataset commit used for every run.

Open dataset →

Evaluation tracks

V, R, and VR templates

V1–V28 are vector-only, R1–R11 are raster-only, and VR1–VR14 combine vector and raster reasoning.

Explore templates →

Leaderboard artifacts

Download result rows

Reuse the normalized, paper-reported scores that power the tables on this page.

Paper-reported evaluation

Rankings by evaluation group

Select a vector answer type or a raster map-algebra class. Each group is ranked only with its own scientifically appropriate metric.

Rank Model / method Verification Benchmark questions Metric Score

Inspect the benchmark

Template explorer

Compare vector-only (V1–V28), raster-only (R1–R11), and raster–vector (VR1–VR14) templates on the same questions.

Model / method Template ID Track and template Benchmark questions Primary metric Score

Evaluation request

Submit an evaluation

Complete one form for a self-service result or a maintainer-managed database, raster, or full evaluation.

Required information

Evaluation request

Required fields are marked with *. Do not include API keys or gated-model tokens.

What is evaluated? Every entry receives the same fixed GS-QA2 questions. The submitted model and selected method are the variables. For managed runs, maintainers pin the benchmark, parser, prompts, software environment, and hardware.
Self-service vector evaluation Run V1–V28 with the lightweight Docker runner, then provide the predictions dataset repository below.
Open runner guide →

What happens after you send it?

  1. Eligibility checkMaintainers confirm the model revision, track, method, and access requirements.
  2. EvaluationSelf-service predictions are validated, or maintainers run the model in the fixed reference environment.
  3. ScoringThe complete selected track is scored with the reference evaluator.
  4. PublicationThe discussion is updated and a verified result is added to the leaderboard.