# GeoGuessr RL: post-train an open-weights VLM with GRPO in one weekend

This document has two parts:

- **Part 1** is for you. It explains what you're building, why each choice was
  made, what it costs, and what to expect.
- **Part 2** is a specification for a coding agent. It contains the repo layout,
  exact configs, reward code, acceptance tests, and a runbook.

Research date: 2026-09-02. Prices and library versions drift, so re-check the
Runpod price page and pin library versions on install day.

---

# Part 1: Read this first

## Open questions for you

None of these block the plan. The plan proceeds on the assumptions in
parentheses. Answer them if you want a different default.

1. **Which post did you see?** The closest match I found is sdan's
   [GeoVLM / Vista3](https://sdan.io/projects/geovlm) work (Qwen3-VL-30B, 50k
   curated Street View pairs, SFT then RL with a log-distance reward, JAX gym).
   That project used far more than $50 of compute. This plan is the
   single-GPU, $30 version of the same idea. (Assumption: you want the
   approach, not a reproduction.)
2. **Do you have a Hugging Face account and a Weights & Biases account?** Both
   free tiers are enough. (Assumption: yes. W&B is used for live training
   curves. Swap to TensorBoard if you don't want it.)
3. **Training framework:** plain TRL with vLLM, or Unsloth? (Assumption: TRL
   with vLLM colocate. Unsloth is the fallback. Reasoning is in the
   "Framework" section.)
4. **Do you care about actual GeoGuessr gameplay (Google Street View
   panoramas) or about the geolocation skill in general?** (Assumption: the
   skill. Training data is Mapillary street imagery, and a small Google Street
   View panorama set is used only for evaluation. GeoGuessr's terms of service
   prohibit bots, so the plan doesn't drive the real game.)

## What you're building

A 4B-parameter vision-language model (VLM) that takes one street-level photo
and outputs a short chain of clues, a country, and a latitude and longitude.
You start from Qwen3-VL-4B-Instruct, which already gets the country right about
45% of the time on hard worldwide imagery, and you use Group Relative Policy
Optimization (GRPO) with a distance-based reward to make it better.

The reward is the actual GeoGuessr scoring function, plus a country bonus and a
format check. No labeled reasoning data is needed. The only labels are
coordinates and country codes, which the dataset already has.

## Expected outcome, honestly

"Extremely good at GeoGuessr" has a ceiling on a weekend budget. Here is the
landscape on the OSV-5M test set, where GeoScore is the GeoGuessr score out of
5000 averaged per image:

| Player | GeoScore on OSV-5M test | Source |
|---|---|---|
| Random guess | 328 | OSV-5M paper |
| Human annotators (OSV-5M study) | 1009 | OSV-5M paper |
| Average GeoGuessr player (world map stats) | ~2109 | Jerry Wei blog |
| Claude 3.5 Sonnet zero-shot | 2269 | Jerry Wei blog |
| OSV-5M best trained classifier baseline | 3361 | OSV-5M paper |
| Expert GeoGuessr player | ~4579 | Jerry Wei blog |

Qwen3-VL-4B zero-shot gets 45.5% top-1 country accuracy on a stratified
50k-image OSV-5M sample and 74.8% on a Google Street View screenshot set
(Where Do VLMs Fail, 2026). Its GeoScore is not published; the first thing
this project does is measure it.

A realistic weekend target:

- Country accuracy on OSV-5M test: **+8 to +15 points** over the base model.
- GeoScore: **+300 to +600** over the base model.
- Beat the "average GeoGuessr player" line. Beating Claude 3.5 Sonnet is a
  stretch goal. Beating experts is out of scope for a 4B model in one weekend.

Published RL-for-geolocation results support these expectations. Geo-R
(Qwen2.5-VL-7B, 8 A100s, 200k RL samples) moved IM2GPS3K accuracy at 25 km
from 31.7% to 41.5%. GeoAgent (Qwen2.5-VL-7B, 8 A40s) reached 76% country
accuracy. You have about 1/50th of that compute, so expect a fraction of the
gain, concentrated at the coarse (country and region) levels.

## Key design choices and why

### Model: Qwen3-VL-4B-Instruct

- Apache 2.0, 4B dense, supported by transformers, vLLM, TRL, and Unsloth.
- On geolocation the 4B **outperforms the 8B** sibling in the one paper that
  tested both (Where Do VLMs Fail, 2026), so the smaller model costs nothing.
- Fits bf16 weights, LoRA optimizer state, 8 rollouts per prompt, and a
  colocated vLLM engine on one 80 GB GPU.

Alternatives considered:

- **Qwen3.5-4B** is natively multimodal and newer, but as of this writing it
  has friction: vLLM registers only the multimodal class and breaks TRL's
  colocate mode for some checkpoints, Unsloth recommends against 4-bit
  training for it, and it needs transformers v5 with a thinking-mode chat
  template. Keep it as a stretch goal after the pipeline works.
- **Qwen3-VL-8B-Instruct** is a drop-in swap if you have budget left. Roughly
  1.8x the compute per step.
- **Gemma 3 4B** works in TRL and Unsloth but has a non-Apache license and
  weaker country recall in the benchmarks I found.

### Framework: TRL GRPOTrainer with vLLM colocate and LoRA

GRPO time is dominated by generation. Over 90% of wall-clock in a well-tuned
run is sampling rollouts, so a fast sampler is the single biggest lever.

- TRL's `GRPOTrainer` supports VLMs natively: the dataset carries an `image`
  column, prompts use the chat format with an image placeholder, and reward
  functions receive every extra dataset column (latitude, longitude, country)
  as keyword arguments. `--use_vllm --vllm_mode colocate` runs vLLM on the
  same GPU and syncs LoRA weights each generation step.
- Unsloth adds memory savings and a "standby" mode, and its Qwen3-VL Vision
  GRPO notebook is a good reference. Its downside is aggressive version
  pinning that breaks often. Use it as plan B if TRL OOMs or is too slow.
- Both paths freeze the vision encoder. vLLM can't serve LoRA on vision
  layers, and for geolocation the language side is where the reasoning lives.

### Data: OSV-5M, subsampled and country-capped

OpenStreetView-5M (CVPR 2024, CC BY-SA 4.0) is 5.1M Mapillary street images
with latitude, longitude, country, region, and city labels, and a test set
that is at least 1 km from any training image. It's 259 GB in full, but it
ships as 98 training zips of about 50k images each, so you download
**two training zips and one test zip** (roughly 8 GB) and never touch the
rest.

Two lessons from the literature shape the sampling:

- **Cap images per country.** Mapillary coverage is heavily skewed to the US,
  Europe, and Japan. Uniform sampling teaches the model to guess Ohio.
  GeoAgent found that hierarchical, population-weighted sampling reduced
  bias significantly. This plan caps each country at 400 training images.
- **Filter for learnable prompts.** GRPO gets zero gradient when all 8
  rollouts for an image score the same (all right or all hopeless). Geo-R
  calls this the "vanishing advantages" problem. Before the main run, the
  base model samples each candidate image 4 times, and only images with mixed
  outcomes go into the training set, plus a 20% random slice to avoid
  narrowing the distribution.

For a second, GeoGuessr-flavored evaluation set, the plan uses
`stochastic/random_streetview_images_pano_v0.0.2` (11k Google Street View
panoramas with coordinates, MIT license). It's out of distribution for the
training data, which is exactly why it's a good check.

### Reward: GeoGuessr score, plus shaping

The GeoGuessr world-map score is `5000 * exp(-d / 1492.7)` with `d` in
kilometers. Normalized to `[0, 1]` it's a fine reward, but it barely
distinguishes 5 km from 100 km, so the plan adds a sharper term and a country
bonus:

```
r_format  = 1 if the answer parses to a valid country string and lat/lon in range, else 0
r_country = 1 if predicted country matches the label, else 0
r_dist    = 0.5 * exp(-d / 1492.7) + 0.5 * exp(-d / 250)
reward    = 0.10 * r_format + 0.30 * r_country + 0.60 * r_dist
```

If the answer doesn't parse, `r_country` and `r_dist` are 0, so the model
learns the format in the first few dozen steps. The 250 km term is the
"region" scale; GeoAgent used `exp(-d/200)` and Geo-R used piecewise-linear
bands at 750 and 2500 km, so this is well within the range that works.

The raw GeoScore is logged as a metric but not used as the sole reward.

### Prompt: short reasoning, strict answer block

The model is asked for at most 150 words of clue analysis inside `<think>`
tags, then a JSON answer inside `<answer>` tags. Completions are capped at
512 tokens. Short completions keep generation cheap, and forcing the model to
name clues (driving side, script on signs, bollards, vegetation, road
markings) is where transfer comes from.

## Compute and budget

Recommended GPU: **1x A100 80GB PCIe** on Runpod. Community Cloud lists it at
about $1.19/hr and Secure Cloud at $1.39/hr. An H100 PCIe at $1.99/hr is
roughly 1.7x faster and costs about the same per unit of work, so pick
whichever is available.

Throughput estimate for the recommended config (4 prompts x 8 rollouts x
~300 generated tokens per optimizer step) is 15 to 25 seconds per step, or
150 to 250 steps per hour. **This is an estimate, not a measurement.** The
runbook has a 20-step smoke test whose job is to replace it with a real
number before you commit budget.

| Phase | GPU hours | Cost at $1.39/hr |
|---|---|---|
| Pod setup, installs, data download | 1.0 | $1.40 |
| Baseline eval (1000 OSV-5M + 500 pano images) | 0.5 | $0.70 |
| Learnability filter (8k images x 4 samples) | 0.75 | $1.05 |
| Smoke test (20 steps) and throughput measurement | 0.25 | $0.35 |
| Main GRPO run | 8.0 | $11.10 |
| Checkpoint evals and final eval | 1.0 | $1.40 |
| Merge LoRA, export, demo | 0.5 | $0.70 |
| Network volume, 100 GB, prorated for one week | | ~$1.60 |
| **Subtotal** | **12.0** | **~$18** |
| Contingency for crashes and re-runs (60%) | | ~$11 |
| **Total** | | **~$29** |

You stay under $50 with room to spare. Ways the budget can exceed $50:

- Swapping to Qwen3-VL-8B and keeping the same step count: roughly +$12.
- Doubling the main run to 16 hours: +$11.
- Adding an SFT warm-up with frontier-model-generated reasoning traces
  (Geo-R1 and GeoAgent both do this): API cost of about $10 to $20 for 3k
  traces, plus an hour of GPU. Worth it if the first RL run plateaus early.

Do all coding, reward unit tests, and data-prep dry runs on your Mac. Only
pay for GPU time when the pipeline is already green on CPU with a tiny
sample.

## Weekend schedule

**Friday evening (no GPU):** Hand Part 2 to the coding agent. It builds the
repo, writes the reward and parser with unit tests, and dry-runs the data
pipeline on the 116 MB OSV-5M `test.csv` plus one test zip. You review the
prompt and the reward weights.

**Saturday morning:** Create a Runpod network volume and pod. Install, download
two training zips, run the baseline eval. Write the baseline numbers down; they
are the whole point of the project. Run the smoke test and read the real
seconds-per-step figure. Run the learnability filter. Launch the main run by
noon with `max_steps` sized to fit your remaining budget.

**Saturday evening:** Check W&B. Reward should rise and format compliance
should be near 100% within the first 50 steps. Evaluate the latest checkpoint
on the 300-image quick set. If country accuracy hasn't moved after 300 steps,
stop and debug rather than burn budget.

**Sunday:** Resume or continue the run if it's still improving. Run the full
eval. Merge the LoRA, push the adapter to the Hub, and try the Gradio demo on
your own photos. Write up before and after numbers.

## Risks and how the plan handles them

- **Version drift.** TRL is at v1.11 and needs transformers 5.8+ and vLLM
  0.22+. The runbook installs `trl[vllm]`, records the resolved versions to a
  lock file, and runs the smoke test before anything else.
- **Reward hacking by always guessing a hotspot.** Country capping plus the
  country bonus make a fixed guess lose on most prompts. The eval logs the
  distribution of predicted countries so you can spot collapse.
- **Degenerate text.** Unsloth reports Qwen VL models emitting repeated junk
  tokens early in GRPO. The format reward is 0 for unparseable output, and
  `max_completion_length` caps the damage.
- **OOM.** Drop `vllm_gpu_memory_utilization` from 0.35 to 0.25, then reduce
  `num_generations` from 8 to 6, then switch to Unsloth.
- **Too slow.** If the smoke test shows over 40 s/step, cut
  `max_completion_length` to 384, reduce image `max_pixels`, or move to an
  H100.

## Stretch goals, in order of value per hour

1. Qwen3-VL-8B-Instruct with the same pipeline.
2. SFT warm-up on 2k to 3k reasoning traces before RL.
3. Multi-view prompts: 2 to 4 images from the same Mapillary sequence
   (OSV-5M has a `sequence` column) to mimic GeoGuessr's ability to look
   around.
4. Test-time zoom tool, as in the GeoVLM agent variant.
5. Qwen3.5-4B once the vLLM and TRL friction is resolved.

---

# Part 2: Specification for the coding agent

You are implementing a single-GPU GRPO post-training pipeline that trains
Qwen3-VL-4B-Instruct to geolocate street-level photos. Everything must run
end to end with one command per phase. Build and test on CPU with tiny
samples first; the GPU steps are marked.

## Constraints

- Language: Python 3.11. Package manager: `uv`.
- Frameworks: `trl[vllm]`, `transformers`, `peft`, `datasets`, `vllm`,
  `torch`. Install `trl[vllm]` and let it resolve compatible `transformers`
  and `vllm` versions. Write the resolved versions to `requirements.lock`.
  Don't hand-pin versions you haven't verified install together.
- One GPU, 80 GB. No multi-GPU code paths.
- Every script takes `--config PATH` pointing to a YAML file and accepts
  `--limit N` to run on a tiny subset for testing.
- Log to Weights & Biases when `WANDB_API_KEY` is set; otherwise log to
  TensorBoard under `runs/`.
- Deterministic sampling everywhere: fixed seeds, sorted inputs.
- Write docs and comments in Google developer documentation style: second
  person, present tense, active voice, sentence-case headings.

## Repository layout

```
geoguessr-rl/
  README.md                  # how to run each phase, with the runbook commands
  pyproject.toml
  requirements.lock
  configs/
    data.yaml
    eval.yaml
    grpo.yaml
    filter.yaml
  geoguessr_rl/
    __init__.py
    prompt.py                # system prompt, user prompt, answer schema
    parsing.py               # extract <answer> JSON, validate, normalize country
    reward.py                # haversine, GeoScore, reward components, TRL reward fns
    countries.py             # ISO-2 <-> name normalization via pycountry
    data.py                  # OSV-5M download, CSV join, country capping, HF Dataset builders
    eval_metrics.py          # GeoScore, acc@km thresholds, country acc, prediction histogram
  scripts/
    prepare_data.py          # phase 1
    run_eval.py              # phase 2 and 5 (vLLM batch inference + metrics)
    filter_learnable.py      # phase 3
    train_grpo.py            # phase 4
    merge_and_export.py      # phase 6
    demo_gradio.py           # phase 6
  tests/
    test_parsing.py
    test_reward.py
    test_countries.py
    test_data_small.py
  runbook/
    runpod_setup.sh          # apt/uv install, env vars, volume paths
    smoke_test.sh            # 20-step GRPO run + throughput printout
```

## Prompt

Define these in `prompt.py`:

```python
SYSTEM_PROMPT = (
    "You are an expert GeoGuessr player. You identify where a street-level "
    "photo was taken from visual clues alone."
)

USER_PROMPT = (
    "Where was this photo taken?\n"
    "First, inside <think></think> tags, list the strongest clues in at most "
    "150 words: driving side, language and script on signs, license plates, "
    "road markings and bollards, utility poles, vegetation and climate, "
    "architecture, soil color, and any place names.\n"
    "Then give your final answer inside <answer></answer> tags as JSON with "
    "exactly these keys: \"country\" (English country name), \"lat\" "
    "(decimal degrees), \"lon\" (decimal degrees).\n"
    "Example:\n"
    "<answer>{\"country\": \"Kenya\", \"lat\": -1.2921, \"lon\": 36.8219}</answer>"
)
```

The conversational prompt for TRL is:

```python
[
  {"role": "system", "content": SYSTEM_PROMPT},
  {"role": "user", "content": [{"type": "image"}, {"type": "text", "text": USER_PROMPT}]},
]
```

The `image` column holds the PIL image. TRL's processor inserts the image
tokens at the placeholder.

## Parsing

`parsing.py` exposes `parse_answer(text: str) -> ParsedAnswer | None`.

- Take the **last** `<answer>...</answer>` block. Ignore `<think>` content.
- Parse JSON leniently: strip code fences, allow trailing commas, allow
  single quotes by attempting `json.loads` then `ast.literal_eval`.
- Validate: `lat` in `[-90, 90]`, `lon` in `[-180, 180]`, both finite,
  `country` a non-empty string.
- Normalize country to ISO 3166-1 alpha-2 with `countries.to_iso2(name)`,
  which uses `pycountry` exact lookup, then `pycountry` fuzzy search, then a
  small alias table (for example "USA", "UK", "Russia", "South Korea",
  "Czechia", "Ivory Coast", "Burma", "Holland", "Macedonia"). Return `None`
  for the ISO code if nothing matches, but still return the coordinates.
- Return `None` for the whole answer only when there is no parseable
  `<answer>` block or the coordinates are invalid.

Tests must cover: valid answer, missing tags, two answer blocks, out-of-range
coordinates, string-typed numbers, code-fenced JSON, unknown country name,
and completions that are only repeated garbage tokens.

## Reward

`reward.py` implements the following and exposes TRL-compatible functions.

```python
EARTH_RADIUS_KM = 6371.0088
GEOGUESSR_SCALE_KM = 1492.7
FINE_SCALE_KM = 250.0

def haversine_km(lat1, lon1, lat2, lon2) -> float: ...

def geoscore(d_km: float) -> float:
    """GeoGuessr world-map score in [0, 5000]."""
    return 5000.0 * math.exp(-d_km / GEOGUESSR_SCALE_KM)

def distance_reward(d_km: float) -> float:
    return 0.5 * math.exp(-d_km / GEOGUESSR_SCALE_KM) + 0.5 * math.exp(-d_km / FINE_SCALE_KM)

def components(completion_text, latitude, longitude, country_iso2) -> dict:
    """Returns format, country, dist, total, d_km, geoscore for one completion."""
```

Weights: `W_FORMAT = 0.10`, `W_COUNTRY = 0.30`, `W_DIST = 0.60`. If parsing
fails, `country = dist = 0` and `d_km = None`. If the predicted ISO code is
`None` but coordinates parse, `format = 1`, `country = 0`, and `dist` is
computed normally.

TRL reward functions receive `prompts`, `completions`, and every dataset
column as a keyword argument. Dataset columns must be named `latitude`,
`longitude`, and `country_iso2`. Implement one combined reward function that
returns the total, and three logging-only functions that return the
components with `reward_weights` set to zero for them in the config, so the
components show up as separate curves in W&B. Completions arrive as a list of
message dicts in conversational mode; take `completion[0]["content"]`.

Tests must cover: `haversine_km` against three known city pairs within 0.5%,
`geoscore(0) == 5000`, `geoscore(1492.7) ≈ 1839.4`, monotonic decrease,
reward for a perfect answer equals 1.0, reward for an unparseable answer
equals 0.0, and reward for a correct-country wrong-coordinate answer lands
between 0.40 and 0.70.

## Phase 1: Prepare data (CPU, then GPU pod)

`scripts/prepare_data.py --config configs/data.yaml`

Downloads from `osv5m/osv5m` on the Hub with `hf_hub_download`:

- `test.csv` (116 MB) and `images/test/00.zip`.
- `train.csv` (2.9 GB) and `images/train/00.zip`, `images/train/01.zip`.

Read CSVs with `polars`, selecting only `id`, `latitude`, `longitude`,
`country`, `region`, `sub-region`, `city`, `sequence`. Inspect the header at
runtime and fail loudly if a column is missing; the `country` column is
expected to be ISO-2 but verify with an assertion that all values are
two uppercase letters, and log a sample if not. Join to the zip contents by
filename stem equals `id`.

Build three parquet files under `data/`, each with columns
`image_path, latitude, longitude, country_iso2, region, city, sequence`:

- `train_pool.parquet`: from the two train zips, cap at 400 images per
  country using a seeded shuffle. Expect 15k to 30k rows.
- `eval_osv.parquet`: 1000 rows from the test zip. Stratify: up to 10 per
  country first, then fill randomly. Fixed seed.
- `eval_quick.parquet`: the first 300 rows of `eval_osv.parquet`.

Also build `eval_pano.parquet` from
`stochastic/random_streetview_images_pano_v0.0.2`: 500 rows, seeded, with
images resized so the long edge is 1024 px and saved as JPEG under
`data/pano/`. Map its country column to ISO-2 with `countries.to_iso2`.

Image preprocessing for all sets: load, convert to RGB, and if either
dimension exceeds 1024 px, resize so the long edge is 1024 px. OSV-5M images
are 512 px tall and about 800 px wide, so most pass through unchanged. Store
the processed images as files and load lazily in the Dataset to keep memory
low.

Data sanity checks the script must print: row counts, number of countries,
top-10 countries by count, and a histogram of image widths.

`--limit 200` must run on a laptop against only `test.csv` and one test zip,
skipping the train download, for the CPU dry run.

## Phase 2: Baseline eval (GPU)

`scripts/run_eval.py --config configs/eval.yaml --model Qwen/Qwen3-VL-4B-Instruct --split eval_osv`

Use vLLM offline inference directly (`vllm.LLM` with
`limit_mm_per_prompt={"image": 1}`), `max_model_len=4096`, greedy decoding
(`temperature=0`), `max_tokens=512`. Batch all prompts in one `generate`
call. Accept `--lora PATH` to evaluate an adapter with `enable_lora=True`.

Write `outputs/eval/<run_name>/<split>/predictions.jsonl` with one line per
image containing the raw completion, parsed fields, `d_km`, `geoscore`, and
per-component rewards. Write `metrics.json` with:

- `format_rate`, `country_acc`, `mean_geoscore`, `median_km`,
- `acc_at_km` for thresholds 1, 25, 200, 750, 2500,
- `pred_country_top10` (a histogram; use it to detect collapse),
- `n`.

Print a one-line summary. Run on both `eval_osv` and `eval_pano`.

Use `max_pixels = 1024 * 28 * 28` in the processor kwargs, which caps an image
at roughly 1000 visual tokens. Don't lower it below `512 * 28 * 28`.

## Phase 3: Filter for learnable prompts (GPU)

`scripts/filter_learnable.py --config configs/filter.yaml`

- Sample 8000 rows from `train_pool.parquet` (seeded, country-capped at 200).
- With vLLM, generate `n=4` completions per image at `temperature=1.0`,
  `top_p=0.95`, `max_tokens=512`.
- Compute the total reward for each completion. For each image record
  `mean_reward`, `std_reward`, and `country_hits` (0 to 4).
- Keep an image if `0 < country_hits < 4` **or** `std_reward > 0.08`. Call
  this set A.
- Add a seeded random 20% of the rejected images. Call this set B.
- Write `train_learnable.parquet` = A ∪ B, shuffled with a seed. Print the
  sizes of A, B, and the country distribution of the result. Target 3k to 5k
  rows. If A is under 1500 rows, lower the `std_reward` threshold to 0.05 and
  rerun without regenerating (cache the completions to disk).

Also write `filter_stats.json` with the base model's mean reward and country
accuracy on the 8000 sampled images, as a second baseline.

## Phase 4: GRPO training (GPU)

`scripts/train_grpo.py --config configs/grpo.yaml`

Build the dataset from `train_learnable.parquet` with columns `prompt`
(conversational list above), `image` (PIL, loaded lazily via
`datasets.Image()` feature), `latitude`, `longitude`, `country_iso2`.

Model and LoRA:

```python
model_name = "Qwen/Qwen3-VL-4B-Instruct"
dtype = bfloat16
attn_implementation = "flash_attention_2"   # fall back to sdpa if unavailable
peft_config = LoraConfig(
    r=32, lora_alpha=64, lora_dropout=0.0, bias="none",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
    # Restrict to the language model. Exclude the vision tower and the
    # merger/projector so vLLM can serve the adapter.
    exclude_modules=r".*visual.*",
    task_type="CAUSAL_LM",
)
```

Verify after wrapping that no LoRA module name contains `visual`. Fail if it
does.

`GRPOConfig` values for `configs/grpo.yaml`:

```yaml
output_dir: outputs/grpo-qwen3vl4b
run_name: grpo-qwen3vl4b-v1
seed: 42
learning_rate: 1.0e-5
lr_scheduler_type: constant_with_warmup
warmup_steps: 20
max_grad_norm: 0.2
bf16: true
gradient_checkpointing: true
per_device_train_batch_size: 8        # completions per micro-batch
gradient_accumulation_steps: 4        # 32 completions = 4 prompts per optimizer step
num_generations: 8
generation_batch_size: 64             # 8 prompts sampled per vLLM call
max_prompt_length: null               # required for VLMs; don't truncate image tokens
max_completion_length: 512
temperature: 1.0
top_p: 1.0
loss_type: dapo
beta: 0.0
scale_rewards: group
mask_truncated_completions: true
use_vllm: true
vllm_mode: colocate
vllm_gpu_memory_utilization: 0.35
vllm_max_model_len: 4096
reward_weights: [1.0, 0.0, 0.0, 0.0]  # total, then logging-only components
max_steps: 1200                       # resize after the smoke test
save_steps: 200
save_total_limit: 3
logging_steps: 1
log_completions: true
num_completions_to_print: 2
report_to: wandb
```

Processor kwargs: `max_pixels = 1024 * 28 * 28`, `min_pixels = 256 * 28 * 28`.

Add a `--resume_from_checkpoint auto` option that picks the latest checkpoint
in `output_dir`. Add a callback that every `save_steps` runs
`run_eval.py --split eval_quick --lora <checkpoint>` in a subprocess
after training releases the GPU, or, if that's too disruptive, write a
separate `scripts/eval_checkpoints.sh` that loops over saved checkpoints.
Prefer the separate script; it's simpler and the run doesn't stall.

Budget arithmetic the script must print at start: given
`seconds_per_step` from `--seconds_per_step` (measured in the smoke test) and
`--budget_hours`, print the recommended `max_steps` and exit if
`--dry_run`.

## Phase 5: Final eval (GPU)

Run `run_eval.py` with `--lora outputs/grpo-qwen3vl4b/checkpoint-<best>` on
`eval_osv` and `eval_pano`. Produce `outputs/report.md` with a before and
after table for every metric, plus five example completions where the
adapter beat the base model by the largest GeoScore margin and five where it
did worst.

## Phase 6: Export and demo (GPU or CPU)

`scripts/merge_and_export.py`: merge the LoRA into the base model with
`merge_and_unload`, save to `outputs/merged/`, and optionally push both the
adapter and the merged model to the Hub under a user-provided repo ID.

`scripts/demo_gradio.py`: a single-image Gradio app that shows the `<think>`
text, a map marker via a static map image or a folium HTML embed, and the
predicted country. Load the merged model with transformers, not vLLM, so it
runs on a Mac with `device_map="auto"` for smoke testing.

## Runbook

`runbook/runpod_setup.sh` must do the following on a fresh Runpod PyTorch
pod (CUDA 12.x image) with a network volume mounted at `/workspace`:

1. Install `uv`, then `uv venv /workspace/.venv` and activate.
2. `uv pip install "trl[vllm]" peft datasets polars pycountry pillow wandb
   gradio flash-attn --no-build-isolation` and freeze to
   `requirements.lock`. If `flash-attn` fails to build, continue without it
   and set `attn_implementation=sdpa`.
3. Export `HF_HOME=/workspace/hf`, `HF_HUB_ENABLE_HF_TRANSFER=1`, and
   `VLLM_WORKER_MULTIPROC_METHOD=spawn`.
4. Run `python -c "import trl, transformers, vllm; print(...)"` and print
   versions.

`runbook/smoke_test.sh` runs `train_grpo.py` with `max_steps=20`,
`save_steps=1000`, `report_to=none`, and `--limit 256`, then prints the
average seconds per step over steps 5 to 20 and the recommended `max_steps`
for an 8-hour budget. It must exit non-zero if any step produced
`format_rate < 0.5` after step 10, since that indicates a broken prompt or
parser rather than a training problem.

Order of operations on the pod:

```bash
bash runbook/runpod_setup.sh
python scripts/prepare_data.py --config configs/data.yaml
python scripts/run_eval.py --config configs/eval.yaml --split eval_osv --run_name base
python scripts/run_eval.py --config configs/eval.yaml --split eval_pano --run_name base
bash runbook/smoke_test.sh
python scripts/filter_learnable.py --config configs/filter.yaml
python scripts/train_grpo.py --config configs/grpo.yaml --seconds_per_step <measured> --budget_hours 8
bash scripts/eval_checkpoints.sh
python scripts/run_eval.py --config configs/eval.yaml --split eval_osv --run_name rl --lora outputs/grpo-qwen3vl4b/checkpoint-<best>
python scripts/run_eval.py --config configs/eval.yaml --split eval_pano --run_name rl --lora outputs/grpo-qwen3vl4b/checkpoint-<best>
python scripts/merge_and_export.py --lora outputs/grpo-qwen3vl4b/checkpoint-<best>
```

## Acceptance tests

Before any GPU is rented, all of the following pass on a laptop:

1. `pytest tests/` passes.
2. `python scripts/prepare_data.py --config configs/data.yaml --limit 200`
   produces `eval_osv.parquet` with 200 rows and prints the sanity checks.
3. `python scripts/run_eval.py --config configs/eval.yaml --split eval_osv
   --limit 4 --backend transformers --device cpu` runs the base model on 4
   images in bf16 or fp32 on CPU and writes a valid `metrics.json`. This
   confirms the prompt, chat template, parser, and metrics code end to end.
   It's slow; that's fine.
4. `python scripts/train_grpo.py --config configs/grpo.yaml --dry_run
   --seconds_per_step 20 --budget_hours 8` prints the dataset schema, the
   LoRA target module list with zero `visual` matches, and
   `recommended max_steps = 1440`.

On the GPU, the smoke test passes and reports seconds per step.

## Failure handling

- **OOM during vLLM init:** lower `vllm_gpu_memory_utilization` to 0.25.
- **OOM during backward:** set `per_device_train_batch_size: 4` and
  `gradient_accumulation_steps: 8`. Same effective batch.
- **vLLM rejects the LoRA adapter (vision modules):** the
  `exclude_modules` regex is wrong for this transformers version. Print
  `model.peft_config` target names and fix the regex.
- **`format_rate` stuck near 0 after 30 steps:** print 10 raw completions.
  Usually the chat template added a thinking block or the model ignores the
  `<answer>` tag. Add a one-shot example to the user prompt.
- **Reward rises but `country_acc` on `eval_quick` doesn't:** check
  `pred_country_top10`. If one country dominates, the model found a hotspot.
  Lower the per-country cap to 250 and raise `W_COUNTRY` to 0.40.
- **Slower than 40 s/step:** reduce `max_completion_length` to 384 and
  `max_pixels` to `768 * 28 * 28`.

## Plan B: Unsloth

If TRL plus vLLM colocate can't be made to fit or run in the smoke test
within one hour of debugging, switch to Unsloth. Follow the structure of the
Unsloth "Qwen3-VL (8B) Vision GRPO" notebook: `FastVisionModel.from_pretrained`
with `load_in_4bit=False`, `fast_inference=True`,
`gpu_memory_utilization=0.6`, `max_seq_length=4096`;
`get_peft_model(finetune_vision_layers=False, finetune_language_layers=True,
r=32, lora_alpha=64)`; then the same `GRPOConfig` values and the same reward
functions. Keep the data, eval, and export scripts unchanged.

---

## Sources

- [Unsloth: Vision reinforcement learning (VLM RL)](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/vision-reinforcement-learning-vlm-rl)
- [Unsloth Qwen3-VL (8B) Vision GRPO notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_VL_(8B)-Vision-GRPO.ipynb)
- [Unsloth Qwen3.5 fine-tuning guide](https://unsloth.ai/docs/models/qwen3.5/fine-tune)
- [Unsloth: Memory efficient RL](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide/memory-efficient-rl)
- [TRL GRPO trainer documentation](https://huggingface.co/docs/trl/main/grpo_trainer)
- [TRL blog: Vision language model alignment](https://huggingface.co/blog/trl-vlm-alignment)
- [TRL releases](https://github.com/huggingface/trl/releases)
- [OSV-5M dataset on Hugging Face](https://huggingface.co/datasets/osv5m/osv5m)
- [OpenStreetView-5M paper (CVPR 2024)](https://arxiv.org/abs/2404.18873)
- [random_streetview_images_pano_v0.0.2 dataset](https://huggingface.co/datasets/stochastic/random_streetview_images_pano_v0.0.2)
- [Qwen3-VL-4B-Instruct model card](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct)
- [Qwen3.5-4B model card](https://huggingface.co/Qwen/Qwen3.5-4B)
- [vLLM issue: Qwen3.5-4B incompatibility](https://github.com/vllm-project/vllm/issues/36275)
- [Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization](https://arxiv.org/html/2604.16248)
- [Geo-R: Vision-Language Reasoning for Geolocalization, an RL approach](https://arxiv.org/html/2601.00388)
- [GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics](https://arxiv.org/html/2602.12617)
- [Geo-R1: Unlocking VLM Geospatial Reasoning with Cross-View RL](https://arxiv.org/html/2510.00072)
- [GeoVLM / Vista3 project page (sdan)](https://sdan.io/projects/geovlm)
- [vlm-gym (sdan)](https://github.com/sdan/vlm-gym)
- [Claude plays GeoGuessr (Jerry Wei)](https://www.jerrywei.net/blog/claude-plays-geoguessr)
- [The maths of GeoGuessr](https://latb.io/geoguessr/articles/the-maths)
- [Epoch AI GeoBench](https://epoch.ai/benchmarks/geobench)
- [Runpod pricing](https://www.runpod.io/pricing)
