← projects

GeoGuessr RL

Post-train an open-weights VLM with GRPO in one weekend

This document has two parts:

  • Part 1 is for you. It explains what you're building, why each choice was made, what it costs, and what to expect.
  • Part 2 is a specification for a coding agent. It contains the repo layout, exact configs, reward code, acceptance tests, and a runbook.

Research date: 2026-09-02. Prices and library versions drift, so re-check the Runpod price page and pin library versions on install day.


Part 1: Read this first

Open questions for you

None of these block the plan. The plan proceeds on the assumptions in parentheses. Answer them if you want a different default.

  1. Which post did you see? The closest match I found is sdan's GeoVLM / Vista3 work (Qwen3-VL-30B, 50k curated Street View pairs, SFT then RL with a log-distance reward, JAX gym). That project used far more than $50 of compute. This plan is the single-GPU, $30 version of the same idea. (Assumption: you want the approach, not a reproduction.)
  2. Do you have a Hugging Face account and a Weights & Biases account? Both free tiers are enough. (Assumption: yes. W&B is used for live training curves. Swap to TensorBoard if you don't want it.)
  3. Training framework: plain TRL with vLLM, or Unsloth? (Assumption: TRL with vLLM colocate. Unsloth is the fallback. Reasoning is in the "Framework" section.)
  4. Do you care about actual GeoGuessr gameplay (Google Street View panoramas) or about the geolocation skill in general? (Assumption: the skill. Training data is Mapillary street imagery, and a small Google Street View panorama set is used only for evaluation. GeoGuessr's terms of service prohibit bots, so the plan doesn't drive the real game.)

What you're building

A 4B-parameter vision-language model (VLM) that takes one street-level photo and outputs a short chain of clues, a country, and a latitude and longitude. You start from Qwen3-VL-4B-Instruct, which already gets the country right about 45% of the time on hard worldwide imagery, and you use Group Relative Policy Optimization (GRPO) with a distance-based reward to make it better.

The reward is the actual GeoGuessr scoring function, plus a country bonus and a format check. No labeled reasoning data is needed. The only labels are coordinates and country codes, which the dataset already has.

Expected outcome, honestly

"Extremely good at GeoGuessr" has a ceiling on a weekend budget. Here is the landscape on the OSV-5M test set, where GeoScore is the GeoGuessr score out of 5000 averaged per image:

Player GeoScore on OSV-5M test Source
Random guess 328 OSV-5M paper
Human annotators (OSV-5M study) 1009 OSV-5M paper
Average GeoGuessr player (world map stats) ~2109 Jerry Wei blog
Claude 3.5 Sonnet zero-shot 2269 Jerry Wei blog
OSV-5M best trained classifier baseline 3361 OSV-5M paper
Expert GeoGuessr player ~4579 Jerry Wei blog

Qwen3-VL-4B zero-shot gets 45.5% top-1 country accuracy on a stratified 50k-image OSV-5M sample and 74.8% on a Google Street View screenshot set (Where Do VLMs Fail, 2026). Its GeoScore is not published; the first thing this project does is measure it.

A realistic weekend target:

  • Country accuracy on OSV-5M test: +8 to +15 points over the base model.
  • GeoScore: +300 to +600 over the base model.
  • Beat the "average GeoGuessr player" line. Beating Claude 3.5 Sonnet is a stretch goal. Beating experts is out of scope for a 4B model in one weekend.

Published RL-for-geolocation results support these expectations. Geo-R (Qwen2.5-VL-7B, 8 A100s, 200k RL samples) moved IM2GPS3K accuracy at 25 km from 31.7% to 41.5%. GeoAgent (Qwen2.5-VL-7B, 8 A40s) reached 76% country accuracy. You have about 1/50th of that compute, so expect a fraction of the gain, concentrated at the coarse (country and region) levels.

Key design choices and why

Model: Qwen3-VL-4B-Instruct

  • Apache 2.0, 4B dense, supported by transformers, vLLM, TRL, and Unsloth.
  • On geolocation the 4B outperforms the 8B sibling in the one paper that tested both (Where Do VLMs Fail, 2026), so the smaller model costs nothing.
  • Fits bf16 weights, LoRA optimizer state, 8 rollouts per prompt, and a colocated vLLM engine on one 80 GB GPU.

Alternatives considered:

  • Qwen3.5-4B is natively multimodal and newer, but as of this writing it has friction: vLLM registers only the multimodal class and breaks TRL's colocate mode for some checkpoints, Unsloth recommends against 4-bit training for it, and it needs transformers v5 with a thinking-mode chat template. Keep it as a stretch goal after the pipeline works.
  • Qwen3-VL-8B-Instruct is a drop-in swap if you have budget left. Roughly 1.8x the compute per step.
  • Gemma 3 4B works in TRL and Unsloth but has a non-Apache license and weaker country recall in the benchmarks I found.

Framework: TRL GRPOTrainer with vLLM colocate and LoRA

GRPO time is dominated by generation. Over 90% of wall-clock in a well-tuned run is sampling rollouts, so a fast sampler is the single biggest lever.

  • TRL's GRPOTrainer supports VLMs natively: the dataset carries an image column, prompts use the chat format with an image placeholder, and reward functions receive every extra dataset column (latitude, longitude, country) as keyword arguments. --use_vllm --vllm_mode colocate runs vLLM on the same GPU and syncs LoRA weights each generation step.
  • Unsloth adds memory savings and a "standby" mode, and its Qwen3-VL Vision GRPO notebook is a good reference. Its downside is aggressive version pinning that breaks often. Use it as plan B if TRL OOMs or is too slow.
  • Both paths freeze the vision encoder. vLLM can't serve LoRA on vision layers, and for geolocation the language side is where the reasoning lives.

Data: OSV-5M, subsampled and country-capped

OpenStreetView-5M (CVPR 2024, CC BY-SA 4.0) is 5.1M Mapillary street images with latitude, longitude, country, region, and city labels, and a test set that is at least 1 km from any training image. It's 259 GB in full, but it ships as 98 training zips of about 50k images each, so you download two training zips and one test zip (roughly 8 GB) and never touch the rest.

Two lessons from the literature shape the sampling:

  • Cap images per country. Mapillary coverage is heavily skewed to the US, Europe, and Japan. Uniform sampling teaches the model to guess Ohio. GeoAgent found that hierarchical, population-weighted sampling reduced bias significantly. This plan caps each country at 400 training images.
  • Filter for learnable prompts. GRPO gets zero gradient when all 8 rollouts for an image score the same (all right or all hopeless). Geo-R calls this the "vanishing advantages" problem. Before the main run, the base model samples each candidate image 4 times, and only images with mixed outcomes go into the training set, plus a 20% random slice to avoid narrowing the distribution.

For a second, GeoGuessr-flavored evaluation set, the plan uses stochastic/random_streetview_images_pano_v0.0.2 (11k Google Street View panoramas with coordinates, MIT license). It's out of distribution for the training data, which is exactly why it's a good check.

Reward: GeoGuessr score, plus shaping

The GeoGuessr world-map score is 5000 * exp(-d / 1492.7) with d in kilometers. Normalized to [0, 1] it's a fine reward, but it barely distinguishes 5 km from 100 km, so the plan adds a sharper term and a country bonus:

r_format  = 1 if the answer parses to a valid country string and lat/lon in range, else 0
r_country = 1 if predicted country matches the label, else 0
r_dist    = 0.5 * exp(-d / 1492.7) + 0.5 * exp(-d / 250)
reward    = 0.10 * r_format + 0.30 * r_country + 0.60 * r_dist

If the answer doesn't parse, r_country and r_dist are 0, so the model learns the format in the first few dozen steps. The 250 km term is the "region" scale; GeoAgent used exp(-d/200) and Geo-R used piecewise-linear bands at 750 and 2500 km, so this is well within the range that works.

The raw GeoScore is logged as a metric but not used as the sole reward.

Prompt: short reasoning, strict answer block

The model is asked for at most 150 words of clue analysis inside <think> tags, then a JSON answer inside <answer> tags. Completions are capped at 512 tokens. Short completions keep generation cheap, and forcing the model to name clues (driving side, script on signs, bollards, vegetation, road markings) is where transfer comes from.

Compute and budget

Recommended GPU: 1x A100 80GB PCIe on Runpod. Community Cloud lists it at about $1.19/hr and Secure Cloud at $1.39/hr. An H100 PCIe at $1.99/hr is roughly 1.7x faster and costs about the same per unit of work, so pick whichever is available.

Throughput estimate for the recommended config (4 prompts x 8 rollouts x ~300 generated tokens per optimizer step) is 15 to 25 seconds per step, or 150 to 250 steps per hour. This is an estimate, not a measurement. The runbook has a 20-step smoke test whose job is to replace it with a real number before you commit budget.

Phase GPU hours Cost at $1.39/hr
Pod setup, installs, data download 1.0 $1.40
Baseline eval (1000 OSV-5M + 500 pano images) 0.5 $0.70
Learnability filter (8k images x 4 samples) 0.75 $1.05
Smoke test (20 steps) and throughput measurement 0.25 $0.35
Main GRPO run 8.0 $11.10
Checkpoint evals and final eval 1.0 $1.40
Merge LoRA, export, demo 0.5 $0.70
Network volume, 100 GB, prorated for one week ~$1.60
Subtotal 12.0 ~$18
Contingency for crashes and re-runs (60%) ~$11
Total ~$29

You stay under $50 with room to spare. Ways the budget can exceed $50:

  • Swapping to Qwen3-VL-8B and keeping the same step count: roughly +$12.
  • Doubling the main run to 16 hours: +$11.
  • Adding an SFT warm-up with frontier-model-generated reasoning traces (Geo-R1 and GeoAgent both do this): API cost of about $10 to $20 for 3k traces, plus an hour of GPU. Worth it if the first RL run plateaus early.

Do all coding, reward unit tests, and data-prep dry runs on your Mac. Only pay for GPU time when the pipeline is already green on CPU with a tiny sample.

Weekend schedule

Friday evening (no GPU): Hand Part 2 to the coding agent. It builds the repo, writes the reward and parser with unit tests, and dry-runs the data pipeline on the 116 MB OSV-5M test.csv plus one test zip. You review the prompt and the reward weights.

Saturday morning: Create a Runpod network volume and pod. Install, download two training zips, run the baseline eval. Write the baseline numbers down; they are the whole point of the project. Run the smoke test and read the real seconds-per-step figure. Run the learnability filter. Launch the main run by noon with max_steps sized to fit your remaining budget.

Saturday evening: Check W&B. Reward should rise and format compliance should be near 100% within the first 50 steps. Evaluate the latest checkpoint on the 300-image quick set. If country accuracy hasn't moved after 300 steps, stop and debug rather than burn budget.

Sunday: Resume or continue the run if it's still improving. Run the full eval. Merge the LoRA, push the adapter to the Hub, and try the Gradio demo on your own photos. Write up before and after numbers.

Risks and how the plan handles them

  • Version drift. TRL is at v1.11 and needs transformers 5.8+ and vLLM 0.22+. The runbook installs trl[vllm], records the resolved versions to a lock file, and runs the smoke test before anything else.
  • Reward hacking by always guessing a hotspot. Country capping plus the country bonus make a fixed guess lose on most prompts. The eval logs the distribution of predicted countries so you can spot collapse.
  • Degenerate text. Unsloth reports Qwen VL models emitting repeated junk tokens early in GRPO. The format reward is 0 for unparseable output, and max_completion_length caps the damage.
  • OOM. Drop vllm_gpu_memory_utilization from 0.35 to 0.25, then reduce num_generations from 8 to 6, then switch to Unsloth.
  • Too slow. If the smoke test shows over 40 s/step, cut max_completion_length to 384, reduce image max_pixels, or move to an H100.

Stretch goals, in order of value per hour

  1. Qwen3-VL-8B-Instruct with the same pipeline.
  2. SFT warm-up on 2k to 3k reasoning traces before RL.
  3. Multi-view prompts: 2 to 4 images from the same Mapillary sequence (OSV-5M has a sequence column) to mimic GeoGuessr's ability to look around.
  4. Test-time zoom tool, as in the GeoVLM agent variant.
  5. Qwen3.5-4B once the vLLM and TRL friction is resolved.

Part 2: Specification for the coding agent

You are implementing a single-GPU GRPO post-training pipeline that trains Qwen3-VL-4B-Instruct to geolocate street-level photos. Everything must run end to end with one command per phase. Build and test on CPU with tiny samples first; the GPU steps are marked.

Constraints

  • Language: Python 3.11. Package manager: uv.
  • Frameworks: trl[vllm], transformers, peft, datasets, vllm, torch. Install trl[vllm] and let it resolve compatible transformers and vllm versions. Write the resolved versions to requirements.lock. Don't hand-pin versions you haven't verified install together.
  • One GPU, 80 GB. No multi-GPU code paths.
  • Every script takes --config PATH pointing to a YAML file and accepts --limit N to run on a tiny subset for testing.
  • Log to Weights & Biases when WANDB_API_KEY is set; otherwise log to TensorBoard under runs/.
  • Deterministic sampling everywhere: fixed seeds, sorted inputs.
  • Write docs and comments in Google developer documentation style: second person, present tense, active voice, sentence-case headings.

Repository layout

geoguessr-rl/
  README.md                  # how to run each phase, with the runbook commands
  pyproject.toml
  requirements.lock
  configs/
    data.yaml
    eval.yaml
    grpo.yaml
    filter.yaml
  geoguessr_rl/
    __init__.py
    prompt.py                # system prompt, user prompt, answer schema
    parsing.py               # extract <answer> JSON, validate, normalize country
    reward.py                # haversine, GeoScore, reward components, TRL reward fns
    countries.py             # ISO-2 <-> name normalization via pycountry
    data.py                  # OSV-5M download, CSV join, country capping, HF Dataset builders
    eval_metrics.py          # GeoScore, acc@km thresholds, country acc, prediction histogram
  scripts/
    prepare_data.py          # phase 1
    run_eval.py              # phase 2 and 5 (vLLM batch inference + metrics)
    filter_learnable.py      # phase 3
    train_grpo.py            # phase 4
    merge_and_export.py      # phase 6
    demo_gradio.py           # phase 6
  tests/
    test_parsing.py
    test_reward.py
    test_countries.py
    test_data_small.py
  runbook/
    runpod_setup.sh          # apt/uv install, env vars, volume paths
    smoke_test.sh            # 20-step GRPO run + throughput printout

Prompt

Define these in prompt.py:

SYSTEM_PROMPT = (
    "You are an expert GeoGuessr player. You identify where a street-level "
    "photo was taken from visual clues alone."
)

USER_PROMPT = (
    "Where was this photo taken?\n"
    "First, inside <think></think> tags, list the strongest clues in at most "
    "150 words: driving side, language and script on signs, license plates, "
    "road markings and bollards, utility poles, vegetation and climate, "
    "architecture, soil color, and any place names.\n"
    "Then give your final answer inside <answer></answer> tags as JSON with "
    "exactly these keys: \"country\" (English country name), \"lat\" "
    "(decimal degrees), \"lon\" (decimal degrees).\n"
    "Example:\n"
    "<answer>{\"country\": \"Kenya\", \"lat\": -1.2921, \"lon\": 36.8219}</answer>"
)

The conversational prompt for TRL is:

[
  {"role": "system", "content": SYSTEM_PROMPT},
  {"role": "user", "content": [{"type": "image"}, {"type": "text", "text": USER_PROMPT}]},
]

The image column holds the PIL image. TRL's processor inserts the image tokens at the placeholder.

Parsing

parsing.py exposes parse_answer(text: str) -> ParsedAnswer | None.

  • Take the last <answer>...</answer> block. Ignore <think> content.
  • Parse JSON leniently: strip code fences, allow trailing commas, allow single quotes by attempting json.loads then ast.literal_eval.
  • Validate: lat in [-90, 90], lon in [-180, 180], both finite, country a non-empty string.
  • Normalize country to ISO 3166-1 alpha-2 with countries.to_iso2(name), which uses pycountry exact lookup, then pycountry fuzzy search, then a small alias table (for example "USA", "UK", "Russia", "South Korea", "Czechia", "Ivory Coast", "Burma", "Holland", "Macedonia"). Return None for the ISO code if nothing matches, but still return the coordinates.
  • Return None for the whole answer only when there is no parseable <answer> block or the coordinates are invalid.

Tests must cover: valid answer, missing tags, two answer blocks, out-of-range coordinates, string-typed numbers, code-fenced JSON, unknown country name, and completions that are only repeated garbage tokens.

Reward

reward.py implements the following and exposes TRL-compatible functions.

EARTH_RADIUS_KM = 6371.0088
GEOGUESSR_SCALE_KM = 1492.7
FINE_SCALE_KM = 250.0

def haversine_km(lat1, lon1, lat2, lon2) -> float: ...

def geoscore(d_km: float) -> float:
    """GeoGuessr world-map score in [0, 5000]."""
    return 5000.0 * math.exp(-d_km / GEOGUESSR_SCALE_KM)

def distance_reward(d_km: float) -> float:
    return 0.5 * math.exp(-d_km / GEOGUESSR_SCALE_KM) + 0.5 * math.exp(-d_km / FINE_SCALE_KM)

def components(completion_text, latitude, longitude, country_iso2) -> dict:
    """Returns format, country, dist, total, d_km, geoscore for one completion."""

Weights: W_FORMAT = 0.10, W_COUNTRY = 0.30, W_DIST = 0.60. If parsing fails, country = dist = 0 and d_km = None. If the predicted ISO code is None but coordinates parse, format = 1, country = 0, and dist is computed normally.

TRL reward functions receive prompts, completions, and every dataset column as a keyword argument. Dataset columns must be named latitude, longitude, and country_iso2. Implement one combined reward function that returns the total, and three logging-only functions that return the components with reward_weights set to zero for them in the config, so the components show up as separate curves in W&B. Completions arrive as a list of message dicts in conversational mode; take completion[0]["content"].

Tests must cover: haversine_km against three known city pairs within 0.5%, geoscore(0) == 5000, geoscore(1492.7) ≈ 1839.4, monotonic decrease, reward for a perfect answer equals 1.0, reward for an unparseable answer equals 0.0, and reward for a correct-country wrong-coordinate answer lands between 0.40 and 0.70.

Phase 1: Prepare data (CPU, then GPU pod)

scripts/prepare_data.py --config configs/data.yaml

Downloads from osv5m/osv5m on the Hub with hf_hub_download:

  • test.csv (116 MB) and images/test/00.zip.
  • train.csv (2.9 GB) and images/train/00.zip, images/train/01.zip.

Read CSVs with polars, selecting only id, latitude, longitude, country, region, sub-region, city, sequence. Inspect the header at runtime and fail loudly if a column is missing; the country column is expected to be ISO-2 but verify with an assertion that all values are two uppercase letters, and log a sample if not. Join to the zip contents by filename stem equals id.

Build three parquet files under data/, each with columns image_path, latitude, longitude, country_iso2, region, city, sequence:

  • train_pool.parquet: from the two train zips, cap at 400 images per country using a seeded shuffle. Expect 15k to 30k rows.
  • eval_osv.parquet: 1000 rows from the test zip. Stratify: up to 10 per country first, then fill randomly. Fixed seed.
  • eval_quick.parquet: the first 300 rows of eval_osv.parquet.

Also build eval_pano.parquet from stochastic/random_streetview_images_pano_v0.0.2: 500 rows, seeded, with images resized so the long edge is 1024 px and saved as JPEG under data/pano/. Map its country column to ISO-2 with countries.to_iso2.

Image preprocessing for all sets: load, convert to RGB, and if either dimension exceeds 1024 px, resize so the long edge is 1024 px. OSV-5M images are 512 px tall and about 800 px wide, so most pass through unchanged. Store the processed images as files and load lazily in the Dataset to keep memory low.

Data sanity checks the script must print: row counts, number of countries, top-10 countries by count, and a histogram of image widths.

--limit 200 must run on a laptop against only test.csv and one test zip, skipping the train download, for the CPU dry run.

Phase 2: Baseline eval (GPU)

scripts/run_eval.py --config configs/eval.yaml --model Qwen/Qwen3-VL-4B-Instruct --split eval_osv

Use vLLM offline inference directly (vllm.LLM with limit_mm_per_prompt={"image": 1}), max_model_len=4096, greedy decoding (temperature=0), max_tokens=512. Batch all prompts in one generate call. Accept --lora PATH to evaluate an adapter with enable_lora=True.

Write outputs/eval/<run_name>/<split>/predictions.jsonl with one line per image containing the raw completion, parsed fields, d_km, geoscore, and per-component rewards. Write metrics.json with:

  • format_rate, country_acc, mean_geoscore, median_km,
  • acc_at_km for thresholds 1, 25, 200, 750, 2500,
  • pred_country_top10 (a histogram; use it to detect collapse),
  • n.

Print a one-line summary. Run on both eval_osv and eval_pano.

Use max_pixels = 1024 * 28 * 28 in the processor kwargs, which caps an image at roughly 1000 visual tokens. Don't lower it below 512 * 28 * 28.

Phase 3: Filter for learnable prompts (GPU)

scripts/filter_learnable.py --config configs/filter.yaml

  • Sample 8000 rows from train_pool.parquet (seeded, country-capped at 200).
  • With vLLM, generate n=4 completions per image at temperature=1.0, top_p=0.95, max_tokens=512.
  • Compute the total reward for each completion. For each image record mean_reward, std_reward, and country_hits (0 to 4).
  • Keep an image if 0 < country_hits < 4 or std_reward > 0.08. Call this set A.
  • Add a seeded random 20% of the rejected images. Call this set B.
  • Write train_learnable.parquet = A ∪ B, shuffled with a seed. Print the sizes of A, B, and the country distribution of the result. Target 3k to 5k rows. If A is under 1500 rows, lower the std_reward threshold to 0.05 and rerun without regenerating (cache the completions to disk).

Also write filter_stats.json with the base model's mean reward and country accuracy on the 8000 sampled images, as a second baseline.

Phase 4: GRPO training (GPU)

scripts/train_grpo.py --config configs/grpo.yaml

Build the dataset from train_learnable.parquet with columns prompt (conversational list above), image (PIL, loaded lazily via datasets.Image() feature), latitude, longitude, country_iso2.

Model and LoRA:

model_name = "Qwen/Qwen3-VL-4B-Instruct"
dtype = bfloat16
attn_implementation = "flash_attention_2"   # fall back to sdpa if unavailable
peft_config = LoraConfig(
    r=32, lora_alpha=64, lora_dropout=0.0, bias="none",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
    # Restrict to the language model. Exclude the vision tower and the
    # merger/projector so vLLM can serve the adapter.
    exclude_modules=r".*visual.*",
    task_type="CAUSAL_LM",
)

Verify after wrapping that no LoRA module name contains visual. Fail if it does.

GRPOConfig values for configs/grpo.yaml:

output_dir: outputs/grpo-qwen3vl4b
run_name: grpo-qwen3vl4b-v1
seed: 42
learning_rate: 1.0e-5
lr_scheduler_type: constant_with_warmup
warmup_steps: 20
max_grad_norm: 0.2
bf16: true
gradient_checkpointing: true
per_device_train_batch_size: 8        # completions per micro-batch
gradient_accumulation_steps: 4        # 32 completions = 4 prompts per optimizer step
num_generations: 8
generation_batch_size: 64             # 8 prompts sampled per vLLM call
max_prompt_length: null               # required for VLMs; don't truncate image tokens
max_completion_length: 512
temperature: 1.0
top_p: 1.0
loss_type: dapo
beta: 0.0
scale_rewards: group
mask_truncated_completions: true
use_vllm: true
vllm_mode: colocate
vllm_gpu_memory_utilization: 0.35
vllm_max_model_len: 4096
reward_weights: [1.0, 0.0, 0.0, 0.0]  # total, then logging-only components
max_steps: 1200                       # resize after the smoke test
save_steps: 200
save_total_limit: 3
logging_steps: 1
log_completions: true
num_completions_to_print: 2
report_to: wandb

Processor kwargs: max_pixels = 1024 * 28 * 28, min_pixels = 256 * 28 * 28.

Add a --resume_from_checkpoint auto option that picks the latest checkpoint in output_dir. Add a callback that every save_steps runs run_eval.py --split eval_quick --lora <checkpoint> in a subprocess after training releases the GPU, or, if that's too disruptive, write a separate scripts/eval_checkpoints.sh that loops over saved checkpoints. Prefer the separate script; it's simpler and the run doesn't stall.

Budget arithmetic the script must print at start: given seconds_per_step from --seconds_per_step (measured in the smoke test) and --budget_hours, print the recommended max_steps and exit if --dry_run.

Phase 5: Final eval (GPU)

Run run_eval.py with --lora outputs/grpo-qwen3vl4b/checkpoint-<best> on eval_osv and eval_pano. Produce outputs/report.md with a before and after table for every metric, plus five example completions where the adapter beat the base model by the largest GeoScore margin and five where it did worst.

Phase 6: Export and demo (GPU or CPU)

scripts/merge_and_export.py: merge the LoRA into the base model with merge_and_unload, save to outputs/merged/, and optionally push both the adapter and the merged model to the Hub under a user-provided repo ID.

scripts/demo_gradio.py: a single-image Gradio app that shows the <think> text, a map marker via a static map image or a folium HTML embed, and the predicted country. Load the merged model with transformers, not vLLM, so it runs on a Mac with device_map="auto" for smoke testing.

Runbook

runbook/runpod_setup.sh must do the following on a fresh Runpod PyTorch pod (CUDA 12.x image) with a network volume mounted at /workspace:

  1. Install uv, then uv venv /workspace/.venv and activate.
  2. uv pip install "trl[vllm]" peft datasets polars pycountry pillow wandb gradio flash-attn --no-build-isolation and freeze to requirements.lock. If flash-attn fails to build, continue without it and set attn_implementation=sdpa.
  3. Export HF_HOME=/workspace/hf, HF_HUB_ENABLE_HF_TRANSFER=1, and VLLM_WORKER_MULTIPROC_METHOD=spawn.
  4. Run python -c "import trl, transformers, vllm; print(...)" and print versions.

runbook/smoke_test.sh runs train_grpo.py with max_steps=20, save_steps=1000, report_to=none, and --limit 256, then prints the average seconds per step over steps 5 to 20 and the recommended max_steps for an 8-hour budget. It must exit non-zero if any step produced format_rate < 0.5 after step 10, since that indicates a broken prompt or parser rather than a training problem.

Order of operations on the pod:

bash runbook/runpod_setup.sh
python scripts/prepare_data.py --config configs/data.yaml
python scripts/run_eval.py --config configs/eval.yaml --split eval_osv --run_name base
python scripts/run_eval.py --config configs/eval.yaml --split eval_pano --run_name base
bash runbook/smoke_test.sh
python scripts/filter_learnable.py --config configs/filter.yaml
python scripts/train_grpo.py --config configs/grpo.yaml --seconds_per_step <measured> --budget_hours 8
bash scripts/eval_checkpoints.sh
python scripts/run_eval.py --config configs/eval.yaml --split eval_osv --run_name rl --lora outputs/grpo-qwen3vl4b/checkpoint-<best>
python scripts/run_eval.py --config configs/eval.yaml --split eval_pano --run_name rl --lora outputs/grpo-qwen3vl4b/checkpoint-<best>
python scripts/merge_and_export.py --lora outputs/grpo-qwen3vl4b/checkpoint-<best>

Acceptance tests

Before any GPU is rented, all of the following pass on a laptop:

  1. pytest tests/ passes.
  2. python scripts/prepare_data.py --config configs/data.yaml --limit 200 produces eval_osv.parquet with 200 rows and prints the sanity checks.
  3. python scripts/run_eval.py --config configs/eval.yaml --split eval_osv --limit 4 --backend transformers --device cpu runs the base model on 4 images in bf16 or fp32 on CPU and writes a valid metrics.json. This confirms the prompt, chat template, parser, and metrics code end to end. It's slow; that's fine.
  4. python scripts/train_grpo.py --config configs/grpo.yaml --dry_run --seconds_per_step 20 --budget_hours 8 prints the dataset schema, the LoRA target module list with zero visual matches, and recommended max_steps = 1440.

On the GPU, the smoke test passes and reports seconds per step.

Failure handling

  • OOM during vLLM init: lower vllm_gpu_memory_utilization to 0.25.
  • OOM during backward: set per_device_train_batch_size: 4 and gradient_accumulation_steps: 8. Same effective batch.
  • vLLM rejects the LoRA adapter (vision modules): the exclude_modules regex is wrong for this transformers version. Print model.peft_config target names and fix the regex.
  • format_rate stuck near 0 after 30 steps: print 10 raw completions. Usually the chat template added a thinking block or the model ignores the <answer> tag. Add a one-shot example to the user prompt.
  • Reward rises but country_acc on eval_quick doesn't: check pred_country_top10. If one country dominates, the model found a hotspot. Lower the per-country cap to 250 and raise W_COUNTRY to 0.40.
  • Slower than 40 s/step: reduce max_completion_length to 384 and max_pixels to 768 * 28 * 28.

Plan B: Unsloth

If TRL plus vLLM colocate can't be made to fit or run in the smoke test within one hour of debugging, switch to Unsloth. Follow the structure of the Unsloth "Qwen3-VL (8B) Vision GRPO" notebook: FastVisionModel.from_pretrained with load_in_4bit=False, fast_inference=True, gpu_memory_utilization=0.6, max_seq_length=4096; get_peft_model(finetune_vision_layers=False, finetune_language_layers=True, r=32, lora_alpha=64); then the same GRPOConfig values and the same reward functions. Keep the data, eval, and export scripts unchanged.


Sources