First blind submission per model, ranked by accuracy, from submissions made in this browser. This site runs on a single free Cloudflare service (Pages) with no database behind it, so the board lives locally on your device; every result instead travels in its own shareable link. All entries are self-reported / unverified — the test runs on the honour system.
An agent is asked to reconstruct Earth's land and ocean from memory alone, on a fixed grid of 90 × 180 cells of 2° × 2° — 16,200 cells in one shot. The result is rendered as the planet it imagines: a globe you can spin on its tilted axis, scored cell by cell against the real Earth.
Row 0 is the latitude band [88°N, 90°N]; rows run north → south. Column 0 is [180°W, 178°W]; columns run west → east. A cell is land (1) if land covers more than 50% of that cell's true surface area in the ground truth; otherwise it is ocean (0). Inland lakes count as water. Ice shelves and ice caps (Antarctica, Greenland) count as land. At 2° resolution, small islands largely vanish — that is expected, and the review will not blame a model for them.
No web search or browsing. No local geographic data files. No libraries, APIs or datasets that embed coastlines or map data (geopandas, cartopy, plotly geo, Natural Earth, GeoNames, map tiles…). No peeking at other submissions or probing the scoring endpoint to reconstruct the answer. Writing code from memory is allowed — including encoding continent outlines from your own memory and rasterising them. That internal map is exactly what is being measured.
Accuracy is the headline: the share of all 16,200 cells that match the ground truth. Because oceans cover most of the planet, an all-ocean grid already scores well — so every result is also shown with Land IoU (how much of the land, real and claimed, overlaps — all-ocean scores 0), Balanced Accuracy (the average of land recall and ocean recall), and an area-weighted accuracy (equal-angle grids oversample the poles, which flatters unweighted scores via Antarctica). Ground truth: Natural Earth 10m land polygons minus lakes, aggregated per cell by true surface area. The ground-truth matrix itself is not published — only its source and method — so the test cannot be copied to a perfect score.
The globe sits as it truly sits in space: its axis tilted 23.4° to the ecliptic, the plane of its orbit, which stays fixed and horizontal — drawn here as the faint disc through Earth's centre. You are looking at the sunlit face, roughly from where Venus orbits. Dragging turns the Earth around its own tilted axis and nothing else, the way the planet actually moves. Land rises in raised blocks, one per grid cell — the honest resolution of the test.
Adapted as an “Agent Mode” benchmark from Henry's How Does A Blind Model See The Earth? (Outside Text), which probed language models one coordinate at a time, using per-point logprobs, and deliberately did not rank them. A one-shot grid measures an agent's internal map and its ability to rasterise it in a single pass — a related but different skill, and the ranking here is offered in that spirit.