Sid Savara / Engineering leadership, personal writing, and notes from Sid Savara Sat, 19 Sep 2026 19:55:10 +0000 en-US hourly 1 https://wordpress.org/?v=7.1.1 /wp-content/uploads/2026/05/sid-savara-favicon-s-512-100x100.png Sid Savara / 32 32 Reflections on Collaborating With an AI Agent /reflections-collaborating-with-ai-agent/ Wed, 09 Sep 2026 19:00:00 +0000 /reflections-collaborating-with-ai-agent/
PART8OF 8

This is Part 8 of 8, and the conclusion of The Projects AI Makes Possible, a series about building a private almanac from a 1992 geography game with a coding agent. In Part 7: Strategic Pruning: Deciding What Belongs in a Minimum Viable Product, I narrowed the finished almanac to what helped us play the game. Here, I want to look back at the collaboration that made it possible.

Building an almanac my family could use during a game gave me a chance to see how far I could take a project with an AI agent. Now that it worked, I wanted to look back at my own involvement: what I needed to understand, what I could leave to the agent, and how we found ways to work together.

I Didn’t Need To Know Every Detail

You’ll notice that throughout this whole series, I haven’t mentioned the specific technologies used for data extraction, hosting the local website, etc. The code was Python (57 separate programs for extraction and reformatting files!), the website is powered via TypeScript, HTML, CSS, the database is SQLite, and review artifacts are stored in JSON, Markdown, and TSV. However, those are details that I never even needed to think about, and only know about due to curiosity when I was watching it work, and just now looking at GitHub as I wrote this.

For production code or something I needed to support longer term, I think more familiarity with the code and architecture is warranted – but for smaller tools such as this, it is possible to be completely removed from the specifics of technologies used and implementation, and via interacting with the agent, be concerned only with the outcomes and end goals.

Reflecting On The Project As A Whole

I began this project because an agent changed the cost of trying. I could give it an uncertain direction, let it spend hours exploring, and decide whether the result was worth another step. The failed approaches did not consume the same amount of my own time, so an idea that would normally have required a more expensive commitment to learning how to decode old DAT files became reasonable to pursue.

What changed after that was not simply the amount of work the agent could do. We gradually found better ways to divide the work between us.

When chat and folders became tedious, we created pages I could review from a phone or tablet. When an image format resisted ordinary inspection, I supplied screenshots from the running game and looked for a recognizable signal among the agent’s experiments. When a polished rewrite lost details, we modeled completeness as a relationship in a database that my agent could both document and later validate. When real gameplay found missing searches, the agent expanded each example into a broader audit.

The Distinction Between Judgment, Decisions and Implementation

Each time, my judgment began as a small observation: this page is hard to review, those arches look familiar, this paragraph lost too much, this phrase should have been found. Collaborating with the agent and delegating (with varying levels of specificity, and sometimes varying degrees of success) could turn that observation into a tool or a repeatable check, then apply it beyond the example I had supplied.

That was the type of collaboration I was excited to have: compared to coding in a high-level language where I’m still specifying the logic myself, here I was directing broadly and delegating implementation decisions. The agent could usually fill in gaps independently from directions given much less precisely than actual code, and then run the experiments. Meanwhile, I provided guidance on whether the process was producing the right kind of answer and clarified the goals when it needed redirecting. I was generally successful in this, as you saw in Part 4 where I gave broad directions and the agent solved the implementation details for the maps by itself, while also learning in Part 5 that I still needed to identify any requirements that I absolutely needed the agent to adhere to (in that case, lossless paraphrasing).

To finally bring the project home, we then went in the opposite direction. After using the agent to explore broadly, I had to narrow the final product back to the original need. Pruning carefully let the core of the product become the entire experience, and helped the almanac disappear and feel like an intuitive part of a broader experience.

We started with an old geography game, a pile of unfamiliar files, and an idea I previously would have been skeptical about investing in, given both the manual work involved and the uncertainty around effort and success. We ended with an almanac my family could open on a tablet, use to crack clues during a case and perhaps even learn a thing or two in the process.

]]>
Strategic Pruning: Deciding What Belongs in a Minimum Viable Product /simplifying-products-built-with-ai-agents/ Tue, 08 Sep 2026 19:00:00 +0000 /simplifying-products-built-with-ai-agents/
PART7OF 8

This is Part 7 of 8 in The Projects AI Makes Possible, a series about building a private almanac from a 1992 geography game with a coding agent. In Part 6: AI Solved the Problem I Pointed It At. I Had Aimed Too Narrowly., testing with gameplay revealed a missing set of source data. After resolving that and successfully using the almanac for a few playthroughs, I had one final step before calling this complete: deciding what should actually make it into the product – and what to cut.

“It is only by selection, elimination, and emphasis that we get at the real meaning of things.”

Georgia O’Keeffe

By the time the almanac worked, the project behind it had become enormous. Two parallel pieces of work had grown throughout: an ever-expanding research workspace and, more slowly, the product itself.

The research workspace held extraction manifests, one-off programs to extract and reformat data, decoder experiments, groupings of images for review, various attempts at maps, a database, data review artifacts, and notes on different directions we could take. Hundreds of image files had accumulated along the way, many of them failed experiments that were still useful as evidence.

The agent had made it inexpensive to create all of that. It could build another review page, preserve another branch, or explore another feature simply because I prompted it in a direction.

In the product itself I similarly had cast a wide net to start, intermingling the almanac with the review databases, pages, and some speculative experiments. Now it was time to narrow the functionality to just what enhanced the experience while playing the game with it.

The product question was no longer what else the agent could build. It was how to whittle what we had built down to the most relevant and useful parts.

The Mess Still Had a Job

The experiments we tried explored many different directions, while the product should be intuitive and almost disappear.

Earlier, when I tested the output and compared it to what we had extracted, I wanted to be able to see everything and easily navigate from the almanac all the way to the sources and reasons for how, where and why the agent had gotten and included each piece of information. Resource IDs helped us connect files. Raw links let us inspect an extraction. Decoder galleries made it possible to compare groupings of images and recognize the one that was getting close. Review tables showed source sentences beside rewritten facts.

While iterating and researching, those artifacts helped us work together, and having a way to drill down to them from the almanac itself helped me audit, test and trace things back to their source and see them in context. The finished almanac, however, only needed them as source material. Once I trusted that the data was good, I could focus on the user experience.

What Survived the Cut

The country page gradually reduced itself to a short list:

  • A flag
  • A modern regional map
  • A small fact box
  • The rewritten country information
  • Additional facts that explained searchable phrases
  • Search

There were also small usability quirks we discovered in testing early iterations, which we gradually polished. Initially countries appeared in the order of their internal IDs. Pages exposed links and labels that tied back to source extraction work. Long sections pushed the search box off the screen.

These usability defects became apparent as we started playing with the almanac, and I worked with the agent to resolve them one at a time – with fixes often straightforward. Countries became alphabetical, with the capital shown before the continent. We made the sidebar collapsible to create more room for the article, and kept its reopen control on the left so it still read as a sidebar control rather than the main menu. We made the search remain visible while the page scrolled.

The search widget also improved through iteration. Pressing Enter originally jumped directly into the first match even when several results were plausible. In the finished version, the search displays multiple matching results but waits for the reader to choose. The selected result then opens the right country and moves to the exact explanation instead of dropping the reader at the top of a long page.

Those details were less technically dramatic than decoding an undocumented image format. But when we used the almanac during a game, they made the experience much more pleasant and brought it closer to the goal of becoming intuitive enough to disappear. It was also a lot of fun asking the agent to fix it as we were playing, and during the course of the case, reloading the page and seeing the changes immediately applied!

The finished almanac page
The final page kept the map, flag, facts, country information, and search.

The resulting site was only about ten megabytes, and most of that was the thirteen regional map images. The facts and search material were small enough to load directly in the browser.

That meant the almanac did not need user accounts, a server-side database, or an API. Initially I thought I would deploy it as an app installed on the tablets, and perhaps in the future I will pursue that. However, the static site we built “just to test” actually performed remarkably well. It was fast, easy to cache, inexpensive to host on a local server (or one of any number of free or inexpensive web hosts), and required no personal information from the children using it.

For a bounded reference collection like this, the static site was totally serviceable as an MVP.

Some Successes Weren’t Features

For all my excitement about our breakthrough decoding the scenic images, they did not make the finished version. Recovering them had been a meaningful part of the project, and the experiments taught us how the old resource files worked. But they did not add anything to the experience, and the game already displayed them when we visited each country. We also had working background audio, but it did not improve the tablet experience enough to justify inclusion and yet another widget on the page, especially since the game itself already played it.

While that work did not make the final product, some of what we learned along the way influenced other parts of the product. The image decoder had answered a technical question and ultimately let us decode the source maps. The audio experiments proved that the old sound records could be understood.

And ultimately, the journey I took in getting to the end goal was enjoyable – unlocking parts of the data files and being able to independently load and play them was a fun reward unto itself.

The experiments and review tools were still there if I wanted to return to them. They just no longer needed to be part of the almanac we opened during a game.

]]>
AI Solved the Problem I Pointed It At. I Had Aimed Too Narrowly. /ai-agent-data-coverage-blind-spots/ Mon, 07 Sep 2026 19:00:00 +0000 /ai-agent-data-coverage-blind-spots/
PART6OF 8

This is Part 6 of 8 in The Projects AI Makes Possible, a series about building a private almanac from a 1992 geography game with a coding agent. In Part 5: The Agent Built What I Asked, But Not What I Meant, the country data set and presentation became complete enough to review sentence by sentence. Playing the game, though, revealed that we still had gaps in clue coverage.

“The first principle is that you must not fool yourself—and you are the easiest person to fool.”

Richard Feynman, 1974 Caltech Commencement Address

In our initial tests, the core functionality worked. The almanac loaded our data and the appropriate maps, and search returned matching results. We had assumed that the country data contained everything needed to play the game. That assumption turned out to be incorrect.

Excited to test it for real, we loaded it up on a tablet and started the game.

During a case, we searched for a phrase the game had given us and found nothing. One of the early misses was paper mill. The source material connected Sweden’s forests with its pulp and paper industry, but none of that information appeared in a form the search could find.

The fix looked small. Add the missing phrase to Sweden and move on.

Then another search failed. And another.

Hari Rud was missing from Afghanistan. A Tamil Hindu temple phrase had been summarized too far. Mongolia needed the words yurt, snow leopards, and wild asses. Buddhist could fail even when the page contained Buddhism. An older spelling such as Szechuan disappeared when the prose used only Sichuan.

The search worked. The index worked. The country pages worked. The almanac was functioning correctly: it just didn’t actually have the data I needed!

Fixing One Wasn’t Enough

Each failed search told us what category of source material to audit next.

At first, the agent repaired the individual entries I found during play. That made the next search work, but it did not answer the more worrying question: what else had gone missing for the same reason?

There wasn’t a single cause behind the failures. They fell into several categories. I asked the agent to research and explain them to me, and their explanations revealed a few different causes.

Some phrases had been removed by an overaggressive summary. Some crossed source lines and made sense only when the neighboring records were read together. Some were historical names or spellings that modern prose naturally replaced. Others described a person’s appearance, hobby, or favorite food and did not honestly belong to any one country.

Each missing result suggested that we might have misclassified a whole group of records.

The Audit Kept Growing

I could give the agent one concrete miss and ask it to look for the pattern behind it. When paper mill was absent, the work expanded beyond Sweden to other country-specific fragments that had been filtered or summarized. When a phrase depended on the line before it, the agent widened the review to neighboring lines and multi-line groups. When a description did not identify a country, it audited the shared evidence instead of forcing the phrase into an arbitrary country page.

The review system began distinguishing among:

  • Country-specific material that needed an explanatory fact.
  • Context recovered from adjacent lines or a multi-line group.
  • Shared descriptions that needed somewhere else to live.
  • Historical or dated language that should remain searchable with context.
  • Edge cases where aggressively pulling in all the data also captured unusable fragments, which we then trimmed rather than turning them into almanac entries.

The agent would scan all our records, group likely failures, generate review tables, and rerun the coverage checks. My contribution was often much smaller and more specific: this phrase appeared in the game, this search result was missing or unhelpful.

Over time, the agent got better at taking one failure and expanding the audit before returning with another supposedly complete pass. The loop grew beyond one missing phrase followed by one patch. A concrete failure became a reason to reconsider the broader approach and fix an entire group of missing clues, including ones we hadn’t explicitly identified.

The Result Had to Teach

We could have made the search return something by dumping the original fragments into, say, a keyword field. The almanac still would not have explained what the phrase meant.

When I searched for Hari Rud, I wanted the result to explain that it is a river flowing west from central Afghanistan past Herat toward Turkmenistan. A page that merely calls it “a geography reference for Afghanistan” technically matches the phrase, but provides very little context or education.

The same applied to paper mill. The explanation needed to connect the phrase to Sweden’s forests and its pulp-and-paper industry while keeping the exact words a player might type.

The Sweden page opened to the paper-industry fact
Opening the result took us to the country context behind the phrase.

Not every phrase could be handled that way. Some clues were specific to helping identify a suspect: ebony-colored hair, malachite-colored eyes. In the game such clues help describe a suspect to allow the player to obtain a warrant for the suspect’s arrest. While I could have finagled it (e.g. put malachite in a country that happens to have malachite), those words were not really attached to any country.

Those descriptions became a separate collection of World Facts. They could remain searchable and understandable without pretending to be country facts, and still be educational.

Every Record Accounted For

Coverage was not complete until every source record had a destination or a reason for exclusion.

The final audit accounted for 5,441 distinct source records. That does not mean the site contains 5,441 polished facts. Some records repeat, some contribute only part of an explanation, and some are structural material rather than content.

Of those records, 4,674 routed to country-specific explanations. Another 217 shared descriptions routed to World Facts. The remaining 550 were records with an explicit reason not to appear as almanac entries. None remained unresolved in the audit.

When I started playing with what I thought was the “complete” almanac, I expected to focus on ease of use and friction. Playing the game also smoke tested our clue coverage, exposing gaps I didn’t even realize we had missed. One failed search could have led to one repaired entry. Instead, it exposed a pattern, and fixing it consistently led to a broader review of the set of clues.

After a few iterations – naturally, while I slept or did other things – the almanac had a much more complete set of facts. The next time we played, search behaved the way we had wanted from the beginning. We could hear an unfamiliar phrase, find it on the tablet, understand why it mattered, and continue the case without leaving the experience.

]]>
The Agent Built What I Asked, But Not What I Meant /clear-requirements-for-ai-agents/ Sun, 06 Sep 2026 19:00:00 +0000 /clear-requirements-for-ai-agents/
PART5OF 8

This is Part 5 of 8 in The Projects AI Makes Possible, a series about building a private almanac from a 1992 geography game with a coding agent. In Part 4: Learning to Trust an AI Agent With the Goal, Not Dictate Every Step, I gave the agent the original regional maps and woke up to a freshly generated matching set with colors and styling of my choosing. Maps could be compared by sight. The country text introduced a less obvious problem: a rewrite could sound good while losing critical information.

“Language is the source of misunderstandings.”

Antoine de Saint-Exupéry, The Little Prince

By this point, we had extracted all the country articles from the game, however they were all individual, discrete fact based sentences. I wanted to turn them into almanac prose that would have consistent sections and ordering across all countries. I needed to preserve the original information and later fill in additional data to give all the country pages roughly the same level of detail.

The agent rewrote the country facts into a consistent structure across all the pages. In the game data, the facts appeared in different orders from one country to another: some led with history, others with geography, and so on. I wanted to rewrite these to naturally be in the same order instead across all countries, under subheadings, in the almanac.

When the rewrites were ready for review, I looked at Greece to see how it had fared. The first version turned several short bullets into a very smooth-sounding overview:

Greece is a mountainous country with a rugged coastline and thousands of islands. Athens is its capital, and ancient Greece made lasting contributions to philosophy, science, drama, and art.

It sounded fine. Nothing in it was obviously false. It was short, clear, and grammatically complete.

Then I looked at the source article.

The original described Greece’s mountains, rugged shoreline, and thousands of islands. It named Crete, Rhodes, Delos, and Mykonos. It included Macedonian rule beginning in 338 B.C., followed by roughly 2,000 years of occupation. It connected Athens and Sparta with philosophy, science, drama, and art. It named the Acropolis and the Parthenon and preserved the words Hellas and Hellenes.

The broad outline had survived. Nearly every name, date, and example was gone.

The Rewrite Sounded Fine… But It Wasn’t

That comparison made me realize I had simply assumed that paraphrasing would keep all the details, but I needed to be explicit: what I really wanted was lossless paraphrasing. Every number, date, place, example, comparison, and factual claim had to survive in the new prose.

This is noteworthy because my broad-strokes instructions to the agent caused friction. Contrast this with the review pages, maps, and image decoding, where I gave high-level goals, the agent stayed quite well aligned, and my unspoken assumptions held. Here, I had omitted a critical assumption and had to backtrack to make it an explicit requirement: the rewritten facts had to contain everything the source noted.

This introduced another issue. I could proofread the Greece entry because I was comparing one paragraph with another. Proving completeness across dozens of almanac entries was going to be difficult, especially given that I didn’t want a word-for-word copy. By definition, I wanted a paraphrase and reorganization of the data.

Mapping Sources to Destinations

Once again, I worked with the agent as a partner and, as funny as it sounds, asked for its opinion about my proposed solution. I proposed giving every source sentence a unique identifier based on its location. A sentence might be identified as:

Greece Fact 1, Sentence 1

Its rewritten destination might be:

Greece, Natural Environment, Sentence 1

The structure depended on the relationship between them. Each old sentence had to point to a new sentence, so if we had a missing destination I would know. Then, for completeness, the agent could check whether each fact from the source was also present in the destination.

Traceability meant logging the changes as structured data that we could inspect and compare.

The review format became:

  • Old location
  • Old fact
  • New fact
  • New location

The agent implemented the relationship in a local database and built a tablet-friendly review page around it.

The sentence review page with Greece source and rewritten facts side by side
The Greece rows made each source sentence and its destination visible together.

Instead of asking whether the new Greece article felt complete, I could inspect one source sentence at a time. I could flag a weak paraphrase, search by country, filter to pending rows, and distinguish source-backed rewrites from newly researched facts.

A newly added fact had no old location and carried its own source reference. I had asked the agent to expand the facts using Wikipedia and other public sources so every country had roughly the same kinds and depth of information, beyond what the game contained. An original fact without a new destination showed up as pending.

In the review snapshot, the system accounted for 858 source sentences across 61 entries, plus 25 added-only facts. No source sentence was left pending.

Initially, the database and review page gave me a way to double-check the agent’s rewrites. After I was satisfied with a few spot checks, I decided once again that I didn’t need to do this work myself! From then on, the same structure gave the agent a way to check its own work.

The same structure that helped me review the agent also taught the agent how to check its own work.

It could query for source sentences without destinations, run coverage checks for names, dates, numbers, and other details, repair the gaps it found, and regenerate the review page. The agent was increasingly finding and resolving omissions independently before I reviewed the next result.

Then and Now, Together

The quick facts exposed a related issue. The game had captured the world in the early 1990s. Greece used the drachma then; today it uses the euro. Populations, flags, capitals, and country names likewise reflected that time.

However, I didn’t want an almanac stuck purely in the past. It also needed to be understandable now. Replacing every old value would remove game-era context, while presenting every old value as current would be misleading.

The agent built another review page that placed extracted game values beside current structured references.

The quick-facts page comparing game values with current reference values
The review separated historical differences from extraction errors.

That comparison let us make field-by-field decisions. The public fact box could show a current population and currency while preserving a former currency or 1990s flag description where it helped connect the almanac to the game.

The same review also exposed fields that looked simple but were not. Head of state changes frequently. Official-language labels can hide federal, regional, or de facto distinctions. As I’ll share in future articles, this ultimately led me to have a small subset of information in a “Fact Box” section with items such as capital, currency, etc. – but omit, for example, head of state, since that changed much more frequently.

Now We Could Trust It

The country rewrite began as a request for better prose. Reviewing the first country made me realize we first needed to deal with data integrity.

After refining the lossless paraphrasing requirement and aligning the sentence-level relationships, the agent turned that idea into a database, review pages, filters, and coverage checks. My review helped clarify what I actually needed, and the agent could then use the review database by itself (themselves?) to audit later batches without waiting for me to identify every gap.

The system could now account for every extracted country sentence. The main country facts had been paraphrased into new prose and reorganized into subsections without dropping the information from the game. Where the passage of time mattered, the almanac could place 1990s context beside present-day facts instead of pretending they described the same moment.

The result was a real almanac we could use. A country page gave us a clear overview, preserved the substance of the game article, and added enough current context to also provide a little bit of education in the process.

But Wait…There’s More…Data??

Playing with it revealed the next gap. We could open a country and learn about it, but searches for some phrases mentioned during a case still returned nothing. In Part 2 I casually mentioned setting aside some “clues”. Those missing clues lived in a much larger collection of fragments, shared descriptions, and lines that only made sense beside the line before or after them and turned out to be critical as well.

The broad set of country facts was complete and usable. Now we needed to do the same with the standalone clues.

]]>
Learning to Trust an AI Agent With the Goal, Not Dictate Every Step /trusting-ai-agents-with-goals/ Sat, 05 Sep 2026 19:00:00 +0000 /trusting-ai-agents-with-goals/
PART4OF 8

This is Part 4 of 8 in The Projects AI Makes Possible, a series about building a private almanac from a 1992 geography game with a coding agent. In Part 3: The Agent Explored Different Paths. I Chose Which To Follow., screenshots from the running game helped us recover the original scenic images. The little regional maps were now visible too, and it was time to figure out the best way to integrate them into the almanac so they remained connected to the game, but didn’t feel out of place on a modern tablet.

“Never tell people how to do things. Tell them what to do and they will surprise you with their ingenuity.”

George S. Patton, War As I Knew It

Once the scenic pictures were decoding correctly, the regional maps looked like they should be straightforward to tackle next. We could already see them, and there were far fewer.

The game did not contain a separate map for every country. It reused the same map across a group of countries in the same region. Afghanistan, Iran, Pakistan, and several of their neighbors all pointed back to the Southwest Asia map. Altogether, the game used thirteen regional maps.

I wanted to preserve that connection. This was a geography game with an almanac, so moving from a country in the game to a familiar regional map on the tablet felt natural. A corresponding version of the same game map would help the two screens feel like parts of the same experience.

The agent extracted the Southwest Asia map. It was recognizable and about 190 by 165 pixels, complete with yellow countries, blue checkerboard water, discrete borders, and white label boxes.

The Afghanistan map as it appeared in the running DOS game
The original map clearly showed the geographic neighborhood around Afghanistan.

On the original screen, those choices made sense, and the map fit perfectly into the layout. On a modern tablet, I did not have to limit my color choices, and simply enlarging the image felt out of place.

Repairing the Pixels

Our first idea was to upscale and recolor the existing maps.

We considered versions at four and six times the original size. Keeping the pixels sharp preserved the coastlines. Smoothing them blurred the borders and labels. Recoloring the water still left the checkerboard pattern underneath.

Once we were replacing the borders, labels, colors, and water, restoration turned out to be more effort than starting over.

The list of changes I was directing the agent to make grew quickly:

  • Remove the checkerboard pattern from the water.
  • Make the ocean lighter near the coast and darker farther away.
  • Smooth the coastlines without losing the borders between countries.
  • Replace the original label boxes with regenerated, crisp text.
  • Keep roughly the same regional view as the original map.

Once we were replacing the water, coastlines, borders, labels, and colors, making a fresh map was easier than using a small image as the starting point for all those changes.

Pivoting and Giving the Agent More Autonomy

By now, I was very comfortable giving the agent broad, general instructions and expecting it to fill in the details. There was also a constant back-and-forth dialogue. I saw the maps being restored, offered critical feedback, and discussed with the agent why this approach felt inefficient. For something as universal as a map, I didn’t have to constrain us to using those exact pixels as a starting point – what I was more interested in was the end-point of an analogous, upscaled version.

So I pivoted: and I asked the agent to pivot. We already had the original maps, and I didn’t care how it arrived at the outcome. I asked it to set aside my original directive to literally upscale the source images. I just wanted higher-resolution maps analogous to them, using what I considered the colors I’d want on the map if I was unconstrained: blue water, green land masses, and reasonable divisions between countries.

I left identifying the map groupings, geographic coordinates, label placement, and rendering details to the agent. The old images already showed which part of the world each map should cover and which countries belonged together. The agent could figure those details out.

We briefly explored a mapping-service approach. The agent then found a route we could reproduce locally using public-domain data from Natural Earth: country boundaries for the geometry and shaded relief for a little texture.

Using a public-domain source also meant the finished almanac could avoid displaying the exact game maps, and avoid inadvertently reusing something I ought not to be. More importantly, the newly rendered maps would simply be flexible enough to look the way I preferred.

This became one of my favorite examples of backgroundmaxxing with an agent. I gave it several goals it could work through independently, including the complete map set, and went to sleep. When I came back the next morning, the work was finished!

The work could take all night, and it didn’t impact my effort at all.

In the final iteration my input was the original map set and the relationship I wanted to preserve. The agent worked out the analogous regions, coordinates, label placement, rendered files, and connections to the existing country data.

The extracted map beside the new Natural Earth version
The new map changed the visual system while preserving the regional view. The agent made the comparison image too, naturally!

The Region Had to Match

The new maps did not copy the old pixels, but they needed to remain recognizable as the same regional maps. The Southwest Asia map should still cover roughly the same part of the world. The Southern Europe map should still feel familiar when someone moved between the game and the almanac.

The agent generated bounding-box comparisons between each extracted map and its modern version. Those made it easy to check that a new map had not zoomed too tightly or drifted to a very different region.

It also validated the relationship between the country records and the shared maps. Every country assigned to one of the thirteen map groups had to appear on the corresponding modern map. The agent ran that check alongside the overnight map work.

I Woke Up to Thirteen Maps

The thirteen modern regional maps
The final set was rendered from Natural Earth data.

In the finished web almanac, countries in the same region reuse the corresponding modern map just as they did in the game. The art is new, but the relationship between the game and the almanac remains familiar.

I could open the almanac on a tablet and see another concrete part of the idea working. I had gone to sleep after describing the outcome and supplying the references. By morning, the agent had turned that direction into a complete set of maps.

Up Next: Realizing We Were Losing Nuance in Text Data

Maps could be checked with our eyes, though frankly, I took the agent at its word that every relevant country name appeared on each map! The recovered country text had a subtler issue. A rewritten paragraph could sound perfectly reasonable after the agent paraphrased and organized the different nuggets into a consistent hierarchy, yet omit crucial details in the process. One early rewrite of Greece did exactly that.

]]>