Critical Thinking | Sid Savara / Engineering leadership, personal writing, and notes from Sid Savara Sat, 19 Sep 2026 19:55:06 +0000 en-US hourly 1 https://wordpress.org/?v=7.1.1 /wp-content/uploads/2026/05/sid-savara-favicon-s-512-100x100.png Critical Thinking | Sid Savara / 32 32 AI Solved the Problem I Pointed It At. I Had Aimed Too Narrowly. /ai-agent-data-coverage-blind-spots/ Mon, 07 Sep 2026 19:00:00 +0000 /ai-agent-data-coverage-blind-spots/
PART6OF 8

This is Part 6 of 8 in The Projects AI Makes Possible, a series about building a private almanac from a 1992 geography game with a coding agent. In Part 5: The Agent Built What I Asked, But Not What I Meant, the country data set and presentation became complete enough to review sentence by sentence. Playing the game, though, revealed that we still had gaps in clue coverage.

“The first principle is that you must not fool yourself—and you are the easiest person to fool.”

Richard Feynman, 1974 Caltech Commencement Address

In our initial tests, the core functionality worked. The almanac loaded our data and the appropriate maps, and search returned matching results. We had assumed that the country data contained everything needed to play the game. That assumption turned out to be incorrect.

Excited to test it for real, we loaded it up on a tablet and started the game.

During a case, we searched for a phrase the game had given us and found nothing. One of the early misses was paper mill. The source material connected Sweden’s forests with its pulp and paper industry, but none of that information appeared in a form the search could find.

The fix looked small. Add the missing phrase to Sweden and move on.

Then another search failed. And another.

Hari Rud was missing from Afghanistan. A Tamil Hindu temple phrase had been summarized too far. Mongolia needed the words yurt, snow leopards, and wild asses. Buddhist could fail even when the page contained Buddhism. An older spelling such as Szechuan disappeared when the prose used only Sichuan.

The search worked. The index worked. The country pages worked. The almanac was functioning correctly: it just didn’t actually have the data I needed!

Fixing One Wasn’t Enough

Each failed search told us what category of source material to audit next.

At first, the agent repaired the individual entries I found during play. That made the next search work, but it did not answer the more worrying question: what else had gone missing for the same reason?

There wasn’t a single cause behind the failures. They fell into several categories. I asked the agent to research and explain them to me, and their explanations revealed a few different causes.

Some phrases had been removed by an overaggressive summary. Some crossed source lines and made sense only when the neighboring records were read together. Some were historical names or spellings that modern prose naturally replaced. Others described a person’s appearance, hobby, or favorite food and did not honestly belong to any one country.

Each missing result suggested that we might have misclassified a whole group of records.

The Audit Kept Growing

I could give the agent one concrete miss and ask it to look for the pattern behind it. When paper mill was absent, the work expanded beyond Sweden to other country-specific fragments that had been filtered or summarized. When a phrase depended on the line before it, the agent widened the review to neighboring lines and multi-line groups. When a description did not identify a country, it audited the shared evidence instead of forcing the phrase into an arbitrary country page.

The review system began distinguishing among:

  • Country-specific material that needed an explanatory fact.
  • Context recovered from adjacent lines or a multi-line group.
  • Shared descriptions that needed somewhere else to live.
  • Historical or dated language that should remain searchable with context.
  • Edge cases where aggressively pulling in all the data also captured unusable fragments, which we then trimmed rather than turning them into almanac entries.

The agent would scan all our records, group likely failures, generate review tables, and rerun the coverage checks. My contribution was often much smaller and more specific: this phrase appeared in the game, this search result was missing or unhelpful.

Over time, the agent got better at taking one failure and expanding the audit before returning with another supposedly complete pass. The loop grew beyond one missing phrase followed by one patch. A concrete failure became a reason to reconsider the broader approach and fix an entire group of missing clues, including ones we hadn’t explicitly identified.

The Result Had to Teach

We could have made the search return something by dumping the original fragments into, say, a keyword field. The almanac still would not have explained what the phrase meant.

When I searched for Hari Rud, I wanted the result to explain that it is a river flowing west from central Afghanistan past Herat toward Turkmenistan. A page that merely calls it “a geography reference for Afghanistan” technically matches the phrase, but provides very little context or education.

The same applied to paper mill. The explanation needed to connect the phrase to Sweden’s forests and its pulp-and-paper industry while keeping the exact words a player might type.

The Sweden page opened to the paper-industry fact
Opening the result took us to the country context behind the phrase.

Not every phrase could be handled that way. Some clues were specific to helping identify a suspect: ebony-colored hair, malachite-colored eyes. In the game such clues help describe a suspect to allow the player to obtain a warrant for the suspect’s arrest. While I could have finagled it (e.g. put malachite in a country that happens to have malachite), those words were not really attached to any country.

Those descriptions became a separate collection of World Facts. They could remain searchable and understandable without pretending to be country facts, and still be educational.

Every Record Accounted For

Coverage was not complete until every source record had a destination or a reason for exclusion.

The final audit accounted for 5,441 distinct source records. That does not mean the site contains 5,441 polished facts. Some records repeat, some contribute only part of an explanation, and some are structural material rather than content.

Of those records, 4,674 routed to country-specific explanations. Another 217 shared descriptions routed to World Facts. The remaining 550 were records with an explicit reason not to appear as almanac entries. None remained unresolved in the audit.

When I started playing with what I thought was the “complete” almanac, I expected to focus on ease of use and friction. Playing the game also smoke tested our clue coverage, exposing gaps I didn’t even realize we had missed. One failed search could have led to one repaired entry. Instead, it exposed a pattern, and fixing it consistently led to a broader review of the set of clues.

After a few iterations – naturally, while I slept or did other things – the almanac had a much more complete set of facts. The next time we played, search behaved the way we had wanted from the beginning. We could hear an unfamiliar phrase, find it on the tablet, understand why it mattered, and continue the case without leaving the experience.

]]>
The Curse of the Worst Acceptable Solution /the-curse-of-the-worst-acceptable-solution/ Mon, 15 Dec 2025 19:00:00 +0000 /?p=91

“Your life reflects what you tolerate.”

Tony Robbins

When I first moved and got my car, I missed the aux input from my old stereo. With an aux input, I could plug in my phone, listen to podcasts, and make my commute feel a little more useful. I figured that once I was settled, I would replace the stereo.

Until then, I would make do.

The First Workaround

At first, making do was fine. I listened to the radio and old CDs. When I wanted to listen to podcasts, I shifted that habit to running instead.

That sort of solved my problem – sure I still had podcast listening time, but not during my drive. It also meant all the music and podcasts on my phone were not available in the car, just the same rotation of CDs.

The workaround was acceptable enough…and that’s exactly what I did, I accepted it.

The Better Bad Solution

Later, my dad gave me an FM transmitter. The sound quality was not good, but it was fine. However, it did let me listen to podcasts in the car again, so I accepted it.

The signal dropped sometimes. The volume was low. The experience was not good. But it crossed the threshold of tolerable, so I kept it.

Eventually I found a way to hang the transmitter cable from the visor, which improved the signal a bit. It looked weird, was a little inconvenient but the sound was better than before – though still worse than the real fix.

However that is how it remained, until years later when I purchased a new car that had an aux input and I hooked up a Bluetooth receiver. Finally – good quality, wireless, and easy for my friends to connect their music as well.

The Worst Acceptable Solution

Why didn’t I simply buy a new car stereo? Why did I let things continue the way they were?

Annoying but workable is where bad systems learn to survive.

Because – the system was not broken. If it were broken, I probably would have fixed it. But it was the worst acceptable solution: annoying, but workable enough.

So I adapted around it, accepted it – and just lived with it.

At Work, It Looks Familiar

In my professional life, I’ve seen the same thing – in software, often tech debt. Worse: tech debt that is continually built upon because the small annoyances and inefficiencies slow things down, but don’t stop enough impactful work.

Meanwhile, we know that the tech debt needs to be addressed – though unlike in my situation, unaddressed tech debt gets worse and worse.

Product quality with weird workarounds is another example. Small quality issues become accepted, and can lead to a new, lower normal.

That is the curse: once a solution is barely acceptable, it can become surprisingly hard to replace with a good one.

]]>
The Four Types of Action /the-four-types-of-action/ Mon, 29 Jan 2001 12:00:00 +0000 /?p=3166

“Action expresses priorities.”

Mahatma Gandhi

I found an old journal entry from a day when I knew I had not used my time well.

Nothing dramatic had happened. I had answered things, clicked through things, handled a few small obligations, and arrived at the end of the day with the uncomfortable feeling that I had been busy without really choosing much.

That is the part that stayed with me. A day can look active from the outside and still feel passive from the inside.

When I looked back at that entry, the actions of the day seemed to fall into four rough categories: inaction, distraction, reaction, and proaction. The names are a little blunt, but they still help me notice what kind of day I am actually having.

Four Ways a Day Can Go

Inaction is the easiest one to recognize after the fact. I did not move the thing forward. I did not make the call, write the page, clean up the decision, or start the work. Sometimes there is a good reason. Sometimes I just avoided the friction.

Distraction is trickier because it often feels like movement. I am reading, checking, organizing, following a thread, adjusting a system, or doing something plausibly useful. The problem is not that the activity is always bad. The problem is that it has captured attention that I meant to spend somewhere else.

Reaction is the mode I can justify most easily. Something came in, and I responded. An email, a question, a small problem, a meeting, a request. Much of life and work requires response, so I do not think reaction is automatically wrong. But if the whole day is reaction, then other people and incoming events have effectively written the day for me.

Proaction is the work I chose before the day chose for me. It is the thing I decided mattered enough to begin, protect, or move forward. It does not have to be grand. Some days it is one useful conversation, one hard paragraph, one clarified decision, one unglamorous task that keeps a larger commitment alive.

What I Know

I know that I am rarely confused at the end of a day about which category dominated it.

I might have explanations. I might have good reasons. I might have been genuinely pulled into something urgent. But usually, if I pause for a minute, I can tell whether I spent the day moving something important forward or mostly orbiting around it.

What part of today was chosen?

That is the useful question for me now: what part of today was chosen?

Not every hour can be proactive. Not every day can be protected. But I do want some part of the day to answer to the things I say matter, instead of only to the things that arrived loudest.

]]>