Seven of my eight cleverest ideas were measuring nothing
Every one had a good name and none of them could change a decision, so they came out. Then I measured what was left against the market, and lost.
Over the summer I built a side project that forecasts game outcomes for a sport I follow. It is not for sale and never will be. I built it to practice the one discipline that matters most in my real work and is hardest to sell: measuring whether something you built actually does anything, and acting on the answer.
By the third version it had eight extra factors layered on the base model, each with a good name: game pace, clutch performance, scheme shock, coaching aggression, referee style, crowd pressure, weather risk, and line movement. Every one sounded like something a careful forecaster would account for, and every one made the readout look more sophisticated.
What the measurement said
Then I replayed the previous season with each factor switched on and off. Seven of the eight changed no prediction at all, not one. When I read the code to find out why, the answer was embarrassing: each of them was computed from a stand-in, a win-loss record or a keyword in a report, rather than from the thing its name described. Three were quietly counting inputs the base model already had. They were labels, not factors.
So the fourth version removed seven of them. That is the whole story, and it is the part most people skip. Adding the factors took a weekend and felt like progress. Deleting them took an hour and felt like loss. Only one of the two made the model better, and it was not the weekend.
The rule that made the deletions honest
None of that measuring means anything unless the record it runs against can be trusted, and this is where the project stopped being a toy. Every forecast is written down before the game and frozen at kickoff. If the run that should have written it did not happen in time, the slot is marked missed and stays empty. It is never filled in afterwards, even though by then the correct answer is sitting right there.
This rule does not come from web work. Nothing on a marketing site asks for it, which is why almost no website has it. It comes from trading systems, where a number written down after the close is not a late number, it is a false one, and where a record that can be quietly amended is treated as no record at all. I brought it with me because I do not know how to trust a record without it.
The number it is not allowed to hide
The base model is also measured against the market's own published number, which is the best forecaster in the room by a wide margin. On last season's replay the market picked the winner 65 percent of the time, my model on its own picked 61, and the two blended picked 66 with slightly worse confidence. On the games where the two disagreed, my side was right 38 percent of the time. That is not an edge, and the app says so on its own front page. It shows the model as a second opinion and never lets it override the number that is actually good.
What this has to do with your website
- Every website has features with good names. Measure whether they do anything before you believe the name.
- If a check or a factor cannot change a decision, it is decoration. Remove it, even when it was expensive to build.
- A record you can edit after the fact is not a record. Freeze what you claimed before you learn the result.
- Compare yourself to the strongest baseline, not the weakest, and publish the comparison even when you lose.
I would rather show a project that lost to the market and said so than one that won on paper. That is the kind of measuring your site gets too, and it is why the findings come in writing.
Narcisa Budworth
Founder of Minarvo. Formerly a senior engineering leader at PEAK6 and Symphony, now building digital platforms for small and mid-sized businesses from Chicago.
Sound like your website? A free 30-minute call will tell you where you stand.
Request a call