How to score your own predictions honestly
To score predictions honestly, write them so they can be checked, add how sure you are, seal them before the outcome is known and grade each one with three marks: Called it, Kind of, Missed. Count one, a half and zero. The sealing step matters most, because it stops hindsight from rewriting what you thought.
Coming soon to iPhone
Why score predictions at all?
People remember being right more often than they were. Once an outcome is known it feels obvious, a bias called hindsight bias. A written record from before the event is the cure. Scoring turns "I knew it" into evidence.
What does hindsight bias look like in a real study?
In the first experiment built to test it, Beyth and Fischhoff asked people to rate the likelihood of several outcomes of Richard Nixon's coming visits to Beijing and Moscow. After the trip, the same people were asked to recall their earlier estimates. According to the Wikipedia summary, they remembered higher odds for the events that had actually happened.
Nobody in that study was lying. Memory simply moved toward the result. A sealed list is the only fix that does not depend on memory.
How do I write a checkable guess?
A checkable guess names something that will be clearly true or false by a date. Compare "Work will be better" with "I will have a new job title by December". Include numbers, names and places where you can. Make one claim per row, so you can grade it without arguing with yourself.
Which wording can be graded and which cannot?
A guess can be graded when a stranger could look at the facts on the date and agree on the mark. Vague guesses fail that test, so rewrite them before you seal.
- Does it name a number, a place, a person or an object?
- Is there one claim, not two joined by "and"?
- Could a photo, receipt or message settle it?
- Will the answer be known by the open date?
| Too vague | Checkable |
|---|---|
| I will be healthier | I will run 5 km without stopping |
| Work will change | I will have a new job title by December |
| The trip will be great | We will take the trip to Porto in May |
| Prices will go up | A coffee near my flat will cost more than it does today |
| I will see my friends more | I will meet Dana at least once a month |
Which kinds of guesses belong in a list?
Mix four kinds. Sealday's starters come in four decks for this reason: Life, People, Around me and Silly. Life guesses test your plans. People guesses test how well you read the people close to you. Around me guesses, such as prices, a trend or the word of the year, are the easiest to grade because anyone can look them up. Silly guesses make the opening fun.
Two or three from each deck makes a balanced list of eight to ten. The starters stay away from health, money outcomes and other people's private matters, and your own list should too.
What about confidence?
Add how sure you are. Sealday offers three tags, Sure, Maybe and Wild guess. Over several capsules you can notice that your "Sure" guesses come true less often than you thought. Forecasters study this formally; see superforecasters on Wikipedia for the research that popularized it.
What is calibration?
Calibration means your confidence matches your results: guesses you call Sure should come true far more often than Wild guesses. People usually fall short. The overconfidence effect describes how confidence tends to run ahead of accuracy, especially on hard questions about unfamiliar topics.
You do not need statistics to see it in your own list. After grading, count how many Sure guesses came true, then how many Maybe and Wild guess ones did. If the three rates are close together, the tags are not telling you anything yet.
What are the three marks?
Use the same marks every time.
| Mark | Meaning | Points |
|---|---|---|
| Called it | It happened as written | 1 |
| Kind of | Close, or partly true | 0.5 |
| Missed | It did not happen | 0 |
Is there a more exact method?
Yes. If you give each guess a probability, such as 70 percent, the Brier score measures accuracy as the average squared difference between your probabilities and what happened, where lower is better. It is overkill for a party game and useful for a yearly review.
How do I work out a Brier score by hand?
Write each guess as a probability, mark the outcome as 1 if it happened or 0 if not, subtract, square the difference and average the results. The Brier score was proposed by Glenn W. Brier in 1950, and in its usual form it runs from 0 to 1, with lower being better.
- Add the squared errors: 0.01 + 0.49 + 0.16 = 0.66.
- Divide by the number of guesses: 0.66 / 3 = 0.22.
- Compare with a baseline: answering 50% to everything scores 0.25, so 0.22 is slightly better than no knowledge at all.
| Guess | You said | It happened? | Squared error |
|---|---|---|---|
| New job title | 90% | Yes (1) | 0.01 |
| Live in Lisbon | 70% | No (0) | 0.49 |
| Half marathon done | 60% | Yes (1) | 0.16 |
How do I avoid grading myself too kindly?
Grade before you read your letter. Judge each guess on the words you wrote, not the intent you remember. Ask a friend for the close calls. In Sealday you can add a note to a mark, and a miss appears in neutral grey, so there is no shame in it. See predictions for the full flow.
What does a scored list look like?
A short example with confidence tags and points.
| Guess | Tag | Mark | Points |
|---|---|---|---|
| I will have a new job title | Sure | Called it | 1 |
| I will live in Lisbon | Wild guess | Missed | 0 |
| I will run a half marathon | Maybe | Kind of | 0.5 |
| I will delete one social app | Sure | Called it | 1 |
| The group chat will be renamed | Maybe | Called it | 1 |
How do I check my calibration after a few rounds?
Here the tags are ordered correctly, but Sure at 60% is overconfident. Next year, call a guess Sure only when you would bet most of a coffee on it. A gap like this is the useful result of keeping score.
| Tag | Guesses | Came true | Hit rate |
|---|---|---|---|
| Sure | 10 | 6 | 60% |
| Maybe | 12 | 5 | 42% |
| Wild guess | 8 | 1 | 13% |
What do I do with the result?
Add up the points: this list scores 3.5 out of 5. Then look at the tags. If your "Sure" guesses scored lower than your "Maybe" ones, you were overconfident. Carry misses into next year's list in a different form, and keep the ones you find funny.
How often should I do this?
Once a year is enough, and New Year or your birthday gives you an easy date. Rolling capsules of three to five guesses let you see patterns after two or three rounds. Sealday supports both, and the yearly ritual does not count toward the free capsule limit.
How do I compare one year with the next?
Keep one line per year in a note and compare the same three numbers each time. One year of data is a story. Three years is a pattern, and patterns are where you learn something about how you think: whether you overrate your plans, underrate your friends, or guess the world around you better than your own life.
- Your total, such as 3.5 out of 5.
- The hit rate of your Sure guesses.
- The one guess that surprised you most, in either direction.
What mistakes spoil a score?
- Grading only the hits and forgetting the misses.
- Changing the wording of a guess after you know the outcome.
- Filling the list with safe guesses, such as "I will still use my phone".
- Predicting private things about other people.
- Counting a guess as Kind of because the idea behind it came true, when the words did not.
- Waiting years to grade. Grade on open day, while you remember the context, because a grade given three weeks later is already shaped by what happened since.
How does Sealday run the whole cycle?
The app follows the same steps, with grading built into open day.
To write guesses worth scoring, read how to get better at making predictions, and see what hindsight bias is for the reason a sealed record helps.
- Pick the open date, then write up to 10 guesses on the free plan, or up to 30 with Plus. Each can carry a Sure, Maybe or Wild guess tag.
- Hold the seal for 2.4 seconds. From then on the list cannot be edited.
- On open day, hold again to break the seal. Each guess appears on its own card.
- Mark each one Called it, Kind of or Missed, with an optional note.
- Read your score, such as "You called 3 of 5", and post the scorecard with any guess blurred.
- Export the opened capsule if you want a text copy: it saves guesses with their marks in a file.
Frequently asked questions
What is a good score?
There is no benchmark for personal guesses. Track your own score across years and watch the trend.
Should I grade partial hits as correct?
Use Kind of and count half. That keeps the full mark honest.
How many predictions should I make?
Five to ten per capsule is enough to be interesting and quick to grade.
Can I hide a missed guess when sharing?
In Sealday, yes. Every guess has a toggle to blur it on the scorecard.
What is the difference between a score and calibration?
A score counts how many guesses came true. Calibration compares that count with how sure you said you were.
Do I need probabilities to score predictions?
No. Three marks and three confidence tags are enough for a yearly review. Probabilities and the Brier score add precision.