Matt's results

How well does Matt estimate the world?

The reasoning score rewards useful facts, a sound approach, correct math, and a final estimate close to the sourced answer. Higher is better.

Headline insight

Strongest area: Public systems and economics

Across 61 scored problems, the most reliable patterns are easier to see by subject and reasoning method than from one overall number.

Strongest established subjectPublic systems and economics11 problems · 72.3 average
Strongest established problem typeHuman time and attention3 problems · 70 average
Overall context60 / 100 reasoning score61 scored first passes
See detailed breakdown

Four parts. One reasoning score.

Useful facts + Reasoning approach + Math + Final estimate

Every reasoning score adds the points earned in four areas. Higher is better in every category.

Useful facts

Quality of remembered constants and domain anchors, sometimes called pegs.

11.3 / 30

Average earned points

Reasoning approach

Whether the conceptual breakdown answered the question.

22.9 / 30

Average earned points

Math

Correct arithmetic, units, and conversions.

6.6 / 10

Average earned points

Final estimate

Closeness of the first pass to the sourced answer.

19.3 / 30

Average earned points

Maximum points: 30 + 30 + 10 + 30 = 100. The same four components appear in each subject and model comparison below.

Published editorial grades, backfilled across the archive. Some older grades need a consistency review; these averages describe this set of problems, not a validated measure of ability. All charts use a 0-100 earned-point scale.

Every scored article is included in these aggregates. 61 of 63 published pages currently have a Reasoning score.

Scored articles 61 63 published problem pages scanned
Average reasoning score 60 100 is best, 0 is worst
Highest score (ties possible) 100 How much sea-level rise comes from 12.5 trillion tons of ice?
Lowest score (ties possible) 10 How far can Mayon volcanic ash travel before falling?

Reasoning score by issue

Higher is better. Issue sequence is a proxy for progress, not a verified attempt date. The unnumbered coffee issue is included in all averages, but omitted here. Changing topics and difficulty can change the trend.

Problem types

Performance by reasoning pattern.

This is the most useful view for learning transfer: the same model family can appear in politics, sports, science, infrastructure, and business news.

Dose, exposure, and health (n=1)
90
Counting, population, and volume (n=1)
80
Human flow and safety (n=2)
75
Human time and attention (n=3)
70
Unit conversion and scaling (n=3)
68.3
Geometry, area, and volume (n=5)
65
Stock-flow and throughput (n=11)
61.8
Logistics and material flow (n=8)
60.6
Energy, power, and physics (n=12)
58.8
Cost and economic scaling (n=7)
58.6
Probability and exponential scaling (n=4)
40
Motion, transport, and distance (n=3)
40
Scale comparison and infrastructure (n=1)
40
Problem typeArticlesAvg scoreUseful factsApproachMathFinal estimateSample note
Dose, exposure, and health 1 90 20 30 10 30 Small sample
Counting, population, and volume 1 80 20 30 10 20 Small sample
Human flow and safety 2 75 15 30 5 25 Small sample
Human time and attention 3 70 16.7 25 8.3 20 Descriptive only
Unit conversion and scaling 3 68.3 13.3 25 10 20 Descriptive only
Geometry, area, and volume 5 65 10 27 6 22 Descriptive only
Stock-flow and throughput 11 61.8 10.9 25.9 6.4 18.6 Descriptive only
Logistics and material flow 8 60.6 12.5 24.4 5 18.8 Descriptive only
Energy, power, and physics 12 58.8 10.8 20 7.9 20 Descriptive only
Cost and economic scaling 7 58.6 8.6 25.7 7.1 17.1 Descriptive only
Probability and exponential scaling 4 40 10 7.5 2.5 20 Descriptive only
Motion, transport, and distance 3 40 10 10 6.7 13.3 Descriptive only
Scale comparison and infrastructure 1 40 0 30 0 10 Small sample

Subjects

Performance by subject matter.

Each article has one primary subject and one model family. These editorial groupings can change; original labels remain in the export. Small groups are especially sensitive to which problems were attempted.

Everyday life and sports (n=2)
80
Public systems and economics (n=11)
72.3
Health, safety, and biology (n=4)
70
Engineering and robotics (n=1)
65
Earth, climate, and environment (n=14)
60.4
Energy and infrastructure (n=10)
57.5
Manufacturing and supply chains (n=5)
54
Space and astronomy (n=5)
53
Computing, media, and attention (n=5)
48
Transportation and aviation (n=4)
41.3
SubjectArticlesAvg scoreUseful factsApproachMathFinal estimateSample note
Everyday life and sports 2 80 20 30 10 20 Small sample
Public systems and economics 11 72.3 15.5 27.3 6.8 22.7 Descriptive only
Health, safety, and biology 4 70 12.5 22.5 7.5 27.5 Descriptive only
Engineering and robotics 1 65 10 15 10 30 Small sample
Earth, climate, and environment 14 60.4 12.1 22.5 5.7 20 Descriptive only
Energy and infrastructure 10 57.5 9 22.5 7 19 Descriptive only
Manufacturing and supply chains 5 54 6 27 6 15 Descriptive only
Space and astronomy 5 53 8 21 8 16 Descriptive only
Computing, media, and attention 5 48 8 21 5 14 Descriptive only
Transportation and aviation 4 41.3 12.5 11.3 5 12.5 Descriptive only

Your version

Your first estimates are tracked separately.

Your Results shows scale accuracy, subjects, problem types, and recent attempts. Optional passkey accounts can carry that history across devices.