Experimentation
Learning log
Learning log is licensed CC BY 4.0. Attribution: Katafacts (katafacts.com).
Customise with AISkill file ↓1 · What it is
What it is
A learning log is a real, checkable answer to 'have we actually tested that, or are we just assuming it' — a fixed set of hypotheses, and for each one, the real tests actually run against it (or the lack of any). It's the same 'management by fact, not opinion' discipline an evidence register runs on claims already circulating, pointed instead at ideas waiting to be tried: a hypothesis with no test logged is named directly as untested, and a test's own recommendation — keep, revise, or kill — is recorded honestly rather than treated as a formality once a test happens to run.
2 · When to use it
When to use it — and when not to
Use it when
- Several hypotheses or ideas are circulating about what would actually improve something, and you want an honest record of which have real tests behind them.
- A test has already run and produced a real result, and you want to record what it actually recommends — keep, revise, or kill — rather than letting a promising-sounding result stand in for a clear verdict.
- You want to know which hypotheses have never been tested at all, not just which ones feel intuitively right.
Not when
- You don't have a real hypothesis to test yet — this log tracks tests against hypotheses already named, it doesn't generate new ones.
- The idea is genuinely trivial or reversible with no real cost to just trying it — the overhead of a formal test isn't worth it yet.
3 · How to fill it in
How to fill it in
- Scope
- What set of hypotheses does this log cover?
- Hypotheses
- List every hypothesis or idea you want tracked, even ones you haven't tested yet.
- Tests run
- For each test: which hypothesis it relates to, the test type, what actually happened, and your real recommendation — keep, revise, or kill — your own facts, never guessed.
- Narrative
- Drafted from the log's own computed coverage.
4 · What good looks like
What good looks like
The example below tests two hypotheses already circulating about Line 2's changeover performance — including the exact sequencing-rule claim the attribution reality check found unverified and the evidence register logged as unproven — with a third hypothesis honestly left untested, continuing the evidence register's own flagged gap.
Same example, as a downloadable xlsx workbook.
Download .xlsxLearning log
Line 2 — learning log
Tests hypotheses about what's actually driving Line 2's changeover performance, instead of letting claims already circulating stand as assumptions.
Priya Nair, Line Supervisor · 2026-04-08
Scope
Tests hypotheses about what's actually driving Line 2's changeover performance, instead of letting claims already circulating stand as assumptions.
Hypotheses
- The sequencing rule causes the on-time shipment recovery
- Redesigning the nozzle cleaning step reduces changeover time
- Queue board update discipline affects changeover variabilityNever tested
Tests run
The sequencing rule causes the on-time shipment recovery
keepSuspended the sequencing rule for one week on second shift only, kept it active on first shift, and compared on-time shipment rates across both — a real, if imperfect, controlled comparison rather than the uncontrolled before/after the attribution reality check already flagged as confounded
Second shift's on-time rate dropped to 81% during the suspension week, down from a recent average of 94%, while first shift held steady at 93%
experiment · Priya Nair, Line Supervisor — shift log comparison · 2026-04-02
Redesigning the nozzle cleaning step reduces changeover time
revisePiloted a redesigned cleaning fixture on 8 changeovers across two weeks
Average changeover time on the piloted changeovers was 41 minutes, versus the 45-minute baseline — a real improvement, but the pilot covered only 8 changeovers and one operator, not enough yet to rule out operator skill as the actual driver
pilot · Marcus Webb, Paint Booth Lead · 2026-04-05
Narrative
The sequencing rule's effect finally has a real test behind it: suspending it for a week on second shift only, while leaving first shift unchanged, showed a real drop when it was off — evidence the rule is doing real work, not just coinciding with a recovery the attribution reality check already found unverified. The nozzle-cleaning redesign shows a promising 41-minute average, but the pilot is too small and single-operator to rule out skill as the real driver — worth a larger, more controlled pilot before committing, not a straight pass. The third hypothesis, whether queue board update discipline affects changeover variability, hasn't been tested at all yet — the same gap the evidence register already flagged with zero evidence logged, now visible here as an idea nobody's actually run a real test against either.
5 · Common mistakes
Common mistakes
Treating a promising-sounding result as an automatic 'keep' without a real recommendation.
A result can be real and still be too small, too narrow, or too confounded to fully validate a hypothesis — the recommendation should name that honestly (often 'revise') rather than rounding a partial result up to a clean pass.
Leaving a hypothesis off the log because you're confident it hasn't been tested.
An unlisted hypothesis can't be flagged as untested — naming it and letting it show up with zero tests is exactly how this tool surfaces the gap plainly instead of hiding it by omission.
Running a test but never recording what it actually recommends.
A test with a result but no recommendation leaves the real decision — keep, revise, or kill — unmade, which is the entire point of running the test in the first place.
6 · What it connects to
What it connects to
upstream
Evidence register
A claim the evidence register logs as unverified is often exactly the hypothesis this log should go test properly, instead of leaving it to stand as an assumption.
downstream
A3 problem solving
A test that recommends 'kill' on a countermeasure hypothesis is a direct signal an A3's own effect-confirmation plan should act on, not quietly ignore.
7 · Where AI helps
Where AI helps
Judgement — stays yours
- Deciding whether a result actually supports keep, revise, or kill
- Deciding which untested hypothesis to prioritize next
Analysis — AI helps
- Tightening a test description from rough notes into a specific, checkable statement
- Drafting the narrative from the log's own computed coverage
Drudgery — automated
- Identifying which hypotheses have no tests logged
- Identifying which logged tests recommend revise or kill
- Exporting to xlsx in the house format
9 · Rate this kata
