Skip to content
KatafactsBeta

Experimentation

Learning log

Learning log is licensed CC BY 4.0. Attribution: Katafacts (katafacts.com).

Customise with AISkill file ↓

1 · What it is

What it is

A learning log is a real, checkable answer to 'have we actually tested that, or are we just assuming it' — a fixed set of hypotheses, and for each one, the real tests actually run against it (or the lack of any). It's the same 'management by fact, not opinion' discipline an evidence register runs on claims already circulating, pointed instead at ideas waiting to be tried: a hypothesis with no test logged is named directly as untested, and a test's own recommendation — keep, revise, or kill — is recorded honestly rather than treated as a formality once a test happens to run.

2 · When to use it

When to use it — and when not to

Use it when

  • Several hypotheses or ideas are circulating about what would actually improve something, and you want an honest record of which have real tests behind them.
  • A test has already run and produced a real result, and you want to record what it actually recommends — keep, revise, or kill — rather than letting a promising-sounding result stand in for a clear verdict.
  • You want to know which hypotheses have never been tested at all, not just which ones feel intuitively right.

Not when

  • You don't have a real hypothesis to test yet — this log tracks tests against hypotheses already named, it doesn't generate new ones.
  • The idea is genuinely trivial or reversible with no real cost to just trying it — the overhead of a formal test isn't worth it yet.

3 · How to fill it in

How to fill it in

Scope
What set of hypotheses does this log cover?
Hypotheses
List every hypothesis or idea you want tracked, even ones you haven't tested yet.
Tests run
For each test: which hypothesis it relates to, the test type, what actually happened, and your real recommendation — keep, revise, or kill — your own facts, never guessed.
Narrative
Drafted from the log's own computed coverage.

4 · What good looks like

What good looks like

The example below tests two hypotheses already circulating about Line 2's changeover performance — including the exact sequencing-rule claim the attribution reality check found unverified and the evidence register logged as unproven — with a third hypothesis honestly left untested, continuing the evidence register's own flagged gap.

Same example, as a downloadable xlsx workbook.

Download .xlsx

Learning log

Line 2 — learning log

Tests hypotheses about what's actually driving Line 2's changeover performance, instead of letting claims already circulating stand as assumptions.

Priya Nair, Line Supervisor · 2026-04-08

Scope

Tests hypotheses about what's actually driving Line 2's changeover performance, instead of letting claims already circulating stand as assumptions.

Hypotheses

  • The sequencing rule causes the on-time shipment recovery
  • Redesigning the nozzle cleaning step reduces changeover time
  • Queue board update discipline affects changeover variabilityNever tested

Tests run

The sequencing rule causes the on-time shipment recovery

keep

Suspended the sequencing rule for one week on second shift only, kept it active on first shift, and compared on-time shipment rates across both — a real, if imperfect, controlled comparison rather than the uncontrolled before/after the attribution reality check already flagged as confounded

Second shift's on-time rate dropped to 81% during the suspension week, down from a recent average of 94%, while first shift held steady at 93%

experiment · Priya Nair, Line Supervisor — shift log comparison · 2026-04-02

Redesigning the nozzle cleaning step reduces changeover time

revise

Piloted a redesigned cleaning fixture on 8 changeovers across two weeks

Average changeover time on the piloted changeovers was 41 minutes, versus the 45-minute baseline — a real improvement, but the pilot covered only 8 changeovers and one operator, not enough yet to rule out operator skill as the actual driver

pilot · Marcus Webb, Paint Booth Lead · 2026-04-05

Narrative

The sequencing rule's effect finally has a real test behind it: suspending it for a week on second shift only, while leaving first shift unchanged, showed a real drop when it was off — evidence the rule is doing real work, not just coinciding with a recovery the attribution reality check already found unverified. The nozzle-cleaning redesign shows a promising 41-minute average, but the pilot is too small and single-operator to rule out skill as the real driver — worth a larger, more controlled pilot before committing, not a straight pass. The third hypothesis, whether queue board update discipline affects changeover variability, hasn't been tested at all yet — the same gap the evidence register already flagged with zero evidence logged, now visible here as an idea nobody's actually run a real test against either.

5 · Common mistakes

Common mistakes

  • Treating a promising-sounding result as an automatic 'keep' without a real recommendation.

    A result can be real and still be too small, too narrow, or too confounded to fully validate a hypothesis — the recommendation should name that honestly (often 'revise') rather than rounding a partial result up to a clean pass.

  • Leaving a hypothesis off the log because you're confident it hasn't been tested.

    An unlisted hypothesis can't be flagged as untested — naming it and letting it show up with zero tests is exactly how this tool surfaces the gap plainly instead of hiding it by omission.

  • Running a test but never recording what it actually recommends.

    A test with a result but no recommendation leaves the real decision — keep, revise, or kill — unmade, which is the entire point of running the test in the first place.

6 · What it connects to

What it connects to

upstream

  • Evidence register

    A claim the evidence register logs as unverified is often exactly the hypothesis this log should go test properly, instead of leaving it to stand as an assumption.

downstream

  • A3 problem solving

    A test that recommends 'kill' on a countermeasure hypothesis is a direct signal an A3's own effect-confirmation plan should act on, not quietly ignore.

7 · Where AI helps

Where AI helps

Judgement — stays yours

  • Deciding whether a result actually supports keep, revise, or kill
  • Deciding which untested hypothesis to prioritize next

Analysis — AI helps

  • Tightening a test description from rough notes into a specific, checkable statement
  • Drafting the narrative from the log's own computed coverage

Drudgery — automated

  • Identifying which hypotheses have no tests logged
  • Identifying which logged tests recommend revise or kill
  • Exporting to xlsx in the house format

9 · Rate this kata

Rate this kata