← AI lab

Design evidence skill

Standardise the evidence, not the story.

Project
Design Evidence
Type
Self-initiated, AI lab
Role
Concept, workflow, safeguards, evaluation
Year
2026
Status
Internally evaluated prototype
  1. Layer 01 The reusable skill

    Instructions and blank templates. No project lives here.

    Shareable
  2. Layer 02 The private ledger

    Real evidence, in a location the designer has approved.

    Never published
  3. Layer 03 The public portfolio

    Only what a person has reviewed and qualified.

    Reviewed by me
Three storage layers, and the boundaries between them

TLDR

I built a Claude skill that records design decisions and their evidence status before hindsight rewrites them. Against the same model without it, it caught every privacy risk in testing, 6 of 6 against 3 of 6. One run each, graded by me, on fictional material.

01

Keeping a trustworthy record

The skill records decisions before hindsight turns uncertainty into outcomes.

Rough notes become a private record of starting conditions, rejected options, contributions and evidence gaps. Six modes cover starting points, weekly capture, milestones, audits, stories and reconstruction. That record can support different formats later. The reusable instructions hold no project evidence. Real material belongs in an approved private ledger, and only reviewed material crosses into a public portfolio.

  • Mode 01 Capture the starting point
  • Mode 02 Record the week
  • Mode 03 Reconstruct a milestone
  • Mode 04 Audit the evidence
  • Mode 05 Build a story
  • Mode 06 Rebuild from old notes

02

Keeping judgement human

Claims keep their evidence status, and publication stays my decision.

Verified, observed, reported, expected and unknown remain distinct. Individual contribution stays separate from team delivery. STAR checks completeness without forcing every story into one template. The skill cannot approve publication or guarantee anonymisation. I rejected automatic collection from workplace tools because it removes the moment of consent. Storage and source selection need approval before capture, rather than a redaction pass afterwards.

Verified Observed Reported Expected Unknown

Walk through all six steps of the Project Tide record

Fictional demonstration

Project Tide, a household collection service

An invented product, team and research round, written to show the workflow. Six steps, in order, from a week of rough notes to a record with its gaps still visible.

  1. Step 01

    What I bring

    Notes from a week redesigning a booking flow, written for me and nobody else. Research, a team decision, my own recommendation and an early prototype signal, all in one pile.

    • Team thought people abandoned the flow because it asked for too many fields.
    • Five moderated prototype sessions on Tuesday. Four hesitated at the collection category step.
    • Three were unsure whether a small electrical item counted as household or specialist.
    • Planned the sessions with the researcher, moderated two, synthesised the category findings.
    • Explored three directions: fewer categories, examples under every category, item before category.
    • Engineer says item-first needs a new classification service and probably misses the pilot.
    • Recommended keeping the category model, renaming two labels, examples only where people hesitated.
    • Prototype analytics: completion 6 of 10 last round, 8 of 10 this round. Same task? Not sure.
    • Still need the final label decision, the next test plan, and who owns pilot measurement.
  2. Step 02

    What it notices

    Four different kinds of thing are sitting in one list, and they do not carry the same weight. Before writing anything it sorts them, and it classifies the material as fictional so it knows which storage rules apply.

    • Research signal

      Four of five hesitated at the category step. Observed, in five sessions.

      Observed
    • Team decision

      Product and operations agreed to test the recommendation next.

      Observed
    • My recommendation

      Keep the category model for the pilot, rename two labels, add examples where people hesitated.

      Expected value
    • Prototype signal

      6 of 10 to 8 of 10, from two rounds of unknown comparability.

      Held back
  3. Step 03

    What it separates and qualifies

    Contribution first, because hindsight blurs it fastest. It puts what I did next to what everyone else did.

    Work Me The team
    Research planning Co-planned Researcher co-planned
    Moderation 2 of 5 Researcher, 3 of 5
    Synthesis Led the category synthesis Researcher reviewed it
    Design response Explored and recommended Content designer co-created labels, engineer advised on feasibility
    Outcome None owned or verified Measurement owner unknown

    Then the completion figures. They are kept, in raw counts, next to the reason they cannot be used: the rounds may not be comparable, and prototype completion is not production data either way. The number stays, labelled.

    6 of 10, then 8 of 10 Comparability unknown Not production data

  4. Step 04

    Where I am still needed

    None of that is approval. The decisions that stay mine are the ones that matter outside the document.

    • Whether this capture is accurate, and whether it may be saved at all.
    • Whether the private destination is the approved one for this material.
    • Whether “led the synthesis” is a claim the researcher would recognise.
    • Whether the completion figures should appear anywhere public, in any form.
    • Whether a fictional reconstruction should replace the real example entirely.
  5. Step 05

    The record it proposes

    A milestone entry, offered for approval rather than saved. In the real workflow it would name the private destination and wait.

    Fictional record

    Category comprehension reframes the problem

    Not saved yet
    Situation

    The team believed the booking flow lost people because it asked for too many fields. Five moderated sessions pointed somewhere else: four participants hesitated at the collection category, and three could not tell where a small electrical item belonged.

    Decision and rationale

    Keep the existing category model for the pilot, rename two labels with the content designer, and add examples only where research showed hesitation. It targets the observed comprehension problem without taking on a classification service that risks the pilot date.

    Outcome status
    • Research signal: observed
    • Team agreement to test it: observed
    • Prototype completion difference: observed, not comparable
    • Improvement to pilot or production completion: unknown
    Evidence horizon, three actions
    • Record the final label decision and its rationale before the prototype changes again.
    • Confirm whether the two test rounds used comparable participants and tasks, before quoting either figure.
    • Agree who owns pilot measurement, and what signal they will collect, before the pilot starts.

    I wouldn’t have written those three actions myself. They’re ordered by how soon the evidence disappears.

  6. Step 06

    What is still unknown

    There’s no outcome yet, and the record says so. At this point a case study could claim a reframing and a decision, and that’s real work.

    • Do the revised labels improve completion in the pilot?
    • Were the two prototype rounds comparable at all?
    • Who owns pilot measurement, and what will they collect?

    Six months later, that record is the difference between remembering that something worked and being able to say what was actually known at the time.

03

What the comparison showed

My fictional tests found a privacy gap, not a better storytelling result.

I wrote the rubric before testing three tasks with and without the skill on the same model. Scores were 20/20 and 16/20; portfolio structure tied. Each evaluation ran once and I graded it myself, so reliability remains unknown. I subsequently prioritised missing evidence by urgency and value, and shortened the trigger description. Neither change proves that real records improve.

  • 20of 20, with the skill
  • 16of 20, without it
  • 3fictional evaluations
  • 1run each, so far
Behavioural checks passed, with the skill and without it
Evaluation With the skill Baseline
Privacy capture 6 of 6 3 of 6
Portfolio structure 7 of 7 7 of 7
Evidence audit 7 of 7 6 of 7
Total 20 of 20 16 of 20

One run per evaluation, graded by me, on fictional material.

Behavioural evaluation scores by task, with and without the skill

04

Taking it to real evidence

Next I would run it on my own milestones, then put it in front of other designers.

Two things matter most: whether what it asks feels proportionate, and whether anyone keeps recording once the novelty wears off. I haven’t yet checked how it behaves on other models. Project Tide above is fictional and fixed, with no model call, and it keeps completion counts qualified because the rounds may not be comparable.