← AI lab

Standardise the evidence, not the story.

A privacy-first AI skill that keeps a trustworthy record of design work while the details are still true, then helps turn that record into whichever story is needed later.

Project
Design Evidence
Type
Self-initiated, AI lab
Role
Concept, workflow, safeguards, evaluation
Year
2026
Status
Internally evaluated prototype
  1. Layer 01 The reusable skill

    Instructions, the evidence model and blank templates. Nothing about any actual project lives here.

    Shareable
  2. Layer 02 The private ledger

    Real evidence from real work, in a location the designer has approved. It never travels upward on its own.

    Never published
  3. Layer 03 The public portfolio

    Only material that has been reviewed, sanitised and qualified. A person decides what crosses this line.

    Reviewed by me
Three layers, and the boundaries between them are the product Most of the risk in a tool like this is storage, not writing. So the separation is designed first: the repository holds no workplace evidence, the ledger holds no public copy, and nothing moves between them without a person deciding it should.

TLDR

Nobody maintains their portfolio until they need it, and by then the evidence has scattered across design files, tickets, chat and memory. The screens survive. The reasoning around them does not.

So I built an AI skill that captures decisions, contribution and evidence status while the project is running, keeps that record private, then helps shape it into a case study or an interview answer. Graded against the same model without it, the whole gap was privacy.

  • 20of 20 behavioural checks
  • 16of 20 without the skill

One run each, graded by me. No other designer has used it and it has never met my own real evidence, so this is a forward test rather than a result.

01

What survives a project, and what does not

The parts that go are the starting condition, the options I rejected, the trade-off I argued for and the actual strength of the evidence. Nobody deletes them. They stop being recoverable somewhere between the last review and the next time anyone asks.

The idea came from a post about tracking work as it happens, and it landed because I recognised the failure mode in myself. Some of those outcomes were never measured in the first place, so there is nothing to go back to even when I want it.

That gap is where hindsight does its damage. An uncertain signal becomes a clean outcome. Team delivery becomes individual ownership. A portfolio template starts deciding which parts of the work mattered. None of that takes dishonesty. It only takes distance between the work and the writing, which is the normal condition of every designer I know, including me.

So the starting question was deliberately narrow. Could an AI skill hold a trustworthy account of design work without turning every project into the same case study?

02

Six modes, and one thing the repository never holds

A reusable skill rather than a saved prompt, because the point is to run the same workflow at the next milestone without rebuilding it from memory. The instructions are Markdown, so they can be adapted across AI products, though installation and behaviour differ between them and I have not checked that they behave identically.

Six modes, covering the moments when a record is worth making.

  • Mode 01 Capture the starting point

    Where the project actually began, before hindsight tidies it.

  • Mode 02 Record the week

    Short entries while the detail is still recoverable.

  • Mode 03 Reconstruct a milestone

    Decisions, alternatives and trade-offs at a point that mattered.

  • Mode 04 Audit the evidence

    What is supported, what is weak, and what is missing.

  • Mode 05 Build a story

    For one audience and one medium, from the record as it stands.

  • Mode 06 Rebuild from old notes

    Without filling the gaps with plausible language.

The input is deliberately low effort: rough milestone notes, links, decisions, feedback, research signals and whatever outcome data exists. The output is not only portfolio writing. One record can produce a case study, a CV achievement, performance-review evidence or an interview answer, which is the whole reason for separating the evidence from the telling of it.

What the repository never holds is any of that evidence. Instructions and blank templates live in the shareable layer, real work lives in a private location the designer has approved, and the public portfolio only ever receives something a person has reviewed. Most of the risk in a tool like this is storage rather than writing, so the separation was designed before anything else.

03

Four rules that keep the record honest

1. Evidence status travels with the claim

Consequential statements are labelled verified, observed, reported, expected or unknown, and that label follows the claim into the story. A result does not get stronger because a portfolio would look better with it, and expected impact cannot quietly become delivered impact during a rewrite.

Verified Observed Reported Expected Unknown

2. Contribution is separate from outcome

The model distinguishes owned, led, co-created, contributed, influenced, facilitated and observed, and it records the designer's action alongside the team's result when both matter. The goal is precise credit, not smaller credit. Precise credit is also the version that survives an interview, because it is the version somebody else on that team would recognise.

3. STAR stays behind the scenes

STAR is a good check on whether an interview answer is complete. It is a poor structure for every portfolio piece. So the skill uses it to test coverage, then picks a shape that fits the evidence: outcome-first, decision-led, capability-led or experiment-led.

The rule that held the system together

Standardise the evidence, not the story. Consistency belongs in what gets recorded. The moment it moves into how the work is told, every project starts sounding the same and the reader stops believing any of them.

4. Privacy is a workflow, not a final redaction pass

Material is classified before it is used, capture is minimised, storage is separated, and anything unclear is treated as confidential. Removing an employer's name is not assumed to make a project safe: a distinctive combination of industry, feature, timing, team and metric can still identify it. That assumption is why these demonstration pages were written from scratch rather than redacted.

04

What I never let it decide

The skill can organise evidence and argue with a claim. It cannot grant permission to publish, interpret an employer's confidentiality rules, guarantee anonymisation, or decide that an uncertain result is true. Those are not gaps to close later. They are the line.

  • Approving what is saved, and where it is saved.
  • Confirming that each claim is accurate and carries the right evidence status.
  • Separating individual contribution from team delivery.
  • Deciding which story represents the work.
  • Reviewing every word that becomes public.
  • Deciding when a real example should be replaced with a fictional one.
  • Following the organisation's policies and contracts.

One thing I decided against on purpose: automatic collection from workplace tools. It would be convenient, and it would quietly remove the moment of consent that makes the rest of this defensible. Source selection stays with the designer unless a future integration is explicitly approved and narrowly scoped.

05

How I graded it, and where the grading was weak

I wrote three fictional tasks aimed at three different failure modes: a confidential weekly capture that the user asks to put in a public repository, a senior portfolio story with strong decision evidence and incomplete outcomes, and an evidence audit full of weak metrics, shared ownership and material that should never reach a portfolio.

Each task had a behavioural rubric written before the run. I ran every task twice, once with the skill and once with the same model working without it. I was not testing whether the model could write. I was testing whether a reusable workflow held the privacy boundaries, evidence status and contribution language in place, instead of relying on the model to rediscover those rules every time.

  • 20of 20, with the skill
  • 16of 20, without it
  • 3fictional evaluations
  • 1run each, so far
Behavioural checks passed, with the skill and without it
Evaluation With the skill Baseline
Privacy capture 6 of 6 3 of 6
Portfolio structure 7 of 7 7 of 7
Evidence audit 7 of 7 6 of 7
Total 20 of 20 16 of 20

One run per evaluation, graded by me, against my own rubric, on fictional material. That is enough to show a difference in designed behaviour and nowhere near enough to estimate how reliably it repeats.

The interesting row is the one with no difference Privacy is where the skill earned its place: it kept confidential evidence out of the repository, held the three storage boundaries, qualified the metric and refused to invent a solution direction. Portfolio structure scored identically both ways, which says the narrative guidance was compatible with good general reasoning and that this particular test proved nothing about it. A comparison that only ever flatters the thing you built is a comparison you designed badly.

What changed because of it

Two things. The evidence audit now ranks missing evidence by value and urgency instead of returning an undifferentiated list, because a list of everything you failed to record is not an action. And the trigger description was shortened so one core skill could run in more than one AI product without maintaining two competing versions of it.

06

Project Tide, from a week of notes to a record

Project Tide is invented, and so is everything in it. My own evidence is private and the implementation is private, so the only way to show you the behaviour is to build a project that has nothing to protect. Every step below is fixed. Nothing calls a model, because what I am arguing for is the designed behaviour rather than whatever a model returns today.

The moment the whole thing exists for

Prototype completion moved from 6 of 10 to 8 of 10 between two rounds. It refuses to call that an improvement, because nobody confirmed the two rounds were comparable. An attractive number stays qualified when the method behind it cannot carry the claim.

Fictional demonstration

Project Tide, a household collection service

An invented product, team and research round, written to show the workflow. Six steps, in order, from a week of rough notes to a record with its gaps still visible.

  1. Step 01

    What I bring

    Notes from a week redesigning a booking flow, written for me and nobody else. Research, a team decision, my own recommendation and an early prototype signal, all in one pile.

    • Team thought people abandoned the flow because it asked for too many fields.
    • Five moderated prototype sessions on Tuesday. Four hesitated at the collection category step.
    • Three were unsure whether a small electrical item counted as household or specialist.
    • Planned the sessions with the researcher, moderated two, synthesised the category findings.
    • Explored three directions: fewer categories, examples under every category, item before category.
    • Engineer says item-first needs a new classification service and probably misses the pilot.
    • Recommended keeping the category model, renaming two labels, examples only where people hesitated.
    • Prototype analytics: completion 6 of 10 last round, 8 of 10 this round. Same task? Not sure.
    • Still need the final label decision, the next test plan, and who owns pilot measurement.
  2. Step 02

    What it notices

    Four different kinds of thing are sitting in one list, and they do not carry the same weight. Before writing anything it sorts them, and it classifies the material as fictional so it knows which storage rules apply.

    • Research signal

      Four of five hesitated at the category step. Observed, in five sessions.

      Observed
    • Team decision

      Product and operations agreed to test the recommendation next.

      Observed
    • My recommendation

      Keep the category model for the pilot, rename two labels, add examples where people hesitated.

      Expected value
    • Prototype signal

      6 of 10 to 8 of 10, from two rounds of unknown comparability.

      Held back
  3. Step 03

    What it separates and qualifies

    Contribution first, because this is the part hindsight damages fastest. It records what I did next to what other people did, rather than compressing a team into a first person singular.

    Work Me The team
    Research planning Co-planned Researcher co-planned
    Moderation 2 of 5 Researcher, 3 of 5
    Synthesis Led the category synthesis Researcher reviewed it
    Design response Explored and recommended Content designer co-created labels, engineer advised on feasibility
    Outcome None owned or verified Measurement owner unknown

    Then the completion figures. They are kept, in raw counts, next to the reason they cannot be used: the rounds may not be comparable, and prototype completion is not production data either way. The number is not deleted, because deleting it would lose real information. It is labelled.

    6 of 10, then 8 of 10 Comparability unknown Not production data

  4. Step 04

    Where I am still needed

    Everything above is organisation. None of it is permission. The decisions that stay mine, in this scenario, are the ones with consequences outside the document.

    • Whether this capture is accurate, and whether it may be saved at all.
    • Whether the private destination is the approved one for this material.
    • Whether “led the synthesis” is a claim the researcher would recognise.
    • Whether the completion figures should appear anywhere public, in any form.
    • Whether a fictional reconstruction should replace the real example entirely.
  5. Step 05

    The record it proposes

    A milestone entry, offered for approval rather than saved. In the real workflow it would name the private destination and wait.

    Fictional record

    Category comprehension reframes the problem

    Not saved yet
    Situation

    The team believed the booking flow lost people because it asked for too many fields. Five moderated sessions pointed somewhere else: four participants hesitated at the collection category, and three could not tell where a small electrical item belonged.

    Decision and rationale

    Keep the existing category model for the pilot, rename two labels with the content designer, and add examples only where research showed hesitation. It targets the observed comprehension problem without taking on a classification service that risks the pilot date.

    Outcome status
    • Research signal: observed
    • Team agreement to test it: observed
    • Prototype completion difference: observed, not comparable
    • Improvement to pilot or production completion: unknown
    Evidence horizon, three actions
    • Record the final label decision and its rationale before the prototype changes again.
    • Confirm whether the two test rounds used comparable participants and tasks, before quoting either figure.
    • Agree who owns pilot measurement, and what signal they will collect, before the pilot starts.

    The three actions are the part I would not have written myself. They are ordered by how quickly the evidence becomes unrecoverable, not by how useful it would be to have.

  6. Step 06

    What is still unknown

    The capture ends without an outcome, and it says so rather than borrowing one. What an emerging case study could honestly claim at this point is a reframing and a decision, which is a real thing to have done.

    • Do the revised labels improve completion in the pilot?
    • Were the two prototype rounds comparable at all?
    • Who owns pilot measurement, and what will they collect?

    Six months later, that record is the difference between remembering that something worked and being able to say what was actually known at the time.

Its team, its research and its numbers were all written for this page. Nothing here comes from a workplace, a client or a research participant.

07

What I would have to prove

What I have so far is a workflow I designed and then graded myself. That is the whole of it.

  • No independent designer has used it.
  • I have not yet run it against my own real project evidence.
  • Nothing shows that regular capture improves a portfolio or an interview.
  • One run per evaluation, so this is a forward test rather than an estimate of variance.
  • Behaviour may differ across models and products.
  • The right capture rhythm has not been established.

The next two tests are about behaviour, not features. First, use it on my own milestones with private evidence I have chosen deliberately, and find out whether a capture really fits inside thirty minutes and whether three months of it makes a case-study outline easier to write. Second, sit with a few designers and work through fictional material together. That does not need the implementation published, because somebody can evaluate a workflow without being handed its source.

Enthusiasm in that session would tell me nothing. What I want to know is whether the evidence statuses make sense to somebody who did not write them, whether the questions feel proportionate to the moment, and whether anybody keeps the record going once it stops being new. That last one is the question the documentation agent already answered for me, and the answer was no.

Why the implementation stays private

A portfolio demonstration does not need the source. The concept, the safeguards, the grading and a fictional walkthrough carry the argument, and the repository can stay private while this is still being tested. Publishing it now would freeze a version I already expect to change.

The thing I keep coming back to is that a design practice is a product too, and I have never maintained mine the way I would maintain somebody else's. This does not hold my life's work and it does not replace my judgement. It holds the record behind the work, so the judgement has something true to act on.