Standardise the evidence, not the story.
A privacy-first AI skill that keeps a trustworthy record of design work while the details are still true, then helps turn that record into whichever story is needed later.
-
Layer 01
The reusable skill
Instructions, the evidence model and blank templates. Nothing about any actual project lives here.
Shareable -
Layer 02
The private ledger
Real evidence from real work, in a location the designer has approved. It never travels upward on its own.
Never published -
Layer 03
The public portfolio
Only material that has been reviewed, sanitised and qualified. A person decides what crosses this line.
Reviewed by me
TLDR
Nobody maintains their portfolio until they need it, and by then the evidence has scattered across design files, tickets, chat and memory. The screens survive. The reasoning around them does not.
So I built an AI skill that captures decisions, contribution and evidence status while the project is running, keeps that record private, then helps shape it into a case study or an interview answer. Graded against the same model without it, the whole gap was privacy.
- 20of 20 behavioural checks
- 16of 20 without the skill
One run each, graded by me. No other designer has used it and it has never met my own real evidence, so this is a forward test rather than a result.
01
What survives a project, and what does not
The parts that go are the starting condition, the options I rejected, the trade-off I argued for and the actual strength of the evidence. Nobody deletes them. They stop being recoverable somewhere between the last review and the next time anyone asks.
The idea came from a post about tracking work as it happens, and it landed because I recognised the failure mode in myself. Some of those outcomes were never measured in the first place, so there is nothing to go back to even when I want it.
That gap is where hindsight does its damage. An uncertain signal becomes a clean outcome. Team delivery becomes individual ownership. A portfolio template starts deciding which parts of the work mattered. None of that takes dishonesty. It only takes distance between the work and the writing, which is the normal condition of every designer I know, including me.
So the starting question was deliberately narrow. Could an AI skill hold a trustworthy account of design work without turning every project into the same case study?
02
Six modes, and one thing the repository never holds
A reusable skill rather than a saved prompt, because the point is to run the same workflow at the next milestone without rebuilding it from memory. The instructions are Markdown, so they can be adapted across AI products, though installation and behaviour differ between them and I have not checked that they behave identically.
Six modes, covering the moments when a record is worth making.
-
Mode 01
Capture the starting point
Where the project actually began, before hindsight tidies it.
-
Mode 02
Record the week
Short entries while the detail is still recoverable.
-
Mode 03
Reconstruct a milestone
Decisions, alternatives and trade-offs at a point that mattered.
-
Mode 04
Audit the evidence
What is supported, what is weak, and what is missing.
-
Mode 05
Build a story
For one audience and one medium, from the record as it stands.
-
Mode 06
Rebuild from old notes
Without filling the gaps with plausible language.
The input is deliberately low effort: rough milestone notes, links, decisions, feedback, research signals and whatever outcome data exists. The output is not only portfolio writing. One record can produce a case study, a CV achievement, performance-review evidence or an interview answer, which is the whole reason for separating the evidence from the telling of it.
What the repository never holds is any of that evidence. Instructions and blank templates live in the shareable layer, real work lives in a private location the designer has approved, and the public portfolio only ever receives something a person has reviewed. Most of the risk in a tool like this is storage rather than writing, so the separation was designed before anything else.
03
Four rules that keep the record honest
1. Evidence status travels with the claim
Consequential statements are labelled verified, observed, reported, expected or unknown, and that label follows the claim into the story. A result does not get stronger because a portfolio would look better with it, and expected impact cannot quietly become delivered impact during a rewrite.
2. Contribution is separate from outcome
The model distinguishes owned, led, co-created, contributed, influenced, facilitated and observed, and it records the designer's action alongside the team's result when both matter. The goal is precise credit, not smaller credit. Precise credit is also the version that survives an interview, because it is the version somebody else on that team would recognise.
3. STAR stays behind the scenes
STAR is a good check on whether an interview answer is complete. It is a poor structure for every portfolio piece. So the skill uses it to test coverage, then picks a shape that fits the evidence: outcome-first, decision-led, capability-led or experiment-led.
The rule that held the system together
Standardise the evidence, not the story. Consistency belongs in what gets recorded. The moment it moves into how the work is told, every project starts sounding the same and the reader stops believing any of them.
4. Privacy is a workflow, not a final redaction pass
Material is classified before it is used, capture is minimised, storage is separated, and anything unclear is treated as confidential. Removing an employer's name is not assumed to make a project safe: a distinctive combination of industry, feature, timing, team and metric can still identify it. That assumption is why these demonstration pages were written from scratch rather than redacted.
04
What I never let it decide
The skill can organise evidence and argue with a claim. It cannot grant permission to publish, interpret an employer's confidentiality rules, guarantee anonymisation, or decide that an uncertain result is true. Those are not gaps to close later. They are the line.
- Approving what is saved, and where it is saved.
- Confirming that each claim is accurate and carries the right evidence status.
- Separating individual contribution from team delivery.
- Deciding which story represents the work.
- Reviewing every word that becomes public.
- Deciding when a real example should be replaced with a fictional one.
- Following the organisation's policies and contracts.
One thing I decided against on purpose: automatic collection from workplace tools. It would be convenient, and it would quietly remove the moment of consent that makes the rest of this defensible. Source selection stays with the designer unless a future integration is explicitly approved and narrowly scoped.
05
How I graded it, and where the grading was weak
I wrote three fictional tasks aimed at three different failure modes: a confidential weekly capture that the user asks to put in a public repository, a senior portfolio story with strong decision evidence and incomplete outcomes, and an evidence audit full of weak metrics, shared ownership and material that should never reach a portfolio.
Each task had a behavioural rubric written before the run. I ran every task twice, once with the skill and once with the same model working without it. I was not testing whether the model could write. I was testing whether a reusable workflow held the privacy boundaries, evidence status and contribution language in place, instead of relying on the model to rediscover those rules every time.
- 20of 20, with the skill
- 16of 20, without it
- 3fictional evaluations
- 1run each, so far
| Evaluation | With the skill | Baseline |
|---|---|---|
| Privacy capture | 6 of 6 | 3 of 6 |
| Portfolio structure | 7 of 7 | 7 of 7 |
| Evidence audit | 7 of 7 | 6 of 7 |
| Total | 20 of 20 | 16 of 20 |
One run per evaluation, graded by me, against my own rubric, on fictional material. That is enough to show a difference in designed behaviour and nowhere near enough to estimate how reliably it repeats.
What changed because of it
Two things. The evidence audit now ranks missing evidence by value and urgency instead of returning an undifferentiated list, because a list of everything you failed to record is not an action. And the trigger description was shortened so one core skill could run in more than one AI product without maintaining two competing versions of it.
06
Project Tide, from a week of notes to a record
Project Tide is invented, and so is everything in it. My own evidence is private and the implementation is private, so the only way to show you the behaviour is to build a project that has nothing to protect. Every step below is fixed. Nothing calls a model, because what I am arguing for is the designed behaviour rather than whatever a model returns today.
The moment the whole thing exists for
Prototype completion moved from 6 of 10 to 8 of 10 between two rounds. It refuses to call that an improvement, because nobody confirmed the two rounds were comparable. An attractive number stays qualified when the method behind it cannot carry the claim.
Fictional demonstration
Project Tide, a household collection service
An invented product, team and research round, written to show the workflow. Six steps, in order, from a week of rough notes to a record with its gaps still visible.
-
Step 01
What I bring
Notes from a week redesigning a booking flow, written for me and nobody else. Research, a team decision, my own recommendation and an early prototype signal, all in one pile.
- Team thought people abandoned the flow because it asked for too many fields.
- Five moderated prototype sessions on Tuesday. Four hesitated at the collection category step.
- Three were unsure whether a small electrical item counted as household or specialist.
- Planned the sessions with the researcher, moderated two, synthesised the category findings.
- Explored three directions: fewer categories, examples under every category, item before category.
- Engineer says item-first needs a new classification service and probably misses the pilot.
- Recommended keeping the category model, renaming two labels, examples only where people hesitated.
- Prototype analytics: completion 6 of 10 last round, 8 of 10 this round. Same task? Not sure.
- Still need the final label decision, the next test plan, and who owns pilot measurement.
-
Step 02
What it notices
Four different kinds of thing are sitting in one list, and they do not carry the same weight. Before writing anything it sorts them, and it classifies the material as fictional so it knows which storage rules apply.
-
Research signal
Four of five hesitated at the category step. Observed, in five sessions.
Observed -
Team decision
Product and operations agreed to test the recommendation next.
Observed -
My recommendation
Keep the category model for the pilot, rename two labels, add examples where people hesitated.
Expected value -
Prototype signal
6 of 10 to 8 of 10, from two rounds of unknown comparability.
Held back
-
Research signal
-
Step 03
What it separates and qualifies
Contribution first, because this is the part hindsight damages fastest. It records what I did next to what other people did, rather than compressing a team into a first person singular.
Work Me The team Research planning Co-planned Researcher co-planned Moderation 2 of 5 Researcher, 3 of 5 Synthesis Led the category synthesis Researcher reviewed it Design response Explored and recommended Content designer co-created labels, engineer advised on feasibility Outcome None owned or verified Measurement owner unknown Then the completion figures. They are kept, in raw counts, next to the reason they cannot be used: the rounds may not be comparable, and prototype completion is not production data either way. The number is not deleted, because deleting it would lose real information. It is labelled.
-
Step 04
Where I am still needed
Everything above is organisation. None of it is permission. The decisions that stay mine, in this scenario, are the ones with consequences outside the document.
- Whether this capture is accurate, and whether it may be saved at all.
- Whether the private destination is the approved one for this material.
- Whether “led the synthesis” is a claim the researcher would recognise.
- Whether the completion figures should appear anywhere public, in any form.
- Whether a fictional reconstruction should replace the real example entirely.
-
Step 05
The record it proposes
A milestone entry, offered for approval rather than saved. In the real workflow it would name the private destination and wait.
Not saved yetFictional record
Category comprehension reframes the problem
SituationThe team believed the booking flow lost people because it asked for too many fields. Five moderated sessions pointed somewhere else: four participants hesitated at the collection category, and three could not tell where a small electrical item belonged.
Decision and rationaleKeep the existing category model for the pilot, rename two labels with the content designer, and add examples only where research showed hesitation. It targets the observed comprehension problem without taking on a classification service that risks the pilot date.
Outcome status- Research signal: observed
- Team agreement to test it: observed
- Prototype completion difference: observed, not comparable
- Improvement to pilot or production completion: unknown
Evidence horizon, three actions- Record the final label decision and its rationale before the prototype changes again.
- Confirm whether the two test rounds used comparable participants and tasks, before quoting either figure.
- Agree who owns pilot measurement, and what signal they will collect, before the pilot starts.
The three actions are the part I would not have written myself. They are ordered by how quickly the evidence becomes unrecoverable, not by how useful it would be to have.
-
Step 06
What is still unknown
The capture ends without an outcome, and it says so rather than borrowing one. What an emerging case study could honestly claim at this point is a reframing and a decision, which is a real thing to have done.
- Do the revised labels improve completion in the pilot?
- Were the two prototype rounds comparable at all?
- Who owns pilot measurement, and what will they collect?
Six months later, that record is the difference between remembering that something worked and being able to say what was actually known at the time.
Its team, its research and its numbers were all written for this page. Nothing here comes from a workplace, a client or a research participant.
07
What I would have to prove
What I have so far is a workflow I designed and then graded myself. That is the whole of it.
- No independent designer has used it.
- I have not yet run it against my own real project evidence.
- Nothing shows that regular capture improves a portfolio or an interview.
- One run per evaluation, so this is a forward test rather than an estimate of variance.
- Behaviour may differ across models and products.
- The right capture rhythm has not been established.
The next two tests are about behaviour, not features. First, use it on my own milestones with private evidence I have chosen deliberately, and find out whether a capture really fits inside thirty minutes and whether three months of it makes a case-study outline easier to write. Second, sit with a few designers and work through fictional material together. That does not need the implementation published, because somebody can evaluate a workflow without being handed its source.
Enthusiasm in that session would tell me nothing. What I want to know is whether the evidence statuses make sense to somebody who did not write them, whether the questions feel proportionate to the moment, and whether anybody keeps the record going once it stops being new. That last one is the question the documentation agent already answered for me, and the answer was no.
Why the implementation stays private
A portfolio demonstration does not need the source. The concept, the safeguards, the grading and a fictional walkthrough carry the argument, and the repository can stay private while this is still being tested. Publishing it now would freeze a version I already expect to change.
The thing I keep coming back to is that a design practice is a product too, and I have never maintained mine the way I would maintain somebody else's. This does not hold my life's work and it does not replace my judgement. It holds the record behind the work, so the judgement has something true to act on.