← All work

AI lab

Three tools for the mess around the work.

Three things kept tripping me up, and none of them was the design: requests that arrive as one line, decisions nobody writes down, and a career record I only build the week I need it.

Built
2025 to 2026
Made with
Microsoft Copilot, Claude
Weights
One case study, one experiment, one note
Demonstrations
Fictional throughout

The short version

I built three working tools, a Claude skill and two Copilot agents, and two of them run inside a workplace. One was graded under controlled conditions, by me, on invented material. One showed early usefulness at work. One never made it into delivery.

The third taught me the most, and it’s why the other two are written so carefully.

01

What keeps going wrong?

They’re the same problem in three places: one person knows something important, and there’s no quick way to write it down while it’s still true.

  • Before the work The one-line request

    A story arrives saying the page needs a button. Design cannot tell what problem it solves, engineering cannot tell what it triggers, and testing has nothing to check it against.

  • During and after The decision nobody kept

    A design review settles something, then the reasoning evaporates. Six months later a new joiner inherits the screens without the argument that produced them.

  • Across a career The record you build too late

    Portfolios get maintained in an emergency. By then the evidence is scattered across files, tickets and memory, and hindsight has tidied an uncertain signal into a clean outcome.

Making messy things usable is the day job. These three are me doing it to my own work.

02

How do the three fit together?

They cover one piece of work from the request to the story you tell about it years later. Each one hands off to the next.

  1. 01 User Story clarifies

    What are we solving, for whom, and what has to be true before this is finished? It works on the request, before anyone designs.

  2. 02 UX Documentation preserves

    What did we decide, what was only a comment, and what is still unresolved? It works on the design and delivery context, and gives the answer somewhere findable to live.

  3. 03 Design Evidence keeps

    What did I actually contribute, how strong is the evidence, and what is still unknown? It works on a private career record, which later supports a case study, a CV line or an interview answer.

I move things between them by hand, on purpose. Delivery notes carry operational detail that has no place in a career record, and a career record carries opinions about people that have no place in a project wiki.

03

Why they are not the same size on this page

Two get a page and one gets a few hundred words further down. That’s how much evidence each one has, and deciding how much room a piece of work has earned is part of the job.

  • Internally evaluated prototype

    Design Evidence. Built, then graded against a rubric I wrote, using invented material. One run per evaluation.

    It does not mean another designer has used it, that it has met real project evidence, or that any outcome improved.

    Gets the full case study.

  • Private workplace tool

    UX Documentation Assistant. It runs inside a company, I use it in my own work, and selected colleagues were given access to try it.

    It does not mean a company-wide rollout, measured time saved, or documentation that improved across a team.

    Gets a shorter page, because the test was smaller.

  • Private workplace prototype

    User Story Agent. Built, tested on real incomplete input, and shared with stakeholders for evaluation.

    It does not mean adoption. Nothing it produced entered delivery.

    Gets a note, because that is the size of the evidence.

04

The two with enough behind them

  1. Case study

    Design evidence skill

    Read it →

    A Claude skill that records decisions, who did what, and how strong the evidence is, at each milestone. The record stays private until I choose what goes into a case study, a CV line or an interview answer.

    20 of 20 against a 16 of 20 baseline One run each, never used on real evidence

  2. Workplace experiment

    Copilot documentation agent

    Read it →

    I joined a team with almost no record of its own decisions, and the reason I was given was time. I think it was the blank page. So I built a Teams agent that turns a short conversation into a structured draft, with a set place to keep it. Removing the blank page turned out to be half the problem.

    One colleague, roughly five minutes, recalled No habit formed, adoption not measured

05

The agent for one-line requests

On a client portal project, story quality depended entirely on who wrote it. Some solution architects handed over full requirements. Other stories gave a designer almost nothing. One amounted to a request for a button: no action, no trigger, no outcome, and nothing QA could later use to tell whether what got built was what was meant.

So I built a Copilot agent for the thin ones. It establishes the story type first, then asks no more than two questions at a time, follows the branch that kind of work needs, and marks anything unresolved To be confirmed. It was only ever meant for the thin requests, not to outwrite the people who already wrote good ones.

Fictional demonstration

Project Willow, a workshop waitlist

An invented service and an invented request, shown as it arrived and as it came back. Fixed content, nothing sent to a model.

  • What arrived

    New feature: add a Join waitlist CTA when a workshop is full.

    That’s the whole story. It doesn’t say who can join, what joining does, or how anyone hears a place has opened.

    One line
  • What came back

    A first draft carrying the user and the problem, expected behaviour, the joining, joined, error and already-on-the-list states, acceptance criteria, dependencies, and a definition of done that includes keyboard and screen reader review.

    Five things it couldn’t know were left open.

    Reviewable draft

The bit I’d defend hardest: asked how someone finds out a place has opened, the stakeholder said to tell them. The agent asked one follow-up, got “email”, and stopped. Nobody had agreed how long a place is held, so that stayed To be confirmed. Two more questions would have got a number out of someone, and an invented number in a delivery ticket gets treated as agreed.

Reservation duration Queue ordering Success measure Priority and estimate Design reference

Project Willow is invented. The workplace project that prompted the idea does not appear here in any form.

What I’d change

I tested it myself on incomplete stories from real work, then shared it with the wider stakeholder group. One experienced stakeholder told me they wrote better stories than the agent did. They were probably right, and they were not who it was for, which was my error in how I framed it rather than theirs in how they read it.

Underneath that, I had shared a prototype before securing any of the things that would let it go anywhere: a committed pilot group who had agreed to try it on live stories, somebody who owned the story-writing workflow, or a route into the delivery tool. No agent-generated story entered delivery on that project. Nothing about time saved or story quality was measured.

Next time I’d pitch it as help for incomplete requests, then run the comparison I skipped: an original story beside the agent-assisted version, with design, engineering and QA marking both against one completeness rubric. One scenario per story type would show whether each branch surfaces the missing delivery detail without becoming heavier than writing the thing by hand.

06

What changed about how I build

Two of these worked and neither changed how a team worked. Same lesson twice, so I now do things in a different order.

Before I build anything for other people now, I want a name. Somebody who has agreed to try it on live work, whose work is the reason the thing exists. Without that, I’m guessing at demand.

  • Design the workflow with the people expected to use it, so it fits the work they already have.
  • Put an original artefact beside an agent-assisted one and let the team judge both against the same rubric.
  • Agree the adoption signal before rollout, so there is something to be wrong about.
  • Introduce it in a live working session, never a link.
  • Land it where the work already happens, not in a new place people have to remember.
  • Treat the behaviour change as part of the product, because it is the part that fails.

Most of that applies to any internal tool. It took me two to learn it.

07

What stays private, and why

Two of these run inside a workplace. What is on these pages does not.

Every demonstration you can open here is fictional. Project Tide, Project Meadow and Project Willow are invented services with invented teams, written so the behaviour that matters can be shown without a single workplace document. The complete instructions, the real inputs and outputs, the internal tools and links, and the private repository all stay where they are.

Why rebuilt rather than redacted

Blur is not anonymisation, and removing a company name does not make a project unidentifiable. A distinctive combination of industry, feature, timing and team still points at one organisation. So the demonstrations were written from scratch in a different domain instead of being scrubbed.

The cost is that you can’t check my workplace evidence. So where something is a recollection I’ve said so, and where nothing was measured I’ve said that too.

The shipped client work is where the harder evidence lives.

All work →