← AI lab

Documentation should not need a blank-page expert.

A Copilot agent that turns a short conversation into a structured record, and the far harder question of why a team that wanted documentation still did not write any.

Project
UX Documentation Assistant
Type
Workplace initiative, AI lab
Role
Agent workflow, documentation model, adoption guidance
Year
2025
Status
Private workplace tool
  1. 01 A short conversation

    Four or five questions in Teams. Fragments are fine. The one-page guide says so explicitly, because people write less when they think they are being marked.

  2. 02 A structured draft

    Decisions kept apart from comments and assumptions, and anything unresolved left standing as To be confirmed with an owner beside it.

  3. 03 A findable record

    A person reviews it and puts it in a page whose name and location were decided in advance, so somebody can find it in six months.

The third box is the one people skip Generating documents is not the same as having documentation. Without a predictable home, a good agent just produces more things nobody can find, which is the original problem in a better font.

TLDR

Teams that want documentation and have none usually say the reason is time. I did not think it was time. I thought it was the blank page.

So I built a Copilot agent that asks a few short questions in Teams and returns a structured draft, plus the Confluence structure to keep it in. I tested it on real work with deliberately incomplete notes, I use it myself, and I gave selected colleagues access. One of them documented part of a real project and recalled it taking around five minutes, then raised it with senior management.

Removing the blank page was the easy half. I never set up a pilot, an owner or a session where we did it together, so the habit never formed.

01

I joined, and there was nothing to read

On my first days I went looking for the things a new designer needs: what the product does, which decisions had already been made, where the files live. The Confluence space existed and was close to empty. The explanation I got was that nobody had time.

The cost of that landed on everyone. Designers could not always reconstruct why an earlier solution had taken the shape it did. Developers built from designs with no record of the intent behind them. Problem framing was hard to find. A new person had no reliable way to build context except by asking people to remember.

It became sharper on a client portal project. Conversations and design iterations moved back and forth with no durable record of what changed or why. On a handover I inherited designs and comments without the reasoning, and the direction changed, changed back, and then raised a more fundamental question: similar functionality already existed at platform level. Documentation would not have made that decision. It would have made the history and the assumptions available at the moment the team needed to weigh it.

None of that was a writing problem. Documentation had never become part of delivery, so it always arrived as an extra task after the work was finished, facing a blank page and an unclear standard. That is a design problem rather than a discipline problem, which is why I thought I could do something about it.

02

What I built, and the half that usually gets skipped

An agent reached through Teams, because Microsoft Copilot, Teams and Confluence were already approved by the company and the company had started training people to build Copilot agents. Working inside approved tools meant the experiment could happen at all, and it shaped the design more than any preference of mine would have.

Four kinds of work, one way in

  • Type 01 Screen and feature design

    What it does, what was agreed, which states exist.

  • Type 02 Prototypes and user flows

    The route through, and what it assumes about the person taking it.

  • Type 03 Research sessions

    What was asked, what came back, what it does and does not support.

  • Type 04 Workshops and ideation

    Who was there, what was generated, what was actually chosen.

The agent asks conversational questions, originally intended to take two or three minutes, then produces a formatted document. The person reviews it and copies it into the matching Confluence page.

Where the output has to land

So I defined the structure the output lands in, with a page-name pattern per document type and one obvious destination for each.

  • UX design space
    • Project
      • Overview and goals
      • Screen and feature docs
      • User flows and prototypes
      • Research and testing
      • Workshops and ideation
The shape, not the space A generalised version of the structure I proposed. The real space, its projects and its contents stay private. The point of showing it at all is that the four areas match the four documentation types exactly, so a person who has answered the questions already knows where the answer goes.

The guide also set out the first adoption steps: create the space, set up the project folders, share the agent through Teams, run a five-minute demonstration, and document one piece of real work together. The agent, the one-page guide and the proposed structure were my own initiative, with AI used to help me brainstorm and build.

Ask, do not hand over another template

A blank template still expects you to know what belongs in it. A short conversation lowers the cost of starting and lets people answer in their own words, roughly. That mattered particularly in a multinational team where English was not everyone's first language. The aim was never to flatten how people write. It was to stop polished English being a prerequisite for contributing to the shared record.

Missing stays visible

Anything the person does not know is marked To be confirmed rather than filled in with something plausible. An agent that answers a gap with confident wording is worse than no agent, because the document then looks settled and the next person treats it as a decision.

Publication stays human, and that is also a limitation

The person copies the output into Confluence instead of the agent publishing it. That creates a natural review point, which I wanted. It is also honestly a constraint: Copilot and Confluence are not connected, so I designed the review step to do useful work rather than pretending the gap was a feature.

The call underneath all of it

Generating a document does not make documentation findable. Without the structure and the naming convention, this becomes a machine for producing pages nobody can retrieve. Faster than before, and no easier to find.

03

Project Meadow, one review, one record

The real agent, its instructions and every document it has produced stay inside the company. What follows is an invented product with an invented design review, written so the behaviour that matters can be shown without a workplace page anywhere near it. It is fixed content. Nothing is sent to a model.

The moment worth watching

Seven things come out of one review and only one of them is a decision. The rest are feedback, an assumption, a constraint and an action. Sorting those is the whole job, and the empty state that nobody agreed stays To be confirmed with a name against it.

Fictional demonstration

Project Meadow, a team scheduling service

An invented product and an invented design review about a notification filter. Six steps, from a handful of fragments to a record another person could act on.

  1. Step 01

    What the designer brings

    Notes taken during a review, in the state notes are actually in. Nothing here is wrong. It is just not yet a record, because everything in it carries the same weight.

    • notification centre filters
    • people need to find urgent shift changes, current list all mixed together
    • proposed chips: All, Action needed, Unread
    • support lead liked Action needed because apparently people miss shift changes
    • product manager asked whether Updates would be easier language
    • engineering can support the three states for pilot
    • agreed keep Action needed after discussion
    • default All so people do not think messages disappeared
    • not sure what happens when a filter has no results
    • need product manager to confirm empty state before handoff
    • think this makes important changes easier to find but no testing yet
  2. Step 02

    What the agent notices

    First, which kind of record this is. A screen and feature document, not a research write-up or a workshop record, which changes the questions it asks and the place the answer will eventually live.

    Screen and feature design Not research Not a workshop

    Then it asks four short questions, and accepts fragments back. What problem, and for whom. What was agreed, and what was only raised. What is unresolved, and who owns it. What evidence exists, and what is still assumption.

  3. Step 03

    What it separates

    Seven things arrived in one list. They are not the same kind of thing, and treating them as if they were is how a comment ends up in a record as a decision.

    • Decision

      Three filters: All, Action needed, Unread. Default All, and newest first stays as it is.

      Agreed in the review
    • Feedback

      The product manager suggested Updates instead of Action needed. The group kept the original after discussion.

      An alternative, not a label
    • Reported signal

      Support hears that people miss shift changes. The underlying tickets have not been reviewed.

      Not research evidence
    • Assumption

      Filtering will make important changes easier to find. Nobody has tested this.

      Untested

    Engineering confirming the three states are feasible is a constraint, and the empty state is an action with an owner. Both are kept, both are labelled as what they are.

  4. Step 04

    Where a person is still needed

    The draft is coherent enough to review and nowhere near approved. What the agent must not do is the useful list here.

    • It must not turn support feedback into user research.
    • It must not state that the filters reduce missed changes.
    • It must not decide the no-results behaviour.
    • It must not present Updates as an agreed label.
    • It must not invent a deadline, an owner or a design link.
    • It must not publish anything without a person reading it first.

    In this scenario the reviewer confirms the attendee roles, adds the private design link only inside the private page, and publishes with the unresolved rule still in it. That is allowed, because the gap has an owner and a deadline attached to it.

  5. Step 05

    The record it produces

    Fictional record

    Notification filter, UX documentation

    One behaviour unresolved
    Problem

    Shift leaders receive several kinds of notification in one chronological list, so changes that need a response can be hard to distinguish from general updates.

    Decision

    For the pilot, filter by All, Action needed and Unread. Default to All so nothing appears to have vanished, and keep newest-first ordering inside each filter. The team considered Updates and kept Action needed, because the label signals that a response may be required.

    States and behaviour
    Filter states and their expected behaviour
    State Expected behaviour
    All Selected by default, every notification, newest first.
    Action needed Notifications that require a response or an action.
    Unread Notifications not yet marked as read.
    No results To be confirmed
    Open action

    Confirm whether a filter with no results reuses the generic empty state or explains the selected filter and offers Clear filter. Owner: product manager. Required before design handoff.

    Evidence and assumptions
    • Reported: support hears about missed changes, tickets not reviewed.
    • Confirmed: engineering can support the three states for the pilot.
    • Assumed: filtering makes important changes easier to find. Untested.
  6. Step 06

    What is still unknown

    The record does not resolve the design. It makes the shape of the remaining work visible, which is a smaller claim and a more useful one.

    • Does the unresolved empty state block handoff, or can it follow later?
    • Do the filters actually help anyone find a shift change?
    • How would the team test that, and who decides it has been answered?

    Six months later, the value of this page is not how well it is written. It is that a new person can see what was decided, what was only suggested and what was never settled, without asking four people who half remember.

Project Meadow is invented. It is not the client portal, it is not any product I have worked on, and its team does not exist.

04

What happened when other people used it

I tested it on a real piece of work rather than an idealised brief, and I fed it deliberately incomplete and awkward input to see whether it would expose the gaps or paper over them. I use it in my own workflow. I could not make it available company-wide, so I gave selected members of the design team access to try it.

One colleague used it to document part of a real project. They recalled that prompting their thoughts and producing a working document took around five minutes, and they later raised the initiative with senior management as an example of how I was applying AI in my own work. Development and early testing ran through roughly November and December 2025.

Claims and the evidence behind them
Claim What is behind it
Around five minutes to a working document One colleague's recollection of one document. Not timed, not repeated.
A comparable document by hand takes a few hours My own estimate from experience. There was no controlled comparison.
Shared with the design team Selected participants only. Not a company-wide rollout.
Time saved across the team Not measured. I am not claiming it.
Documentation quality improved Not measured. I am not claiming that either.

Two useful numbers exist and both are recollections. Putting five minutes against a few hours would make a much better slide than it would make a finding.

Other things I do not know: how often the selected participants used it, whether the original two-to-three-minute expectation holds across document types, whether the four templates covered the team's real needs, how much correcting the output usually needed, and whether the Confluence structure was adopted, piloted or quietly ignored.

05

The lesson I took

The agent cut the effort of turning my own thoughts into a coherent document, and another designer produced a working document fast enough to mention it upward. Both true. Neither is the finding.

Removing the blank page did not create a documentation habit.

I shared a link and expected the tool to carry the rest. What I had not set up was anything that actually carries a new practice: somebody who had agreed to try it on live work, a person who owned documentation as part of delivery, a session where we wrote one together, and a signal we had agreed to look at afterwards. I built the half I could build on my own and treated the other half as rollout.

Next time it arrives as a guided pilot instead of a link. One live session, one active piece of work documented together in the room, one person who has agreed to keep going after the session, and a measure agreed before we start. Speed on its own tells me almost nothing. What I want is accuracy, how much editing the draft needed, and whether the items left To be confirmed were ever confirmed.

Further out, the workflow I would want connects the reviewed document to a dedicated space, then links that record into the delivery ticket alongside the design file, so there is a traceable path from problem to decision to build to design QA. Those are directions, not features. The current version stops at a document a person copies into place.

What changed for good is how I think about putting AI into a team. The quality of the agent is maybe a third of it. Facilitation, landing it where the work already happens, and one person with a reason to keep using it are the rest, and I design for those now instead of hoping for them.