Prototyping - the key to "worked the first time"

Use the KIT spreadsheet and prompt to run the workflow manually, inspect the results and improve the rules before you automate it.

8 min read
On this page (11)
  1. What happened when I tested KIT
  2. What the prototype is designed to prove
  3. Build and run the prototype
  4. Run it in chat
  5. Decide whether the results are good
  6. Fun’s over, now comes the work
  7. Improve the part that failed
  8. What changes when the workflow is automated
  9. How can you enrich this system? More data
  10. What my real agent looks like
  11. What you should do now

The fastest way to find out whether an AI agent is worth building is to run the workflow manually first. So we’re going to build KIT in ChatGPT or Claude, test the results, improve the rules and then look at what changes when that workflow becomes a real agent.

What happened when I tested KIT

I ran the first version of KIT with five people from my network.

  • 3 messages were good enough to send with some basic changes.

  • 1 found a good story, but gave me a bad reason for sending it.

  • 1 returned poor story choices and should have said, ‘nothing worth sending today’.

Three of the five results were good enough to send. The articles were relevant, the connection to each person made sense and the messages gave me a useful starting point.

I’d still make some basic changes because the wording sounded slightly more AI-ee than I naturally speak.

That’s fine for a first prototype. KIT did most of the research and drafting. I handled the last bit of judgement.

The fourth result found a good article. The reason it gave me for sending it to that person was bad.

The research was ok, but the message crafting logic needs work. Sending it as written would have felt like sharing a random link and hoping that my valued contact worked out why.

A random link with ‘thought you might find this interesting’ attached to it is basically homework.

The final result returned stories that were too general to justify a conversation. KIT should have said, ‘nothing worth sending today’.

So we have two specific things to improve before the next test: how KIT explains relevance and when it should return nothing.

What the prototype is designed to prove

For KIT, the first version is a spreadsheet plus ChatGPT, Claude or another AI chat tool with the capabilities we need.

The prototype needs to answer four questions:

  • Does the AI generate messages that sound like me?

  • Are the web results interesting, or could I find better things with one Google search?

  • Can we develop a reliable prompt?

  • Can we refine that prompt so it works with a cheaper model and doesn’t rely on the latest, most expensive frontier model?

Everything else stays manual for now. We’re still trying to figure out if the idea works well enough to build out. Can KIT find topics and draft messages that are good enough for us to send?

If it can’t do that, automation just makes a useless tool faster.

Doing the work manually also highlights the nuances in what the AI needs to do. We can see where issues arise, improve the instructions and solve problems as we go.

So let’s learn quickly and cheaply first. Once we know the workflow works, we can tackle automation.

Build and run the prototype

You’ll use the KIT spreadsheet and prompt in your AI chat app of choice. I recommend doing this on a laptop or desktop because you will be moving between the spreadsheet, the AI and WhatsApp Web.

The prompt also includes space for voice profile notes. If you haven’t created a voice profile for yourself or your business, use the short voice-profile interview provided with the KIT resources. This gives the AI a clearer idea of how you naturally communicate.

**I’ll break down how this voice profiling works in a future edition.

Run it in chat

  1. Open the spreadsheet and add five important contacts. Five keeps the review manageable and avoids creating a day full of replies if several messages start conversations.

  2. Open ChatGPT or Claude. The tool needs to accept spreadsheet uploads, search the web and produce a downloadable output.

  3. Upload the completed spreadsheet.

  4. Paste the KIT prompt.

  5. Download the output.

  6. Log in to WhatsApp Web or open the desktop app.

  7. Review each result and click the link only when you are happy with the message.

Use an AI tool approved by your organisation. Don’t upload confidential CRM information unless that use is permitted. If you’re unsure, start with fictional, public or low-sensitivity contact information.

Decide whether the results are good

Fun’s over, now comes the work

Now that we’ve got the results from version one of this AI tool, we need to see what can be improved.

Ask yourself:

  • Am I happy with the drafted language?

  • Did the AI pick good articles?

  • Does the reason for sharing each article make sense for that person?

  • Did the messages start in a way that I would?

  • Did they end in a way that allows a conversation to ensue?

  • Would I actually send this? If not, why not?

KIT should be allowed to return nothing when the available stories are not good enough. Otherwise, the AI will always find something, or make something up, to say.

Improve the part that failed

Do not assume a weak output is a prompting problem. First work out where the failure started:

  • Input data: The person’s interests, role or relationship information is too broad.

  • Research: The search returned boring, old or irrelevant sources.

  • Quality rules: KIT accepted a result when it should have chosen not to answer.

  • Writing: The idea is good, but the message doesn’t sound like you.

Fix the part that failed.

Update the spreadsheet when the contact information is wrong. Tighten the search or source rules when the research is weak. Refine the prompt when KIT’s judgement or writing is the problem, using AI of course. Ask the AI tool to improve the prompt based on the specific weakness you found.

The important part is to change one thing at a time and run the sheet again. If you update the input data, change the source, rewrite the rules and alter the writing style together, you will have no idea what actually fixed it.

Keep a simple record of the contact, the failure, the change and the new result.

This turns random prompt tinkering into a proper test. When the same change improves several different cases, it is ready to become part of the next version.

If you use Claude Code or Codex, you can automate some of this prompt improvement. Give the tool clear examples, a way to judge the results and a stopping condition. You still need to review what it changes.

By now you have something that is working well. It’s still quite manual to use, which is where automation comes in.

Note: the iterative nature of this work will highlight one annoying fact of llms at the moment - they take forever to use tools, in this case, the spreadsheet reading/writing tools. You could just get the llm to print the output on the screen instead of manipulating the spreadsheet. Just ask it.

Me waiting for AI to use a tool

What changes when the workflow is automated

Once the manual workflow is producing useful results, automation has a clear job: collect the context, run the process consistently and keep track of what happened.

How can you enrich this system? More data

Now imagine piping relationship data into the AI. It can see what has been happening in that relationship, so the process becomes much richer.

That gets us closer to the likes of an executive assistant.

Useful context could include:

  • Name, role, company, industry and likely business needs.

  • Relationship tier and the nature of the relationship.

  • The last 20 touchpoints and the date of the last meaningful contact.

  • A summary of the previous conversation.

  • Topics, priorities or projects discussed.

  • Outstanding promises.

  • Preferred communication channel.

  • Appropriate personal or professional notes.

  • Trusted profile or company links.

  • Reasons to delay or avoid contact.

This is the value of context.

We’ll unpack context properly in a future edition because, in my opinion, it is one of the most important AI concepts to understand. It’s how you give the AI a much better understanding of the work it is doing for you.

Collecting all this data manually would be a mission. For my live version, I’ll be using my dev kungfu to collect it from the CRM I built.

What my real agent looks like

My agent is connected to my custom CRM, so the next five contacts and their relationship data can be retrieved automatically. It uses tools and API’s (application programming interfaces) to access the CRM, search the web and update the records.

LangGraph coordinates the steps and keeps track of where the workflow has reached. With persistence configured, the workflow can continue after an interruption instead of starting again.

The code manages the process and calls the LLMs using prompts similar to the one you copied into your chat.

The automated version brings in the relationship history, runs the searches, returns each recommendation in the same structure and records what should happen next.

The approval still sits with me. KIT proposes the contact, source and message, then I decide whether anything gets sent.

So the prototype and the agent are the same workflow with different levels of automation. Your job is to prove that the workflow creates enough value to deserve the automation.

What you should do now

Run KIT with five people. Record what failed. Improve one part. Then run the same five people through it again.

After a few batches, you should understand which context KIT needs, which rules improve the recommendations and where your judgement still matters.

That becomes the design for the system you can automate. The process is simple, but as usual, the doing is on you.

Get the next one

Subscribe free

Free. One edition a week. Unsubscribe in one click.