Taking an AI agent from PoC to production: a documented process, system access and a decision after testing

For:
CTOs, CIOs, IT managers and architects planning AI agent PoCs before a production decision
Reading time:
10 min
Episode:
25 min

This episode is in Polish. Full Polish version with transcript

After you click, the video loads from YouTube (Google). Google may store data on your device and process it in the USA as well. More (PDF, in Polish) · Watch on YouTube

In brief

You have an idea for an AI agent and need to decide whether to build the full solution. A PoC prepares that decision. Before it, the team writes the process down from what employees know and checks access to systems; after testing, it rates the result as scale, iterate or kill against criteria set in advance.

Key takeaways

  • Before the PoC starts, the person who runs the process writes down in Word what they do and why, or records their screen. This preparation saves a lot of work that should never be left to debugging.
  • Business people who want to test something should get a playground with the same bare model the agent will be built on, because that model behaves differently from ChatGPT.
  • In most of Protopia's PoCs, access to data and systems causes the most problems, and building the agent itself is mostly tedious integration work.
  • Decide at the start what triggers the agent and at which point a human enters the process. An approval matrix sets what can run automatically. Processes can run without a human when any problem would be very small and would harm neither reputation nor finances.
  • Success criteria come before testing. Tests stay focused on the goal: the UI is ugly on purpose, and a log records the agent's run, its duration and its decisions.
  • After testing, the decision is scale, iterate or kill, and a kill with well-documented findings is a success too, because it keeps the team out of a project that would not have worked anyway.

A PoC tests an AI agent before you invest in a full solution

A proof of concept (PoC) is a feasibility study that checks whether a concept will work and whether it will work consistently. Protopia often calls it a proof of value (PoV) as well, because it has to check whether the solution will deliver value and whether it is feasible.

One client wrote that the prompt is very precise, yet the answers differ in substance. For a PoC to succeed, you sometimes need to approach the process a little differently and use business knowledge differently than when you simply drop a task into GPT, Copilot, Gemini or Claude.

Diagram of an AI agent's path from idea to decision. Preparation covers writing the process down from what employees know, a playground for the business and a check of access to systems. The actual PoV starts with success criteria, followed by tests with logging. After testing, a decision follows: scale leads to a pilot and production, iterate goes back to new tests, and kill closes the initiative with documented findings.
Diagram: from writing the process down to the decision after PoC testing

The first step is writing the process down from employee knowledge

Before the PoC, you need to know exactly what the process to automate looks like: what people do by hand today and how it runs in the systems. Protopia has the person who holds this knowledge write down in Word, step by step, what they do and why, or record their screen. Sometimes two employees or experts can do it.

Tribal knowledge is the knowledge of the people who run the process. Writing it down is the first step, and in Protopia’s projects it is the step that contributes most to success. Often nobody updates the instructions when business conditions, the system configuration or the approach change.

At one insurance client, an employee wrote down that they first take data from the system, then search the internet to find out who the customer really is and what they could be offered, and finally try to match it all and do an analysis. A Word document like this is a great starting point for planning, because it shows what actually happens and what the person does. A week of preparation and data gathering saves a great deal of work that should never be left to debugging.

Conway’s law shows that a process mirrors the communication structure of the company, not the logic it should have. Writing the process down can therefore be a chance to check whether it is worth simplifying and cleaning up.

A playground shows the business how the model really behaves

After the process is written down, you need to set expectations. Non-technical people should be able to check for themselves how agents behave: on the playgrounds the team later builds the solution from, not in Copilot or ChatGPT. Protopia works with Azure most of the time. A playground is a chat-like place that gives access to the bare language model, and this model later goes into agents and applications.

A bare model behaves quite differently from ChatGPT in agent mode. For example, in your own OpenAI model hosted on Azure, web search works somewhat differently: the Bing Search underneath does not behave the way it does in the public consumer ChatGPT. On the playground, the business sees how the model actually works, including with a prompt like the one in the client’s message.

If business people want to test something, they should get access to a playground instead of having the tools blocked. They then form realistic expectations and build trust. These tests happen before the actual PoC, at the stage of play and discovery.

Access to systems makes the analysis hardest

After the playground tests comes the analysis, the worst phase. In most of Protopia’s PoCs, automation and agent behavior cause no problems. Łukasz Kałużny: “Most of the problems come down to access to systems.” (translated)

Someone has to do the painstaking work: find out where the needed data comes from, from which system, how to get access to it and whether the system has an API. If there is no API, check whether access can be automated through the browser or whether access to the database is enough. This is an inventory of the systems the automated process uses. Today it is most often about pulling data and acting on it or, much as with RPA (robotic process automation) in the past, writing back to another system or calling an API.

Decision tree for an agent's access to data. The team finds out which data is needed and which system it comes from, then checks whether that system has an API. If it does, the agent calls the API. If it does not, the team checks whether access can be automated through the browser or whether access to the database is enough.
Diagram: how to set up an agent's access to data

Once the process is written down, most of the remaining technical work is tedious integration. Building agents is very much a programming job, not data science.

The trigger and the human’s role are set at the start

At the start, you need to answer two questions. The first is what triggers the agent. The agent will not figure it out by itself, so something has to trigger its work, for example an email in a mailbox, an event or a human click. The second is human in the loop: at which point a human enters the process, approves the agent’s work or uses it.

An approval matrix is a set of clear rules for when something can run automatically and when it cannot. It sets the human’s role in the process. Protopia often advises having someone approve at the end, and in many cases that is enough.

Diagram of how an AI agent starts and where a human fits into the process. An email in a mailbox, an event or a human click starts the agent. The result of the agent's work goes to the approval matrix. The matrix sends it to a human who approves the result at the end, or to automatic execution when a possible problem is very small and harms neither reputation nor finances.
Diagram: the agent trigger and the approval matrix
The human’s role in processes with an agent
Process What the agent does Human role
Customer complaint or inquiry Works through the case, prepares the justification and the reply, shortens delivery time Checks what goes out to the customer before it is sent; may change nothing and click “send”
PoC that gathers data and prepares an analysis Prepares a finished report with all sources Checks on their own whether the report is correct
Price quote based on web search The agent searches the web and proposes a price quote; an extra judge agent named Judge screenshots the source pages and passes the price on when the basic information on those pages checks out Approves the result at the end
After-sales handling of open training courses under Protopia’s Patoarchitekci brand (sending invitations, adding people to the course) Handles the course purchase No human: any problem would be very small and would harm neither reputation nor finances

Protopia always leaves the interpretation of the AI Act to the client, because it is not Protopia’s specialty. When you set the human’s role, what counts is the notion of high-risk systems and how the company’s lawyers and compliance interpret it. The first idea of where the human belongs may turn out wrong, and that comes out during testing. That is one of the things a PoV is for.

Success criteria are written down before the first test

The actual PoV starts with clear success criteria. In the PoV, Protopia usually tries a first-pass automation of the whole process. You need to agree on what exactly will be tested: for example the most important and most problematic elements, or the whole process without edge cases and corner cases, that is, the happy path with branches that cause few problems. A sample criterion: in the PoC phase, 80% of cases are handled correctly.

The client’s business expert rates results such as reports. Some quality traits are very hard to measure: whether a letter reads well always depends on human judgment and is highly subjective.

Tests stay focused on the goal, with an ugly UI and logging

Tests should stay focused on the goal and in most cases should not be about checking technology or playing with a new framework. For testing, Protopia most often provides a piece of UI, usually as ugly as possible, so that nobody is tempted to deploy it to production the next day. This UI is only minimally usable.

From the start, you need reasonably good logging across the whole process: how the agent worked, how long it took, what decisions it made and what happened next. In a PoC, a text file is enough for this.

Logs are a catalog of errors and a roadmap: they show what needs fixing and what not to fix. One example is a discovered edge case that must be fixed if the project moves to a pilot or to production, but not at this point. Tests also show run time. At another client, the agent cannot respond within 30 seconds because there is too much data, and there is no way around that. Generating a report can take 5 minutes, because the agent has to go through several dozen pages before it makes a decision. Such limits are fine, because they help set expectations and build the roadmap, so that the decision to go to production brings no surprises.

After testing, one of three decisions follows: scale, iterate or kill

Three states after PoC testing
State When What next
Scale The PoC met its goals: it reached the written success criteria, or the business subjectively confirmed that it wants to go further Pilot and production; the experiment becomes a real project
Iterate The result is close to success, and the team knows what to improve and that it needs time for it; a gut feeling that it will be fine is not enough Another iteration and new tests
Kill The criteria are not being met and the team is going in circles Consider closing the initiative; write down the findings

In the pricing process, the iteration was to add a judge agent that verifies the result. This change and another full round of tests needed 3–4 more days.

Kill is the most hated status, because you cannot announce a success. If the findings are well documented, you know why it failed, and that is a success too. Łukasz Kałużny: “The operation was carried out successfully. The patient died.” (translated)

The reason may be an idea that was not worth automating, or missing access to an API. Unrealistic expectations are another reason, though this happens less and less often. Adding automation to a system may also turn out too expensive when the cost of building the API eats up the potential gains. Closing protects the team from drowning in a project that would not have worked anyway.

With good analysis, most of these PoCs end in success, even when the cost summary shows that the project makes no sense.

Speakers

  • Łukasz Kałużny

    Łukasz Kałużny

    Founder, Managing Partner, Technology Advisor.

    Łukasz Kałużny co-founded Protopia and is a Microsoft MVP in the Microsoft Foundry category. He has co-hosted Patoarchitekci since 2019, talking about IT architecture, GenAI and AI agents without the marketing spin.

    All posts by this author
  • Mikołaj Szczerbicki

    Mikołaj Szczerbicki

    Head of Sales & Business Development.

    Mikołaj Szczerbicki is Head of Sales & Business Development at Protopia and co-hosts the Powered by Protopia podcast. He scopes and prices projects, so he asks about the cost, risk and timeline of AI, Azure and Kubernetes work.

    All posts by this author

FAQ

What if the procedures are outdated or missing?

You still have to get the tribal knowledge from people. Some people will say that the procedures should be up to date, but they are not.

Do the success criteria have to stay fixed?

The criteria may change later and will be subjective, but they show where the team is heading.

Does an iterate result mean the tests failed?

An iterate result is a good test result. It shows what still needs work and what needs rethinking so that the solution meets expectations.