The AI agent factory: how to start and reach production

For:
CTOs, CIOs, architects and IT managers planning a platform for AI agents
Reading time:
9 min
Episode:
25 min

This episode is in Polish. Full Polish version with transcript

After you click, the video loads from YouTube (Google). Google may store data on your device and process it in the USA as well. More (PDF, in Polish) · Watch on YouTube

In brief

You want to roll out AI agents in your company, and you must decide what to build them on and how to give them safe access to data and systems. An agent factory consists of a ready-made platform (if you already have a cloud, your provider's service), one shared gateway to company systems and a closed test space from which an agent goes to production like any other application. You start with a simple case where you know what the agent receives and what it must return.

Key takeaways

  • An AI agent is the same LLM-based process as a chatbot, except that it works longer on data and makes its own decisions.
  • At the start, one orchestrator agent with small, specialized agents usually handles a business process, because a long context is expensive and the agent then follows instructions less well.
  • Integrations exposed through one gateway serve every agent you build later and cost less to maintain.
  • When you already have a cloud, you use your provider's ready-made service, and you can later swap the model and the platform without rebuilding the integrations.
  • The sandbox enforces a secure configuration. From there, an agent kept as code in a repository goes through CI/CD, first to a non-production environment and, after tests, to production.
  • The first stage is usually the hardest and the longest: choosing the case and working out which systems the agent must integrate with.

An AI agent decides on its own, within the limits you set

An AI agent is a process that uses a large language model (LLM), works longer on data and makes its own decisions within the limits you set. Some definitions describe an agent as an independent, living entity. In practice, it is the same chatbot or language model plugged into a process, for example in an application, just in a different form.

A runtime framework runs the agent. It is either a ready-made cloud service or code in Python or another language with a framework that connects to an LLM, for example one from OpenAI. Under the hood, an agent is a loop that invokes itself based on the input data.

Agents at Protopia’s clients handle documents, reports and processes with no fixed path

Document processing keeps coming up with Protopia’s clients. The agent receives a document after OCR, for example from an email in a mailbox or from a system where an employee uploaded it. It recognizes the PESEL (the Polish personal ID number), the customer number and other identifying data, tries to gather information from various company systems and supports the employee. A typical document is a complaint, probably the most common case of this kind. An example from one client: someone disputes an invoice, the agent finds that the discount was calculated incorrectly and prepares a draft email for the sales rep that explains it.

The second case is reports built from data spread across systems. A few times, clients needed a merged view of a customer across several systems (CRM, ERP, help desk or another domain-specific system), to stop jumping between them or to find inconsistencies in the data.

The third case is processes whose path is hard to set in advance. The agent can then, depending on the situation, collect data or take actions in systems. If something is missing, it fills in the data in several places.

At the start, one orchestrator and small agents usually handle a business process

The starting setup is one orchestrator agent per business process, with small agents attached to it that take actions in individual systems. Technically, this scenario is called connected agents. The orchestrator acts like a conductor or a team leader: it distributes the work and combines the results. Other scenarios exist too, but they are not recommended at the start, because they can be a nightmare to test and the results can be very disappointing.

Clients who build agents on their own first try to make a super agent and connect everything to it. The opposite trend resembles microservices: a very large number of small agents.

The reason is the context window. A context window is the limit on how much text an LLM accepts as input. In automation, a long context turns out to be expensive, and the agent does not follow instructions. That is why the context is kept to a minimum and tasks are split among smaller, specialized agents.

Besides the budget, the technology itself is a limit for now. An agent does not always follow its instructions. Anthropic says its goal for the coming months and years is for its language model to follow the instructions it gets and the execution order exactly.

One integration gateway exposes systems to all agents

A system connected once to the integration gateway is then available to every agent the company builds. Protopia recommends Azure API Management to clients as the central gateway that exposes various systems to agents in a uniform way. The language model and the runtime code are universal, and the company must standardize how it integrates with its own systems.

An orchestrator agent handles a business process and distributes work among small agents. The small agents call company systems through one integration gateway, Azure API Management, which holds the call descriptions. Behind the gateway are CRM, ERP, help desk and an HR system. A next agent uses the same integrations, connected once.
Diagram: agents reach company systems through one gateway

Łukasz Kałużny: “Because the heart of it, contrary to appearances, is not the agent, but the way you deliver knowledge to it and the ability to take those actions through integration with other systems.” (translated)

An agent’s integration with a system is a natural-language description of the calls, optimized for the LLM. You write the description the way you would explain the call to a person: what data the method reads or writes and what you must pass to it. Example: this call returns customer data by PESEL number or by customer ID.

A team that writes agent code without a ready-made platform can try to build the integration itself. At a larger scale, source systems may need changes before they can be connected. A larger organization also wants to reuse an integration it built once in other places, as cheaply as possible. It is possible without a gateway, but in Protopia’s experience it is more expensive and slower.

The gateway also holds the central documentation of the descriptions, for example: in the HR system, employee data is here and their leave is there. Other people who build and test agents later use these descriptions. The gateway makes integrations cheaper to maintain later.

The agent factory works like platform engineering for agents

An agent factory is a platform that lets you move quickly from testing an agent to running it in production. You can look at it as platform engineering: the path goes through a proof of concept (PoC) and a proof of value and, if everything is fine, ends with the agent released for use.

At a small scale, a data scientist or a developer writes an agent in Python with frameworks from Microsoft, OpenAI or Google. That is how a PoC starts and the first agents appear. Protopia’s clients, often companies in finance, in regulated industries or larger organizations, want to bring order to this process.

Tests usually run on data as close as possible to the data in production systems, to see the agent’s value.

With a cloud in place, choose your provider’s ready-made PaaS service

If your organization already has a cloud in place, your provider’s ready-made platform as a service (PaaS) is usually the best and most cost-effective choice, especially when you want to show value and start testing. Every hyperscaler has an equivalent of Azure AI Foundry. From the point of view of Protopia, which works with Azure every day, AI Foundry is “good enough”.

Platform options for agents
Option What it means
A cloud provider’s PaaS service, e.g. Azure AI Foundry For organizations with a cloud in place. It has some limits, but you can use it right away once the platform is rolled out.
A ready-made open-source agent platform, e.g. Dify An alternative when you have no cloud. You still have to supply the language model somehow.
Your own platform built on open source Probably several months before it fully works. That time does not go into building agents.

A ready-made service limits the amount of code and the number of frameworks. The team focuses on two things: designing agents, and connecting systems and data sources along with testing. The work shifts from writing code to writing prompts. Do not bury yourself in open-source frameworks and gorgeous diagrams from LinkedIn.

There is no major vendor lock-in when you move away from Azure AI Foundry. You can swap the language model. System prompts, the agent’s instructions, need a check and an adjustment after a model change, because every model has its nuances, but this is more like polishing. Every framework, open source included, consumes the integrations in Azure API Management in exactly the same way. The exit strategy can therefore be described as copy-paste plus tests.

After leaving Azure AI Foundry for another framework, open source included, you can swap the language model. You check the system prompts and adjust them to the new model, which is more like polishing. Every framework consumes the integrations exposed in Azure API Management the same way. The exit strategy is copy-paste plus tests.
Diagram: what changes when you leave Azure AI Foundry

AI Sandbox enforces a secure configuration on every experiment

An AI Sandbox is a prepared space, for example in Azure, where the team experiments with agents safely. That is Protopia’s name for this concept in its work with clients. The preparation covers security controls, the platform configuration, access management and control of inbound and outbound network traffic.

The cloud platform’s mechanisms let you enforce a configuration in which neither a developer nor a data scientist can click an “expose to the internet” button.

The sandbox also limits configuration changes, so the team can focus on testing the agent. At one client, the team spent time working out how to connect to an AI service from an open-source library and sign in securely.

Łukasz Kałużny: “So the sandbox has two goals. One is security and giving room to experiment, but on the other hand also forcing a focus on business value, not on playing with technology.” (translated)

An agent goes to production like any other application

An agent moves from the sandbox to production just like another application on Kubernetes. When it rolls out the platform, Protopia prepares a set of deployment scripts for the whole process: the agent goes from the sandbox to a non-production environment and, after tests, to production. The platform is built in this order: first a secure AI Sandbox from PaaS or ready-made services, then the process of moving to the environments where the agent runs.

The AI Sandbox is a shared space, divided much like namespaces in Kubernetes. The same development practices apply in it as for regular applications. The agent lives as code in a repository and is deployed through CI/CD, with no magic and no special clicking. The sandbox has the same configuration that the agent later gets in production.

The agent is built in the AI Sandbox. Its code goes to a repository as agent as code, and CI/CD deploys it to a non-production environment. After tests, the agent goes to production. All environments have the same configuration, as with a regular application.
Diagram: the agent's path from AI Sandbox to production

The first stage is a simple case with a defined input and output

To start, choose simple cases for testing agents. Next, define what goes into the agent, for example a piece of information, a request or a document, and what the result must be. Then work out which systems you must integrate with the agent. Only with this knowledge do you move on to implementation. This stage is usually the hardest and the longest.

Speakers

  • Łukasz Kałużny

    Łukasz Kałużny

    Founder, Managing Partner, Technology Advisor.

    Łukasz Kałużny co-founded Protopia and is a Microsoft MVP in the Microsoft Foundry category. He has co-hosted Patoarchitekci since 2019, talking about IT architecture, GenAI and AI agents without the marketing spin.

    All posts by this author
  • Mikołaj Szczerbicki

    Mikołaj Szczerbicki

    Head of Sales & Business Development.

    Mikołaj Szczerbicki is Head of Sales & Business Development at Protopia and co-hosts the Powered by Protopia podcast. He scopes and prices projects, so he asks about the cost, risk and timeline of AI, Azure and Kubernetes work.

    All posts by this author

FAQ

Does Azure meet the regulatory requirements of Polish companies?

In most cases, yes: Azure covers the regulatory side for Polish companies.

Does the AI Sandbox prevent a data leak caused by sharing a cloud resource?

In the AI Sandbox, a developer will not cause the kind of leak you read about on IT news sites, where someone shared something in the cloud.

Does a large context window let you build one super agent?

No. Claims that the context window is almost unlimited can be marketing talk. Agents are dedicated to one business case, process or problem.