AI agent factory: a shared platform in a large organization

For:
CTOs, CIOs, architects and IT managers who plan a growing number of AI agents in a large organization
Reading time:
12 min
Episode:
36 min

This episode is in Polish. Full Polish version with transcript

After you click, the video loads from YouTube (Google). Google may store data on your device and process it in the USA as well. More (PDF, in Polish) · Watch on YouTube

In brief

If your company is going to build more and more AI agents, you have to decide whether to build each one separately or to set up a shared platform. An agent factory is that kind of platform, built on Azure AI Foundry, API Management and Logic Apps: separate environments for experiments, tests and production, security approved once, and integrations you connect once and share with later agents. A team with an approved idea gets a secured environment quickly.

Key takeaways

  • When the company strategy expects more and more agents and APIs, a shared platform is probably the only sensible direction, while the first agent for the board can likely be built without it.
  • AI Sandbox is a standardized, secured place to test a business hypothesis, and the agent moves from there through the pre-production environment to production; it is deployed from a Git repository through CI/CD like any other software.
  • In the enterprise variant, thread data sits in resources in the organization's own subscription, encrypted with its keys.
  • API Management exposes company systems to agents on one facade, with a description the model understands, so you do not have to write custom MCP gateways or change the source systems.
  • When a process is truly deterministic, Logic Apps very often turns out more predictable than custom code or agents, and the agent comes in where generative analysis is needed.
  • Before platform work starts, Protopia tries to hold a workshop that picks, from the client's business cases, the ones that fit the platform well.

A platform makes sense when agents and APIs keep growing

If the company strategy expects more and more agents and more and more APIs, a shared platform is probably the only sensible direction. The first agent, meant to show the board a short-term win, can likely be built in a couple of weeks.

Here, an agent factory means a rollout in a large organization, including one constrained by regulations.

The team does not set up the whole infrastructure every time, so it starts using agents in a standardized way sooner. Once the platform is approved, there is no need to open network access and recertify the environment for each business case. Security and consistency, for example logs, are solved once.

A shared platform also keeps knowledge in one place. At several clients, teams test agents instead of yet more open-source frameworks, which keep multiplying and keep changing how they work.

An agent moves through three environments

The platform has three environments: AI Sandbox, a pre-production environment and production. AI Sandbox is the place for safe experiments and innovation. In many companies a sandbox means a free-for-all where users do whatever they want. Here it is a secured, shared environment, standardized and subject to corporate rules, possibly with access to production data. It is hard to build agents, and AI in general, on anonymized and synthetic data.

In the sandbox, the team checks the proof of value. Proof of value is a test of the business hypothesis: whether the scenario will run on agents and work. Writing and testing code takes a back seat.

A scenario that proves itself moves to the pre-production environment and from then on follows a production lifecycle. There the team checks that all deployments and integrations work and that the agent is fully integrated and ready for production.

Each environment has four elements: access channels, the agent runtime, the integration part, and the operations and maintenance part.

Architecture of one agent factory environment in four parts. Access channels, that is a button in a system, an event or queue, and a chat on a website or in Teams, call the agent in an Azure AI Foundry project. The agent uses APIs and MCP through API Management, which exposes backend systems and external MCP on one facade with a description the agent understands. Logic Apps calls APIs in API Management and can also call the agent. The operations and maintenance part covers the Git repository, CI/CD scripts, Application Insights and content filtering.
Diagram: the elements of one agent factory environment

The access channel is separate from the agent

The agent does not depend on how you reach it. It can be standalone, part of a mobile app, a chat on a website or a background process. The platform lets you call the agent in any way, and the call itself is an external integration:

  • synchronously from a business system, for example with a “generate a report for me” button;
  • asynchronously through an event, for example a message, a queue or an email to a monitored mailbox;
  • through a chat on a website, in Teams or in another application.

The calling method stays outside the platform so that it does not impose a solution. A proof of value can show that the agent needs to be called in a completely different way. Keeping the channel separate is meant to avoid hard rewiring and a later rebuild.

One of the recent client cases is an analysis of a counterparty’s financial results. In tests, an analyst pastes the results into a chat window and checks whether the agent answers as expected. The target is a “Run the analysis for me” button in the system that starts the agent. The agent synchronously queries the backends and APIs, collects and analyzes the data, and returns only the final result.

Agents run in Azure AI Foundry projects

The agent runtime is Azure AI Foundry, in the Microsoft ecosystem.

The unit of work in AI Foundry is a project, the equivalent of an initiative. Technically it resembles a namespace in Kubernetes or a resource group in Azure. One project can hold many agents built by one team, and agents in a project can talk to each other.

A project uses models from a defined list: OpenAI and the supported open-source models from the marketplace. Competing vendors’ models are not on the list. Protopia works with clients on OpenAI 99% of the time.

An orchestrator agent connects small agents

Because of how language models work, Protopia usually builds small agents, each for a specific task or part of a business process. Above them runs an orchestrator agent that calls the smaller agents. The orchestrator can get a more expensive reasoning model, and the small agents a cheaper one, for example GPT-4o mini.

In most cases Protopia uses the connected agents scenario instead of a chain of fully independent agents: a central agent plus small agents, with the steps orchestrated deliberately. In AI Foundry you literally specify which agent connects to which.

Connected agents in one Azure AI Foundry project. The access channel passes the thread to the orchestrator agent, which has a more expensive model with reasoning. Through a handoff, the orchestrator passes the thread to small agents with a cheaper model, for example GPT-4o mini, each for a specific task, and waits for their answer.
Diagram: an orchestrator agent and small agents in one project

A thread is an agent’s conversation about one matter. A handoff is a temporary transfer of the thread to another agent: the agent waits for the answer and continues processing.

A2A (agent to agent) is an open-source protocol for communication between agents, started by Google. It is available in AI Foundry, but it is not worth tying yourself closely to it, because agents usually connect within one project.

In the enterprise variant, thread data stays in your subscription

Thread history lets you later resume an asynchronous thread and collect its results. In a less enterprise-grade variant, when you do not need this, the service keeps the data itself. In the enterprise variant, which Protopia usually deploys, the data goes to separate services that you have to deploy:

  • storage for files;
  • Cosmos DB for conversations and agent data;
  • Azure AI Search, with inverted indexes and a vector database for text or semantic search across texts, documents and attachments, the basis of RAG.

Bring your own resources is a model in which the organization brings its own storage, Cosmos DB and Azure AI Search, deployed in its own subscription. You see them and control them. The resources are deployed almost entirely on the organization’s terms, so they can match its internal regulatory requirements, for example secure communication and independent logs.

Conversation history can contain sensitive and protected data, for example in healthcare, pharma or banking. That is why the resources are encrypted with the organization’s own keys (customer-managed keys, CMK), in line with the rules the organization has set.

API Management is the central integration point

In this architecture, API Management is the heart of the platform, because an agent without APIs is just a chat. It exposes the organization’s APIs consistently, on one facade, so that agents can consume them easily.

A business system connected once through an API is available to every agent you connect it to. An agent does not automatically get access to everything: authentication and authorization decide which API a given agent talks to.

The API has to be exposed differently from the backend system. Systems delivered by outside vendors or by an internal department are often poorly described. An agent does not know from memory what a given method does, and the method description alone does not guarantee it. That is why you add a description in API Management that the agent understands, so it knows which endpoint solves its problem. This works for both OpenAPI Schema and MCP (Model Context Protocol): every API, call and expected response is described in natural language tailored to the model.

A chat model writes the description, with a system prompt that tells it the text is for its own later use.

API Management also translates on the fly, for example it exposes a SOAP API as REST. With API Management you do not need to write your own MCP gateway for a system or, worse, modify the source system to expose endpoints for agents.

API Management has two new capabilities. Any existing API, including one that systems outside the agent world may already use, can be turned into MCP with practically one switch. Recently you can also move external MCP servers onto the facade: third-party ones, or ones someone in the company built as MCP from the start, without an API. The organization controls them like its own, with extra security, policies and orchestration.

Logic Apps handles the predictable steps

If a scenario is truly deterministic, meaning you know the process and how to connect it, Logic Apps very often turns out more precise and predictable, and also simpler to build, than custom code or agents.

Teams often end up with narrowly specialized, low-level APIs and an API Management instance that exposes them, but no glue to orchestrate them, for example: call an API, get the response, send three emails or a text message, wait for approval and confirm on another endpoint. In this architecture the glue is Logic Apps, Microsoft’s no-code tool, which you can also manage from code. Logic Apps together with API Management make up iPaaS (Integration PaaS).

You can build a Logic Apps flow by clicking, as in Power Apps, but there are more system and business connectors: for Oracle, SAP, CRM or Salesforce. A flow can also call an API in API Management or an agent. There are, roughly, hundreds of ready-made integrations and actions. Control flow covers conditions (IF), loops (FOR), error handling and resuming integrations.

The agent comes in only where you need generative analysis, something new created, or facts combined on the fly.

In some scenarios the entry point is a workflow in Logic Apps. Marek Grabarz: “it’s not that the agent triggers the workflow, it’s the workflow that calls the agent.” (translated) Based on the agent’s result, the workflow then does something predictable. The access channel is then the integration in iPaaS.

Agents are deployed like code

An agent is a piece of software like any other, so the operations and maintenance part rests on the DevOps approach. There are no special rules; only some technical details change.

Agents live in a Git repository: GitHub, Bitbucket, GitLab or another one the organization already has. Prepared CI/CD scripts deploy the agents to the next environments.

The agent file in the repository holds the agent definition with the system prompt, plus the connected integrations and other agents. You can treat it like an infrastructure as code manifest. In this case these are Python scripts, so the file is also the agent definition in Python code. Once the platform runs, a deployment often comes down to copy and paste, or help from GitHub Copilot or another coding agent.

The agent file in a Git repository holds the agent definition with the system prompt, the connected integrations and the connected other agents. CI/CD scripts deploy it to three environments. In AI Sandbox the team tests the business hypothesis. A scenario that proved itself moves to the pre-production environment, where the team checks deployments and integrations, and then to production.
Diagram: an agent from a Git repository to the next environments

Observability shows how each thread went

When something goes wrong in an agent, or someone reports a wrong answer, the team checks the thread. AI Foundry stores the full conversation threads of every agent in a database until you delete them: the input, all the tools called, the content of requests to other systems, integrations and agents, and the responses. The thread is stateful. In Azure AI Foundry Studio you can find it, go back to it, view it in the playground and check what happened.

You turn on observability in a project with almost no deployment work. Application Insights, a standard APM tool from Microsoft, tracks the agent lifecycle: questions, answers and calls, across the whole conversation thread from a chat or the call thread from another system.

AI Foundry also has a quality evaluation area: another model judges whether the whole flow went well. Response evaluation metrics are starting to appear, but the market is still at an early stage, because one non-deterministic thing is measured by another.

Content filtering works from the start and protects against users putting something dangerous into the agent. The organization decides which content it allows. Some companies may want to loosen the filters, for example in e-commerce with adult content.

Usage and complaints measure the platform

The first measure of the platform and the agent use cases is whether anyone uses them after rollout.

The second measure applies to each agent separately: how many complaints and bug reports there are, for example about answer quality, relative to the number of uses. A good result is usage that does not drop and no complaints about quality.

A workshop comes first where possible, and the sandbox should be ready the next business day

Before work on the platform starts, Protopia tries to organize a workshop. The client prepares business cases for it. Business architects are invited, not necessarily operations people, because it is not that stage yet. At the workshop, the team points out the cases that fit the platform well and will integrate well. With 5, 10 or 15 such examples, the incentive to build the platform is much bigger. Łukasz Kałużny: “it’s an important point, so that we don’t build a platform for the sake of building a platform.” (translated)

Once an initiative is approved for tests in AI Sandbox, the team should have access the next business day. The platform team can, for example, provide such an environment in a dozen or so minutes. Self-service is possible, but it is not the default here. The time counts from approval and does not include the company’s organizational processes.

Speakers

  • Łukasz Kałużny

    Łukasz Kałużny

    Founder, Managing Partner, Technology Advisor.

    Łukasz Kałużny co-founded Protopia and is a Microsoft MVP in the Microsoft Foundry category. He has co-hosted Patoarchitekci since 2019, talking about IT architecture, GenAI and AI agents without the marketing spin.

    All posts by this author
  • Marek Grabarz

    Marek Grabarz

    Founder, Managing Partner, Technology Advisor.

    Marek Grabarz co-founded Protopia and is a Microsoft MVP in the Microsoft Azure category. He co-hosts the Powered by Protopia podcast, talking with IT leaders about API Management, integrations and secure AI adoption.

    All posts by this author

FAQ

Does AI Foundry also cover Microsoft's older AI services, such as speech to text?

In theory, AI Foundry also gives access to AI services from before OpenAI, for example speech to text.

Does Protopia work only with Microsoft technologies?

Protopia also has cases with other technologies, but it specializes mainly in Microsoft solutions.

Is it worth building each agent as a separate initiative?

In Protopia's experience, building single agents as separate initiatives makes no sense, apart from a first agent built as a quick win for the board.