<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Protopia — publications</title><description>Episode summaries, lessons from projects and articles on what we implement: AI agents, APIs, observability, FinOps, Azure and Kubernetes.</description><link>https://protopia.tech/</link><language>en</language><atom:link href="https://protopia.tech/en/publications/feed.xml" rel="self" type="application/rss+xml"/><item><title>Virtual data center: offloading your own data center to Azure virtual machines</title><link>https://protopia.tech/en/publications/virtual-data-center-azure/</link><guid isPermaLink="true">https://protopia.tech/en/publications/virtual-data-center-azure/</guid><description>Quotes for server room hardware stay valid only briefly. Instead of expanding your server room right away, you can move some less critical machines to Azure.</description><pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;When your on-premises data center runs short of resources and the cost of expanding it is uncertain, you can offload it by moving some of your virtual machines to Azure. A virtual data center is a standardized, shared space in Azure where machines are built from templates that meet the same requirements as your data center. Development and test environments are probably a big win: when they shut down automatically outside working hours, you pay for their compute only while they run.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;A virtual data center treats Azure as a place for virtual machines (IaaS), not for PaaS or SaaS.&lt;/li&gt;&lt;li&gt;You split systems by business risk: less critical ones go to Azure, and critical ones stay in the on-premises data center.&lt;/li&gt;&lt;li&gt;Machines that run around the clock can use a savings plan or a reservation, depending on your needs. Windows Server and SQL Server licenses with active Software Assurance can, depending on their type, be used in the cloud through Azure Hybrid Benefit.&lt;/li&gt;&lt;li&gt;Decide how to respond to oversized machines: an alert to the owner, who downsizes them, or a more aggressive option in which the platform resizes them automatically or after prior notice.&lt;/li&gt;&lt;li&gt;You move existing machines with Azure Migrate, which requires downtime, or restore them from backup in Azure, if your tool supports this and you have a license for it.&lt;/li&gt;&lt;li&gt;The first step is a set of workshops: requirements, the network concept, and a check of whether the company already has Azure.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;what-a-virtual-data-center-is&quot;&gt;What a virtual data center is&lt;/h2&gt;
&lt;p&gt;The speakers point out that quotes for expanding an on-premises data center stay valid only briefly, and that prices have risen sharply in some places. Timing can be a problem too: whether an expansion can be done this year depends on the size of the order.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A virtual data center is a standardized landing zone in Azure where virtual machines are created automatically on shared subscriptions.&lt;/strong&gt; The machines share subscription resources in one place instead of being scattered, which is meant to lower costs. The team gets simple guidelines on how to add another machine or move one from the on-premises data center, so it does not choose a service and a resource group for each machine.&lt;/p&gt;
&lt;p&gt;The platform mirrors the requirements you apply on premises, including those from your security policies. It has several layers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a set of subscriptions in the form of a landing zone,&lt;/li&gt;
&lt;li&gt;a network that matches your on-premises segmentation, that is, the split into VLANs and zones,&lt;/li&gt;
&lt;li&gt;governance: security policies, tags, budgets and alerts,&lt;/li&gt;
&lt;li&gt;infrastructure as code (scripts or Terraform, depending on preference) and machine templates that install the same agents as on premises, e.g. for scanning, DLP or monitoring,&lt;/li&gt;
&lt;li&gt;identity: the access model,&lt;/li&gt;
&lt;li&gt;backup: native Azure Backup, Veeam (Protopia’s preference) or another vendor’s tool you already use, if it supports Azure.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The platform can also connect to a ticketing system. Teams then order machines themselves through forms (self-service).&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-vdc-architecture-desktop.Bn_dbqbr.webp&quot; alt=&quot;Diagram: less critical and nonproduction machines move from the on-premises data center to a virtual data center in Azure. The virtual data center is a landing zone with shared subscriptions and five layers: a network that matches the on-premises segmentation into VLANs and zones, governance with policies, tags, budgets and alerts, infrastructure as code with machine templates, identity with the access model, and backup. Optionally, machines are ordered through forms in a ticketing system.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: virtual data center architecture.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;which-systems-to-move-to-the-cloud&quot;&gt;Which systems to move to the cloud&lt;/h2&gt;
&lt;p&gt;You move less critical systems to the virtual data center: applications of lower business importance, back-office applications and nonproduction environments, meaning development and test. In this scenario the cloud is an extra source of compute that frees up resources in the on-premises data center.&lt;/p&gt;
&lt;p&gt;The selection criterion is business risk. Łukasz Kałużny: “We don’t migrate everything, only what may be getting in our way, or what could free up local resources safely for the business, without much risk.” (translated)&lt;/p&gt;
&lt;h2 id=&quot;how-to-match-a-pricing-mechanism-to-a-machine&quot;&gt;How to match a pricing mechanism to a machine&lt;/h2&gt;
&lt;p&gt;You pay for the disk separately, even when the machine is off. The amounts in this section cover compute and do not include the disk.&lt;/p&gt;

&lt;div class=&quot;table-scroll&quot; role=&quot;region&quot; tabindex=&quot;0&quot; aria-label=&quot;Mechanisms that lower the cost of virtual machines in Azure&quot;&gt;&lt;table&gt;&lt;caption&gt;Mechanisms that lower the cost of virtual machines in Azure&lt;/caption&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;For which machines&lt;/th&gt;
&lt;th&gt;Savings or effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Automatic shutdown outside working hours (the machine runs, e.g. 8 a.m.–6 p.m.)&lt;/td&gt;
&lt;td&gt;Nonproduction: development, test&lt;/td&gt;
&lt;td&gt;71% off the list price&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Savings plan or reservation&lt;/td&gt;
&lt;td&gt;Noncritical, running 24 hours a day&lt;/td&gt;
&lt;td&gt;45–61%, depending on the configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Azure Hybrid Benefit&lt;/td&gt;
&lt;td&gt;Windows Server and SQL Server with an active Software Assurance agreement&lt;/td&gt;
&lt;td&gt;A license bought for the on-premises environment can be used in the cloud, to an extent that depends on its type&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;A machine that is on only during working hours runs about 210 hours a month instead of an average of 730. Łukasz Kałużny: “We have a very mistaken belief that it has to run 24 hours a day.” (translated)&lt;/p&gt;
&lt;p&gt;An example from the price list: compute for a popular machine with an Intel or AMD processor costs about 160–170 USD a month. With automatic shutdown, its cost drops to 46–51 USD.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-pricing-mechanism-desktop.3LTS1vcz.webp&quot; alt=&quot;Decision tree. If a machine does not have to run around the clock, it shuts down outside working hours and runs, e.g. 8 a.m.–6 p.m.: it runs about 210 hours instead of 730, and the cost of its compute is 71% lower than the list price. If it has to run around the clock, a savings plan or a reservation saves 45–61%. Machines with Windows Server or SQL Server covered by Software Assurance can also use their own license through Azure Hybrid Benefit, to an extent that depends on the license type. The disk is billed separately.&quot; width=&quot;736&quot; height=&quot;552&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: which pricing mechanism for which machine.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;how-to-keep-costs-in-check-after-launch&quot;&gt;How to keep costs in check after launch&lt;/h2&gt;
&lt;p&gt;Two layers keep costs in check, and both are built into the platform from day one. The first is the standard FinOps tooling in Azure: budgets and Azure Advisor. The second is a set of alerts and scripts that Protopia adds to the platform. The scripts regularly look for unused and oversized resources.&lt;/p&gt;
&lt;p&gt;The speakers often come across machines sized for headroom, with unused CPU. In the milder response, an alert goes to the machine’s owner, who downsizes it. In the more aggressive one, if the organization agrees to it, the platform regularly checks usage metrics, e.g. every week. It then resizes machines automatically or after prior notice, according to the agreed policy.&lt;/p&gt;
&lt;p&gt;An example policy: a machine gets 2 cores instead of 4, and the RAM stays the same. For one machine that runs around the clock, the cost drops from 177 to 117 USD.&lt;/p&gt;
&lt;h2 id=&quot;how-to-move-existing-machines&quot;&gt;How to move existing machines&lt;/h2&gt;
&lt;p&gt;You deploy new systems with scripts and templates as soon as the platform is built. Existing machines move in one of two ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Azure Migrate assesses machines from VMware, Hyper-V and other platforms, then converts and moves them. The migration requires downtime because the machine has to be transferred.&lt;/li&gt;
&lt;li&gt;A backup tool that integrates with Azure, e.g. Veeam, restores the machine from an existing backup directly in the virtual data center. The tool must support this scenario, and you must have a license for it. This kind of migration is faster or, depending on the approach, runs without downtime.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-migration-paths-desktop.nm_ADyBL.webp&quot; alt=&quot;Diagram of migration to a virtual data center. Machines from VMware, Hyper-V and other platforms reach Azure through Azure Migrate, which assesses, converts and moves them, and this requires downtime. The second path restores the machine from an existing backup, e.g. in Veeam, faster or with no downtime, if the tool supports it and you have a license. New systems are built right away from scripts and templates.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: two paths for migrating existing machines.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;how-the-rollout-works-and-how-long-it-takes&quot;&gt;How the rollout works and how long it takes&lt;/h2&gt;
&lt;p&gt;Protopia estimates 4–8 weeks to build a working platform.&lt;/p&gt;
&lt;p&gt;The platform build has three stages:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Discovery in workshops: whether the company already has Azure, what the network concept looks like, what the requirements are and which approach to choose.&lt;/li&gt;
&lt;li&gt;Design: adapting Protopia’s standardized infrastructure as code to the company’s needs.&lt;/li&gt;
&lt;li&gt;Setup and configuration, the main phase: the subscription, the network, tests of the scripts and machine templates, and the rollout of the governance, identity and backup layers described above.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;After these stages, you deploy new virtual machines on the platform. Migrations are a separate stage.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Pilot migrations of the first machines.&lt;/li&gt;
&lt;li&gt;After successful pilots, migration of less critical systems at a larger scale.&lt;/li&gt;
&lt;/ol&gt;</content:encoded><dc:creator>Łukasz Kałużny, Mikołaj Szczerbicki</dc:creator><category>Cloud</category></item><item><title>Rolling out AI in the SDLC: process standard, measurement and pilot</title><link>https://protopia.tech/en/publications/ai-in-the-sdlc-rollout/</link><guid isPermaLink="true">https://protopia.tech/en/publications/ai-in-the-sdlc-rollout/</guid><description>Before you roll out AI in your development teams, prepare the process and measure the starting state. After the pilot, compare the result with that starting state to judge whether the rollout had an effect.</description><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;Your teams already use AI tools and you see no effect beyond the cost, or you are deciding which processes get AI first and how to measure success. The gain from AI shows only after you redesign the software development life cycle: a specification before code, a shared standard for working with the agent and automatic feedback, all tested in a pilot on a real application. Measure the current state before the pilot, because only the comparison with it shows whether to scale or to fix the foundations first.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;Speeding up only one stage clogs the process somewhere else, so you standardize the whole development cycle: the specification before code, the work with the agent, skills, documentation in the repository and an automatic feedback loop, and you choose the tools last.&lt;/li&gt;&lt;li&gt;The rollout starts with an assessment of process readiness: CI/CD, release cadence, testability and deployment automation.&lt;/li&gt;&lt;li&gt;The pilot works best with a team that volunteered, on an application the team really develops and maintains.&lt;/li&gt;&lt;li&gt;The measurement before the pilot is a diagnostic with no set targets: you measure delivery, flow and quality, not tokens, lines of code or the share of AI-generated code.&lt;/li&gt;&lt;li&gt;The pilot can reveal a bottleneck outside AI, such as a queue at acceptance, and that is also a result: the process gets measured and improved.&lt;/li&gt;&lt;li&gt;Scaling does not have to cover the whole organization: depending on company size and criticality, part of it or a few development teams are sometimes enough.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;handing-out-ai-tools-alone-brings-no-visible-speedup&quot;&gt;Handing out AI tools alone brings no visible speedup&lt;/h2&gt;
&lt;p&gt;Handing teams tools such as GitHub Copilot, OpenAI Codex or Anthropic’s Claude Code does not by itself speed up the work in any real way, and the only visible effect is the cost. The speakers hear this in the market and when they exchange experience with peers.&lt;/p&gt;
&lt;p&gt;One of the main topics clients raise is AI wired into processes that are neither measured nor standardized. In conversations with the management of several organizations, two situations come up. In the first, the company is still working out which processes get AI first and how to measure success. In the second, developers start using AI on their own and the manager agrees. Paid, more expensive and supposedly more secure versions of chat apps then appear, next to custom interfaces wired to the API. None of it is standardized, and at some point the organization sees neither an effect nor a gain.&lt;/p&gt;
&lt;h2 id=&quot;speeding-up-one-stage-clogs-the-process-elsewhere&quot;&gt;Speeding up one stage clogs the process elsewhere&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The SDLC (software development life cycle) is the whole path from gathering a business need to working, deployed software.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Reports from companies that focused on writing code alone speak of gains of around 4–5%, sometimes less, depending on whom you ask. In some very small organizations you hear of much higher numbers, because they have no organizational overhead. In a small product, such as a small SaaS, a developer can write a change, test it and deploy it within a day. In the large organizations Protopia works with, that is very often not possible.&lt;/p&gt;
&lt;p&gt;Łukasz Kałużny: “if we speed up just one stage, then realistically it will get clogged somewhere” (translated)&lt;/p&gt;
&lt;p&gt;That is why, when you roll out AI in the SDLC, you standardize the whole process.&lt;/p&gt;
&lt;h2 id=&quot;the-gain-from-ai-comes-only-after-a-process-redesign&quot;&gt;The gain from AI comes only after a process redesign&lt;/h2&gt;
&lt;p&gt;In 1987 the Nobel laureate Solow observed that you could see computers everywhere, but nowhere could you see a rise in value from them. Organizations waited about 15 years for a return on their investment in computers. With the electric motor in factories it took about 40 years, because processes had to adapt to the new technology.&lt;/p&gt;
&lt;p&gt;AI follows the same path: the gain will show only after the process and the approach are redesigned and the organization matures. The payback cycle will probably be much faster than with computers or the electric motor.&lt;/p&gt;
&lt;p&gt;If you want to control this return, you have to measure it, that is, assess the state before and after you improve the process. You plan what to measure and at which points before the rollout starts, whether it runs all at once or in stages.&lt;/p&gt;
&lt;p&gt;The rollout starts from the current process and evolves it gradually: it adapts what exists to working with AI tools. The goal is a repeatable, scalable effect for the whole organization. The success of one team may sometimes be enough, but then you have to check whether you can repeat it later.&lt;/p&gt;
&lt;p&gt;The 2025 DevOps Research and Assessment (DORA) report describes AI as an amplifier that lifts mature teams and intensifies the dysfunctions of weak ones.&lt;/p&gt;
&lt;h2 id=&quot;the-standard-covers-the-specification-agent-work-and-the-feedback-loop&quot;&gt;The standard covers the specification, agent work and the feedback loop&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Spec Driven Development is a method in which the team gathers requirements in a repeatable way, generates a specification from them, then a plan for the agent’s work on a given feature, and only at the end the code.&lt;/strong&gt; It is the opposite of vibe coding. It builds on open standards such as GitHub Spec Kit or OpenSpec, or on a lighter version prepared for a specific organization.&lt;/p&gt;
&lt;p&gt;The goal is one consistent standard for working with the agent across the organization’s repositories. The team gets a base it can reuse in a new project or in a project that the standard covers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Skills are sets of commands and prompts that tell the agent how to do something.&lt;/strong&gt; They can be company-wide, for example “review the requirements against the global policy”, or project-specific. One of the popular skills that Protopia rolls out at clients analyzes a bug report from production to locate the bug faster.&lt;/p&gt;
&lt;p&gt;Documentation goes into the code repository, stays up to date and is readable by both the agent and a human.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Harness and guardrails are the whole environment around the agent that gives it a short feedback loop.&lt;/strong&gt; Just as a developer uses unit tests, the agent gets automatic formatting, linting and testing, blocking checks and quality gates. In some places the gate is test coverage or cognitive complexity (cognitive load), a measure of how hard the code is to understand and how complex it is. You choose the metrics so that the code is easier to remove or maintain in the future, and you wire them in automatically. This way the agent gets feedback as fast as possible and works in the direction you set.&lt;/p&gt;
&lt;p&gt;You choose the tools last, because many of these practices carry over between models and tools. Tools currently change every quarter, and some say the trends change even every month. That is why you choose them together with the models to fit your tech stack and your compliance requirements.&lt;/p&gt;
&lt;p&gt;The standard covers the whole SDLC. Spec Driven Development handles requirements. The agent’s tools, together with Spec Driven Development, handle code, tests and review. Maintenance relies on integration and unit tests that run locally and in CI/CD.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-ai-sdlc-shared-standard-desktop.CplZgH9y.webp&quot; alt=&quot;Layer diagram. The top row shows three stages of the software development life cycle: requirements, code with tests and review, and maintenance. Spec Driven Development sits under requirements, agent tools together with Spec Driven Development sit under code, tests and review, and integration and unit tests that run locally and in CI/CD sit under maintenance. Under all stages lies a shared standard for working with the agent: skills, documentation in the repository, and harness and guardrails that give the agent a short feedback loop.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the standard for working with the agent covers the whole development cycle, not only writing code&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-rollout-starts-with-a-readiness-assessment-and-the-choice-of-a-pilot-team&quot;&gt;The rollout starts with a readiness assessment and the choice of a pilot team&lt;/h2&gt;
&lt;p&gt;The first step is an assessment of the current state and of readiness. It checks whether the process is ready for AI and assumes from the start that it will need adjusting. It shows what your SDLC standard looks like and whether it has bottlenecks already on paper: how CI/CD and the build servers work, what the release cadence is and whether the company has standardized it. It also covers testability hygiene and deployment frequency. If frequent deployments are not needed, it checks whether deployment is at least automated. The assessment draws on the DevOps approach and its feedback loop.&lt;/p&gt;
&lt;p&gt;The second step is choosing the pilot: one or two teams, systems, modules or applications. Teams that volunteer are the best choice, because the pilot needs the commitment of the whole team. The pilot team should be active in testing and later carry the solution further in the organization as its evangelist.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-ai-sdlc-rollout-steps-desktop.B4jTcBf8.webp&quot; alt=&quot;Flow diagram in five steps. The rollout starts with an assessment of process readiness: CI/CD, release cadence and testability. Next comes the pilot choice, ideally teams that volunteer, and measuring the state before the pilot as a diagnostic without KPIs. Then follows the pilot on a real application, and at the end the comparison with the state before. A dashed arrow labeled baseline links the measurement of the state before directly to the comparison.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the steps of rolling out AI in the SDLC, from the readiness assessment to the comparison with the baseline&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;measuring-the-state-before-the-pilot-is-a-diagnostic-with-no-set-targets&quot;&gt;Measuring the state before the pilot is a diagnostic with no set targets&lt;/h2&gt;
&lt;p&gt;The third step is measuring the state before the pilot, without KPIs. KPIs bring the risk of window dressing and gaming the metric. The measurement is a diagnostic: it shows the bottleneck and the place to start, so you speed up wisely.&lt;/p&gt;
&lt;p&gt;You do not measure vanity metrics: tokens used, lines of code generated, the number of pull requests or the share of code generated by AI.&lt;/p&gt;
&lt;p&gt;Łukasz Kałużny: “This is not productivity, and we will not reduce it all to one number; it will be a multidimensional assessment.” (translated)&lt;/p&gt;
&lt;p&gt;The measurement has 3 levels, or 2.5, depending on how you count. The goal of the whole measurement is faster or more stable delivery.&lt;/p&gt;

&lt;div class=&quot;table-scroll&quot; role=&quot;region&quot; tabindex=&quot;0&quot; aria-label=&quot;Three levels of measurement before and after the pilot&quot;&gt;&lt;table&gt;&lt;caption&gt;Three levels of measurement before and after the pilot&lt;/caption&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;What it shows&lt;/th&gt;
&lt;th&gt;How you measure it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Delivery&lt;/td&gt;
&lt;td&gt;Whether you deliver faster and with more stability&lt;/td&gt;
&lt;td&gt;4 DORA metrics, among them how often you deploy, the share of failed deployments and the time to restore after a failure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flow&lt;/td&gt;
&lt;td&gt;The time of the whole process&lt;/td&gt;
&lt;td&gt;From the arrival of a request for a feature or a change, such as a new screen, to its release in production&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quality&lt;/td&gt;
&lt;td&gt;Whether the agentic approach made quality worse, made it better or kept it stable (in a mature team)&lt;/td&gt;
&lt;td&gt;You set the method with Protopia for your case, for example the number of reported bugs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2 id=&quot;the-pilot-runs-on-a-real-application-and-ends-with-a-comparison-to-the-state-before&quot;&gt;The pilot runs on a real application and ends with a comparison to the state before&lt;/h2&gt;
&lt;p&gt;The fourth step is the pilot with the chosen team or teams on a real application under active development and maintenance, not on made-up cases tested on the side. At the start, Protopia supports it with mentoring and consultations for 1–2 production release cycles, and fits the length to the sprint, so the team goes through and tests the whole path end to end. Once everything is configured, the approach can start working after 1–2 sprints.&lt;/p&gt;
&lt;p&gt;After the pilot, you compare its data with the state you collected earlier (the baseline) and check whether anything moved forward. You have to be ready for nothing to move. Some research found, for example, that very mature teams saw no significant speedup.&lt;/p&gt;
&lt;p&gt;The pilot can also reveal areas to improve that do not depend on AI, such as acceptance processes. For example, the team builds and tests changes faster on its side, and the queue clogs at acceptance because someone cannot keep up with testing. At the start, this is a very welcome scenario: you then consider how to speed up or improve that stage. If the gain turns out small where AI was supposed to help, you may in the end decide against AI in that place. Success then means that the process got measured and improved in another part.&lt;/p&gt;
&lt;h2 id=&quot;the-pilot-result-decides-on-scaling-and-its-reach&quot;&gt;The pilot result decides on scaling and its reach&lt;/h2&gt;
&lt;p&gt;After the comparison, you have two paths. If you do not accept the improvement, you do not scale: you diagnose the constraints and perhaps go back to the foundations. You check which processes in the organization block the change, where the bottlenecks are, whether the problem lies in the pilot setup, and how review, testing and observability looked in this area. If you accept the improvement, you turn what worked into standards, templates and a definition of done.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-ai-sdlc-post-pilot-decision-desktop.C67hufpA.webp&quot; alt=&quot;Decision tree after the pilot. Two paths lead from the comparison with the state before. If the improvement is not accepted, you do not scale: you diagnose the constraints, such as bottlenecks or the pilot setup, and perhaps go back to the foundations. If the improvement is accepted, you scale: you turn what worked into standards, templates and a definition of done, prepare an adoption plan for part of the organization or all of it, and ambassadors and mentoring support the teams.&quot; width=&quot;736&quot; height=&quot;552&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the pilot result decides whether you scale or go back to the foundations&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;A rollout to the whole organization may not be needed. Depending on company size and criticality, part of the organization or a few development teams are sometimes enough. You update the standard and prepare an adoption plan for the projects that are ready and willing to adopt it.&lt;/p&gt;
&lt;p&gt;One way to scale is ambassadors: selected people from the teams attend dedicated ambassador workshops, then show their teams day to day how the process should work. Teams that adopt the standard get expert support: ad hoc consultations and mentoring.&lt;/p&gt;
&lt;p&gt;The standard sets the general direction and the main guidelines. Every project still needs final polishing, so the team always has some work to do before it gets value from the standard.&lt;/p&gt;</content:encoded><dc:creator>Łukasz Kałużny, Mikołaj Szczerbicki</dc:creator><category>AI in the SDLC</category></item><item><title>An AI assistant at a large company: system access, knowledge bases and tickets</title><link>https://protopia.tech/en/publications/ai-assistant-technical-side/</link><guid isPermaLink="true">https://protopia.tech/en/publications/ai-assistant-technical-side/</guid><description>You are building an AI assistant that should read company knowledge bases and open tickets. How do you give it access to systems so that it does not reveal other people&apos;s data after a prompt injection?</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;You are planning an AI assistant for a large company. It should answer questions from knowledge bases in SharePoint and ServiceNow, use data in SAP and open tickets in ServiceNow. Most of the work then goes into access to these systems: the assistant should read only selected sources and act with the permissions of the signed-in user, and even a successful prompt injection generally cannot take it beyond them. The assistant gets the hidden rules of ticket forms as plain text, so it handles many ticket types without a rewrite of the frontend.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;Before the team starts implementation, it must know which systems have an API, what that API gives and which licenses and limits restrict it.&lt;/li&gt;&lt;li&gt;The assistant should query systems with the signed-in user&apos;s token, because a service account or impersonation through a header lets a successful prompt injection pull out other people&apos;s data.&lt;/li&gt;&lt;li&gt;If a system requires impersonation, an API Gateway adds it based on the identity it reads from the token.&lt;/li&gt;&lt;li&gt;Knowledge sources need narrowing and cleanup: an integration with all of SharePoint shows too much, and raw HTML and icons from the knowledge base use up tokens.&lt;/li&gt;&lt;li&gt;The agent gets the rules hidden in ticket forms from a custom API as plain sentences, which the topic owner, for example the HR department, can edit.&lt;/li&gt;&lt;li&gt;The team needs people with long experience in integrations and security who can also carry an integration through talks with many stakeholders.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;most-of-the-work-is-reaching-the-systems-and-their-apis&quot;&gt;Most of the work is reaching the systems and their APIs&lt;/h2&gt;
&lt;p&gt;Marek Grabarz: “at the end of the day it turns out that 90% of the problems, or rather of the work, is about breaking through all these barriers and doing the integration, I would say, from start to finish” (translated)&lt;/p&gt;
&lt;p&gt;That is why the project starts with discovery: which systems the assistant integrates with, whether they have an API, what kind of API it is and what scope of knowledge it gives. The assistant connects to systems through APIs, without a user interface, so the team must first find the API and learn what it can do.&lt;/p&gt;
&lt;p&gt;Existing systems are often maintained with the frontend in mind. For example, the knowledge base team in ServiceNow keeps the content up to date and has a documented way to write it, but nobody has asked about access through the API before, so nobody knows the answer. Maintenance teams spread across Asia, the Americas and Europe do not know either.&lt;/p&gt;
&lt;p&gt;API availability is often selective: whether an API exists depends on the deployed modules and the license plans. The APIs of individual modules can be completely different from each other. It can also turn out that there is no API at all, or that it supports only a service account or impersonation.&lt;/p&gt;
&lt;p&gt;Discovery also covers licenses, access terms and rate limits, which the system architecture must then handle. One of the integrations in the project leads to SAP. The recording was made at the very start of May 2026. Shortly before, SAP announced that the use of its API by AI agents would be charged for or forbidden, and the analysis of the impact was still in progress. The speakers suppose that SAP wants to push everyone to use its own assistant, Joule, this way.&lt;/p&gt;
&lt;h2 id=&quot;a-service-account-lets-a-prompt-injection-reach-other-peoples-data&quot;&gt;A service account lets a prompt injection reach other people’s data&lt;/h2&gt;
&lt;p&gt;A service account has very limited use when the assistant’s access must be narrowed. In the project described here, access was one of the first problems. In most organizations the intuitive choice is a technical account with administrative access to the source system, for example SAP. The assistant uses it to fetch the data of one employee, and then of any other.&lt;/p&gt;
&lt;p&gt;Such an account is very easy to manipulate. A successful prompt injection on the chatbot side is enough to convince the assistant to hand over, for example, the data of the company’s CEO: their plans, and maybe also their salary and benefits.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Impersonation is&lt;/strong&gt; adding the user’s ID or full name to the technical account’s request, for example in a header or in the query string. It is still easy to manipulate, so it can leave a vulnerability in the integration and expose sensitive data.&lt;/p&gt;
&lt;h2 id=&quot;the-assistant-queries-systems-with-the-signed-in-users-token&quot;&gt;The assistant queries systems with the signed-in user’s token&lt;/h2&gt;
&lt;p&gt;The assistant should use the user’s real token and the user’s permissions in the target system. &lt;strong&gt;Delegated access is&lt;/strong&gt; a model based on delegated OAuth flows: the user signs in as themselves, gets an access token that represents them to the target system, and the assistant queries the system with that token. The assistant then generally cannot go beyond the user’s permissions. It does not re-create RBAC (Role-Based Access Control) and does not restrict access with system prompts or parameters: it uses full permissions, but only those of the signed-in person.&lt;/p&gt;
&lt;p&gt;This model requires identity federation, and its design needs knowledge from several fields. SAP, SAP SuccessFactors and every other large platform in the organization have their own user identity. Federation means that the agent’s user does not get a sign-in window for every system the assistant uses. SAP may also expect a completely different identity. Then a SAML federation is needed between Microsoft Entra ID and the source system, and in large organizations its configuration is not always simple.&lt;/p&gt;
&lt;p&gt;If access must be impersonated, you need an API Gateway that translates one model into the other. The assistant reaches the gateway with delegated access. The gateway reads the user’s identity from the token, and this path cannot be spoofed. Only then does the gateway add impersonation inside the system, based on its own deterministic logic.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-ai-assistant-system-access-desktop.Bxz2yjrN.webp&quot; alt=&quot;A comparison of three ways an AI assistant connects to a source system, for example SAP. With a service account, the assistant uses a technical account with administrative access and adds the user&apos;s ID in a header, so a successful prompt injection can pull out another person&apos;s data. With delegated access, the user signs in on their own, and the assistant queries the target system with the user&apos;s token and has only that user&apos;s permissions. If the system requires impersonation, the assistant passes the token to an API Gateway, which reads the identity from it and only then adds impersonation based on its own deterministic logic.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: service account, delegated access and impersonation through an API Gateway&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-assistant-reads-only-selected-sharepoint-sites&quot;&gt;The assistant reads only selected SharePoint sites&lt;/h2&gt;
&lt;p&gt;When you connect SharePoint to the assistant, you must extract the knowledge it should show, limit what it sees and defend it against injected bad content. SharePoint Online connects through Microsoft Graph API, and technically this works. The speakers see, however, that clients tend to turn on the integration for all of SharePoint. The assistant then gets access to many sites, including ones it should not read, even with delegated access.&lt;/p&gt;
&lt;p&gt;A user has access to the sites of their projects. These can also be sites where someone once turned on public access, uploaded very sensitive data and forgot about it. The access existed before, but a person does not go through sites one by one unless they do it on purpose. The assistant synthesizes everything it finds, so a question about budgets can surface projects that the user in theory had no access to. In the project described here, the assistant reads only selected, moderated sites with HR knowledge.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Content poisoning is&lt;/strong&gt; injecting bad content into the sources that the assistant reads. A person with write access to such a site can generate, for example with an LLM, a correct-looking HR procedure that says they are entitled to a Porsche as a company benefit, and the assistant starts to repeat it in its answers. They can also replace the form for reporting misconduct or workplace bullying with their own and collect the reports, even outside the organization.&lt;/p&gt;
&lt;p&gt;Graph API works only with SharePoint Online. SharePoint on-premises, which is also common, requires scraping and a vector database, and delegated access does not work there. The whole content of the selected site then goes into the database as vectors.&lt;/p&gt;
&lt;h2 id=&quot;servicenow-knowledge-base-content-needs-cleanup-before-the-model&quot;&gt;ServiceNow knowledge base content needs cleanup before the model&lt;/h2&gt;
&lt;p&gt;In the project described here, most of the knowledge sat in knowledge bases that the HR departments built in ServiceNow. Authors write the articles in a WYSIWYG editor, similar to the old Pajączek tool, in which you built a web page by clicking, so the API returns them as HTML with inline styles, divs and nested tables. Loading such content without extracting the text made token use 5 times higher and can end in a serious hallucination.&lt;/p&gt;
&lt;p&gt;Converting HTML to Markdown or plain text seems like a solved problem. The articles, however, often contain images, diagrams, attachments and links to articles in SharePoint. An agent that fetches only text does not know, for example, which procedure flow an attached image shows. So the model must be multimodal and handle images, text, PDFs and Word files, and the fetching of these elements must be orchestrated.&lt;/p&gt;
&lt;p&gt;Adding all images to the context causes other problems. Someone on the editorial team used icons with a green check mark or a red X instead of bullet points. One fetched article had 20 image attachments, and 19 of them, or maybe all of them, added nothing: they were icons or a banner with the company logo at the bottom of the page. You need exception lists that block the fetching of such attachments, for example at the API level, in the integration or in the agent’s instructions. Without them, token use grows.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-ai-assistant-knowledge-sources-desktop.D9z4UYDO.webp&quot; alt=&quot;Three knowledge sources of the assistant and what happens to them before the model. SharePoint Online connects through Graph API, but the assistant reads only selected, moderated sites with HR knowledge. SharePoint on-premises requires scraping, and the content of the selected site goes into a vector database. The ServiceNow knowledge base articles are stored as HTML, so the text must be extracted, because raw HTML made token use 5 times higher, and an exception list blocks the fetching of useless images such as icons or a logo strip. Everything goes to a multimodal model.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: how the assistant narrows and cleans its knowledge sources&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;knowledge-base-search-needs-rules-that-the-business-describes&quot;&gt;Knowledge base search needs rules that the business describes&lt;/h2&gt;
&lt;p&gt;Knowledge base search needs priorities, because local procedures exist next to global ones. An employee from Poland who asks about leave wants an answer under Polish law, not under the global policy. ServiceNow searches through the API by keyword: a query for “company car” does not find a Polish procedure that uses the Polish term “samochód służbowy”, so such differences need a description. Nested cases are harder, for example an HR employee with access to everything, or a manager who has people in five locations and asks how much leave their employee in France has. The permissions and the query then need proper orchestration.&lt;/p&gt;
&lt;p&gt;Search alone does not give 100% certainty that the answer is correct. The assistant also needs text understanding. The way to search and summarize the knowledge base must be described in user stories that the business and the analysts prepare.&lt;/p&gt;
&lt;h2 id=&quot;the-agent-gets-ticket-form-rules-from-a-custom-api&quot;&gt;The agent gets ticket form rules from a custom API&lt;/h2&gt;
&lt;p&gt;The goal of the project is fewer tickets, so the assistant first answers from the knowledge bases. When it cannot help, it falls back to a ticket. If the user says from the start that they are sick and want to file a ticket, the assistant starts the request at once.&lt;/p&gt;
&lt;p&gt;A ticket in ServiceNow is usually a form with drop-down lists, where the choice of one option shows more fields. There are many ticket types. Marek Grabarz: “We try, and I have said this before, to follow ‘orchestrate, do not replicate.’ So we try to orchestrate the handling of such a ticket, and not necessarily every ticket type.” (translated) The agent fetches the categories, matches a category to the user’s request, fetches the ticket types in that category and then the template: the fields, their types and whether they are required.&lt;/p&gt;
&lt;p&gt;This is not enough, because a lot of logic sits in the ticket form in ServiceNow, not in the backend: validators, lookups to tables (for example City ID) and dependencies between fields. ServiceNow refers to tables through SYS_ID identifiers, separate for the country table, the user table and every other table, and this is a big problem. Even if these dependencies were in the API, the agent would not understand them, because nobody has described them. There is no documentation: an external company built the ticket forms in the ServiceNow UI 10 years ago, and the HR staff know roughly how they work.&lt;/p&gt;
&lt;p&gt;An example: in a ticket that requests a training course, the training cannot take place earlier than 2 weeks from now. This rule is not in the API entity; it exists only in the form logic. The team solved this with a custom API in the middle that adds new properties to the ticket fields on the fly, especially to the dynamic ones: validation rules, descriptions and dependencies. The agent’s instructions tell it to fetch the template, check which fields are required and read the descriptive and validation rules. The rules are plain sentences, without JSON Logic, for example: if field A has a given value, field B is required.&lt;/p&gt;
&lt;p&gt;An LLM handles such a description very well, because it is not technical jargon or an IF this then THAT rule. The topic owner, for example the HR department, can also edit the description. The source of the rules for the custom API can be an attachment or another set of information, and the API applies the fields, validations and descriptions on the fly as text elements. The agent gets the information it needs dynamically, without a rewrite of the frontend or a step-by-step description in the prompt.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-ai-assistant-ticket-rules-desktop.C7O3UtgX.webp&quot; alt=&quot;The flow of a ServiceNow ticket handled by an AI agent. The agent fetches the categories, matches a category to the request, fetches the ticket types in that category and then the template with the fields and their types. A custom API in the middle adds validation rules, descriptions and dependencies to the fields on the fly, written as plain sentences, for example that a training date cannot be earlier than 2 weeks from now. The topic owner, for example the HR department, can edit the descriptions.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: ticket form rules through a custom API&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-team-needs-integration-experience-and-soft-skills&quot;&gt;The team needs integration experience and soft skills&lt;/h2&gt;
&lt;p&gt;The challenges of such a project lie mainly in integrations, security and authentication. The team needs people with long experience in IT who understand authentication and authorization, how APIs work and how to call them, networking, protocols and secure access to APIs over the network.&lt;/p&gt;
&lt;p&gt;Integration also means meetings with many stakeholders: the network people and the platform owners. Sometimes you have to push support to get information, and maybe also escalate the case to the system vendor. This work requires leadership, soft skills and the ability to translate needs into the language of different teams.&lt;/p&gt;
&lt;h2 id=&quot;maintaining-the-agent-requires-automated-answer-tests&quot;&gt;Maintaining the agent requires automated answer tests&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;An eval strategy is&lt;/strong&gt; an automated assessment of the agent’s answers against an expected answer, prepared from business cases and customer stories. The team is still working on it. LLM-based and non-LLM mechanisms automatically run a set of use cases and assess whether the agent answered sensibly and reliably, and whether it was proactive and friendly or curt and unhelpful.&lt;/p&gt;</content:encoded><dc:creator>Marek Grabarz, Szymon Warda</dc:creator><category>AI and agents</category></item><item><title>An AI assistant for more than 60,000 employees: HR, integrations and a staged rollout</title><link>https://protopia.tech/en/publications/ai-assistant-for-60000-employees/</link><guid isPermaLink="true">https://protopia.tech/en/publications/ai-assistant-for-60000-employees/</guid><description>Employees at a global company did not know which HR system to use for what. The AI assistant built for them knows each user&apos;s context and sees only the data that user may access. Its scope grows in stages.</description><pubDate>Wed, 20 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;When employees at a large company do not know which HR system holds the answer or which ticket to file, an AI assistant can be a single entry point to those systems. This rollout at a global pharmaceutical company shows how to narrow the first scope, when the assistant should send the user to a form, why identity and integrations take the most work, and how to measure the effect as the assistant reaches more groups of employees.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;The project started because HR came to IT with a problem: employees got lost in scattered HR systems.&lt;/li&gt;&lt;li&gt;The assistant knows the user&apos;s country, role and contract, while the process logic stays in the source systems.&lt;/li&gt;&lt;li&gt;The scope grows in stages: first answers from knowledge bases, then help with tickets and annual plans.&lt;/li&gt;&lt;li&gt;For the most complex ticket types, the assistant deliberately sends the user to a form and says from the first contact what it does not handle.&lt;/li&gt;&lt;li&gt;Identity, integrations and identifying the APIs take the most work. The first obstacle: the assistant may see only the data the user has access to in the source system.&lt;/li&gt;&lt;li&gt;The rollout covers ever larger groups of employees, and the effect is measured by ticket count, returning users and feedback.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;the-project-started-with-hrs-problem-of-scattered-systems&quot;&gt;The project started with HR’s problem of scattered systems&lt;/h2&gt;
&lt;p&gt;The assistant came from an initiative of the HR department at a large global pharmaceutical company that operates in several dozen countries. HR matters at this company are spread across many systems, for example knowledge bases, payroll systems and benefits systems. Employees did not know which system to turn to, what they would find there and what they could do there.&lt;/p&gt;
&lt;p&gt;HR wanted to reduce the complexity of this setup and put it plainly: if IT did not do it, HR would find a way and solve the problem on its own.&lt;/p&gt;
&lt;p&gt;The speakers fairly often see the opposite situation at clients. The board or a director says “we need AI”, and IT wonders what it can do. IT usually starts with its own knowledge bases instead of discovery with the business. This is the source of an opinion heard in the market: 90% of AI rollouts fail because they bring no value.&lt;/p&gt;
&lt;p&gt;When the Protopia team joined the rollout, the project was already fairly well described, had a sponsor and business owners, and the company ran an AI platform. The client decided it needed someone from outside who had seen agentic rollouts. With these systems, experience means a year or two, and context problems or hallucinations appear sooner or later.&lt;/p&gt;
&lt;h2 id=&quot;the-assistant-knows-who-the-user-is&quot;&gt;The assistant knows who the user is&lt;/h2&gt;
&lt;p&gt;In this project, &lt;strong&gt;the assistant is a chatbot frontend that knows the user’s context and makes HR processes easier for them&lt;/strong&gt;. The name “assistant” sets it apart from an agent, understood as an autonomous component that connects and automates processes in the background. Unlike a classic chatbot without context, it knows the user’s country, role and contract type: an employment contract or a B2B contractor agreement. This decides what the person can do in HR matters.&lt;/p&gt;
&lt;p&gt;The assistant integrates with systems that have run at the company for years, for example the HR ticketing system. Tickets cover, among other things, maternity or sick leave, a training request, a role change and an employee transfer. Around them sits a huge, scattered knowledge base (KB): some procedures are local, some are global. This raises questions such as whether a manager with employees in three different locations has access to all the KBs or only to some of them.&lt;/p&gt;
&lt;h2 id=&quot;the-scope-grows-in-stages-along-the-roadmap&quot;&gt;The scope grows in stages along the roadmap&lt;/h2&gt;
&lt;p&gt;The assistant takes on one area after another, because the product roadmap does not try to put everything in at once. The client has many systems, so the starting scope had to be narrowed.&lt;/p&gt;
&lt;p&gt;Knowledge base integration came first, probably the most urgent need: the assistant was to answer in the user’s context. Example questions: how many days of leave do I have, am I eligible for a given benefit, can I get a company car. The assistant searches these knowledge bases for the user and builds the answer in the language of the question.&lt;/p&gt;
&lt;p&gt;The second area is tickets. The assistant first tries to resolve the matter from the KB. When the user still does not know what to do, or already knows that a ticket is needed, the assistant offers to fill in the right ticket with them and shows where it is. There are many ticket types.&lt;/p&gt;
&lt;p&gt;The third area is employees’ annual plans. Together with the manager and the assistant, the employee drafts a plan with SMART goals and reviews it. The roadmap beyond that is broad, and for now it is a set of intentions: learning, recruitment, employee development and internal promotions. The target is for the assistant to be a central entry point that ties the HR systems together.&lt;/p&gt;
&lt;p&gt;The speakers see that clients expect a superagent that solves every problem from its first day in production. From experience they know that this does not work. In this project the work is iterative, split into tasks and separate agents.&lt;/p&gt;
&lt;h2 id=&quot;the-assistant-says-plainly-what-it-does-not-handle&quot;&gt;The assistant says plainly what it does not handle&lt;/h2&gt;
&lt;p&gt;From the first contact, the assistant states what is possible and what is not. The speakers see that clients are often afraid to say what an agent can and cannot do. Here, with each new MVP release, the organization holds a large town hall meeting and tells the first users what the assistant can do.&lt;/p&gt;
&lt;p&gt;Some ticket types have complex forms with dynamic drop-down menus and too many dependencies. The team decided that the assistant does not handle them. It is more efficient for the user when the assistant says it cannot help with this matter and gives a link to the form, which the user then fills in on their own. This way the team avoids hallucinations. The Pareto principle held for tickets:&lt;/p&gt;
&lt;p&gt;Marek Grabarz: “We can handle 90% of the functionality with 10% of the effort, and the other way around.” (translated)&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-hr-assistant-ticket-path-desktop.h554t4lR.webp&quot; alt=&quot;An employee asks a question, and the assistant first answers from the knowledge bases. If that is enough, the matter is resolved. If the matter is still unresolved, a ticket is needed. The employee fills in a supported ticket type together with the assistant. For a complex form with dynamic menus, the assistant says it cannot help and gives a link to the form, which the employee fills in alone. The team applied the Pareto principle in a 90/10 ratio.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: when the assistant sends the user to a form&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-assistant-connects-systems-and-processes-stay-at-the-source&quot;&gt;The assistant connects systems, and processes stay at the source&lt;/h2&gt;
&lt;p&gt;The assistant orchestrates processes and does not have to duplicate them. When a process is very complex and locked inside the source system, the team describes it to the assistant in general terms, without step-by-step instructions. The substance of the process stays on the system side.&lt;/p&gt;
&lt;p&gt;Existing systems such as SAP or SharePoint do not integrate in an obvious way. You have to work out how to connect to them, where their maintenance teams are and where in the organization the knowledge of the target systems sits.&lt;/p&gt;
&lt;p&gt;Delegated access is an example of a low-level problem. If the system is SAP, for example, a user sees their own annual goals and e-learning in it, and a manager also sees their employees’ data and what the employees have declared. The assistant may see only the data that the user has access to, so it must replicate the roles and permissions from the source system (RBAC). This was the first obstacle the team had to solve with the client.&lt;/p&gt;
&lt;p&gt;Marek Grabarz: “Most of the work is around identity, around integration, around identifying those APIs.” (translated)&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-hr-assistant-architecture-desktop.Cry5h3e0.webp&quot; alt=&quot;An employee, whose context is country, role and contract, asks the assistant&apos;s chatbot a question. The assistant orchestrates processes through integrations and APIs in the source systems: knowledge bases, the HR ticketing system, SAP and SharePoint. The process logic and permissions stay in those systems. The assistant replicates the roles and permissions (RBAC) from the source system, so it sees only the data the user has rights to.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the assistant as an entry point to HR systems&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-rollout-reaches-ever-larger-groups-of-users&quot;&gt;The rollout reaches ever larger groups of users&lt;/h2&gt;
&lt;p&gt;The stage called the first pilot run is production with limited reach. At the time of recording, it had more than 1,000 active unique users.&lt;/p&gt;
&lt;p&gt;Probably the day before the recording, the team got approval to go to production with the full scope described above. This is still not a full rollout: the next stage covers 10% of employees, which means thousands of people. The target, probably at the start of July, is for the assistant to reach all countries and all employees. At the time of recording, no major obstacles to this plan were visible.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-hr-assistant-rollout-desktop.BhVukS_a.webp&quot; alt=&quot;The assistant rollout in three stages. The first pilot voyage is production with limited reach and more than 1,000 active unique users. After production approval, the next stage is the full feature scope for 10% of employees, which means thousands of people. The target, probably at the start of July, is for the assistant to reach all countries and all employees.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: rollout stages of the assistant&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;tickets-returning-users-and-feedback-measure-success&quot;&gt;Tickets, returning users and feedback measure success&lt;/h2&gt;
&lt;p&gt;The first measure is the number of tickets, especially user requests, meaning questions to HR like “how do I do this or that”. The assistant should resolve the matter at the knowledge base stage, so that it does not end as a ticket and HR staff do not answer such questions in chats. The team tracks changes in the number of tickets and in their types. The second measure is return visits: whether users are willing to come back to the assistant. The third is whether the assistant really solves problems. The feedback built into the platform and the assistant shows this.&lt;/p&gt;
&lt;p&gt;The feedback has two dimensions: business and technical. The user rates whether the session solved their problem, and whether it did so quickly or only after many questions and steps. From the feedback, the team finds the cause: a problem on the assistant side, a misunderstanding of how the assistant works, or gaps in the knowledge bases. A team fixes the knowledge bases on an ongoing basis, because they sometimes hold old or incorrect versions of content. Errors happen, but a lot of the feedback says: great, just add more features.&lt;/p&gt;
&lt;p&gt;When errors come up, the team uses a monitoring platform with tracing. At the pilot stage, the team can pull the full chat content, the call context and the called functions, and the session ID is kept. After the pilot, tracing in production will be narrower, because the assistant handles sensitive HR data, sometimes about pay, sometimes about health.&lt;/p&gt;
&lt;h2 id=&quot;cost-covers-the-project-and-tokens-and-some-benefits-are-hard-to-measure&quot;&gt;Cost covers the project and tokens, and some benefits are hard to measure&lt;/h2&gt;
&lt;p&gt;The cost of an agentic solution is the sum of many items, including the project cost and the tokens that the whole system uses. You can try to measure the benefits through organizational efficiency and user satisfaction, but the value of many such processes cannot be measured in money.&lt;/p&gt;
&lt;p&gt;In a large organization, it is not always clear where the right ticket is and which one to file. This is what people call tribal knowledge. Take a financial analyst with an HR matter. Instead of waiting for days, writing emails and pinging someone on Microsoft Teams, they can resolve it quickly with the assistant and return to their own work. Such uninterrupted work time is a benefit that is hard to measure.&lt;/p&gt;</content:encoded><dc:creator>Marek Grabarz, Mikołaj Szczerbicki</dc:creator><category>AI and agents</category></item><item><title>FinOps is a process, not a tool</title><link>https://protopia.tech/en/publications/finops-is-a-process-not-a-tool/</link><guid isPermaLink="true">https://protopia.tech/en/publications/finops-is-a-process-not-a-tool/</guid><description>A FinOps tool will not fix your cloud cost problems, and cost cuts handed to teams can come back bigger. You will learn how to build a FinOps process and roll it out in your organization in phases.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;Your cloud costs are spread across vendor subscriptions, licenses and several versions of the same services, and you want to get them under control without adding work for your teams. In this approach, FinOps is an ongoing process in the organization: it first gives cost visibility, then adds shared rules, automation and monitoring, and savings come as a side effect. The rollout runs in phases: visibility, optimization, and then maintaining the process, which the organization handles on its own.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;Fragmentation hides costs: vendor resources without budgets or alerts, licenses on a separate invoice, and many versions of the same service, each of which needs its own policies.&lt;/li&gt;&lt;li&gt;Neither another tool nor asking teams to cut costs delivers lasting results: a new tool can end up ignored, and the cut costs can come back.&lt;/li&gt;&lt;li&gt;The process starts with a review of vendors and of your own resources, then adds consistent budgets, automated policies, incentives for teams, and monitoring whose conclusions feed back into the documents from the first step.&lt;/li&gt;&lt;li&gt;The first rollout phase asks little of the client: mainly access and named contact people.&lt;/li&gt;&lt;li&gt;In the second phase, the savings from an optimization are weighed against the vendor&apos;s cost of the change, because that cost can exceed the gain.&lt;/li&gt;&lt;li&gt;The maintenance phase never ends: the organization keeps the FinOps culture going on its own.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;finops-is-a-process-that-gives-cost-visibility&quot;&gt;FinOps is a process that gives cost visibility&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;FinOps is a process in the organization that gives predictability and insight into cloud costs.&lt;/strong&gt; Along the way it improves governance and compliance, and savings are its side effect. FinOps is not a one-time check, a tool installed in CI/CD or a cost review once every three months. The speakers often see the reverse order: a company introduces checks once a year or once a quarter and wants to cut costs right away. That will not work.&lt;/p&gt;
&lt;p&gt;With visibility, the organization knows what it spends money on, controls costs more easily, has fewer resource types and reviews them more easily. A lower cloud bill must not come from extra work for people.&lt;/p&gt;
&lt;p&gt;Szymon Warda: “It’s not about spending less. It’s about spending more wisely.” (translated) As a rule, spending more wisely also means spending less.&lt;/p&gt;
&lt;h2 id=&quot;costs-leak-through-scattered-resources-and-the-long-tail-of-systems&quot;&gt;Costs leak through scattered resources and the long tail of systems&lt;/h2&gt;
&lt;p&gt;Costs spread wherever each vendor, region or department has its own underused resources. The description comes from work with one client, but the speakers see this situation in very many organizations, to different degrees and in different areas.&lt;/p&gt;
&lt;p&gt;A typical setup is many subscriptions and accounts where vendors deploy software and, in a way, spend money that is not theirs. Looser control is not bad in itself: it gives agility and shortens delivery time, and organizations often do not know how to do it differently. But this setup has no reservations, no budget control, no alerts, no forecasts and no rules on which resources to use. Some resources can be centralized: more reuse makes control easier, cuts staff overhead and improves compliance.&lt;/p&gt;
&lt;p&gt;Software licenses are a separate problem. At this client nobody controlled them, and they can run into serious dollar amounts. Licenses are often easy to overlook because they land on a different invoice, and in some organizations they are a very significant cost.&lt;/p&gt;
&lt;p&gt;Repeated solutions that differ slightly from each other also create a hidden cost. Say one service exists in six versions: each one needs backup, security and network policies and has to be maintained. Every system a vendor deploys extends the long tail of costs, and the IT department’s headcount has to grow with the number of systems it maintains. This model does not scale well.&lt;/p&gt;
&lt;p&gt;Hours add up too. Over the years, three hours here, two there and four somewhere else can turn into, for example, two or three full-time positions. Getting that time back later is expensive, because saving four hours a week pays off only moderately. It may be better not to let those four hours appear at all and to stay at, say, 20 minutes. That is why FinOps also watches future costs: the organization sees what they will look like and how they will be distributed.&lt;/p&gt;
&lt;h2 id=&quot;neither-a-new-tool-nor-cost-cuts-by-teams-give-lasting-savings&quot;&gt;Neither a new tool nor cost cuts by teams give lasting savings&lt;/h2&gt;
&lt;p&gt;Buying a FinOps tool or license will not fix the problems, yet the speakers often see companies take this path. It is easier to give in to marketing that promises the product will remove them. Szymon Warda: “no magic powder has ever cured anyone” (translated)&lt;/p&gt;
&lt;p&gt;A new tool lands on a team that already maintains many products. Nobody knows whether it will be configured well, whether it will work and whether people will want to use it. At worst it can become another obstacle, at best another ignored tool that raises the cost.&lt;/p&gt;
&lt;p&gt;A set of best practices is a better direction because it addresses behavior. But the organization has to create conditions in which people want to adopt them, have time for it and know how to apply them.&lt;/p&gt;
&lt;p&gt;Asking every team to find unneeded resources and cut costs on its own is a trap the speakers see very often. Six months later the costs can come back even higher, a yo-yo effect: they were shifted onto people, the company paid incorrectly, or the savings only had to look good in Excel for the right quarter. Even eight hours a month per team adds up to a fairly large budget across the organization.&lt;/p&gt;
&lt;p&gt;A good share of the savings and visibility has to happen at the organization level. Protopia proposes to do this work once, centrally, and to give teams automation and ready-made processes instead of a daily morning review of a spreadsheet on their to-do list. &lt;strong&gt;A pit of success is a setup in which the right things are easy, cheap and simple to do.&lt;/strong&gt; Teams hand over part of the work, and if they act as agreed across the organization, their work gets easier and cheaper. Then nobody has to fight the organization, because optimal behavior pays off.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-finops-three-paths-desktop.BN3DEYiV.webp&quot; alt=&quot;Three paths to lower cloud costs. Buying a tool adds a new product for the team, which at best becomes another ignored tool. Team cuts work only for a short time: after half a year the costs come back, and bigger, like a yo-yo effect. Work done once, centrally, gives teams automation and ready-made processes, so the right things are easy, cheap and simple to do.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: a tool, team cuts and a central process&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-first-step-puts-vendors-and-the-organization-in-order&quot;&gt;The first step puts vendors and the organization in order&lt;/h2&gt;
&lt;p&gt;Building the process starts with vendors, in two ways. For the future, the organization sets architecture guidelines: which resources to use and how vendors should report budgets and forecast costs. This way the long tail of costs stops growing at its previous pace. For systems already running in production, the organization takes an inventory: it checks which resources are in use and what can be optimized. The inventory produces a list of best practices and changes that make the system meet regulatory, security, governance and compliance requirements.&lt;/p&gt;
&lt;p&gt;The organization itself also needs order. As a rule, a cleanup stage comes first: licenses, unused resources, bad tags, budgets and different approaches to the same FinOps work. The processes have to be cleaned up and unified. Then the organization consolidates the resources it uses in many places. The vendor review also shows which resources vendors use and what is worth taking under the organization’s own control, for example DevOps agents, a Kubernetes platform or an observability platform.&lt;/p&gt;
&lt;h2 id=&quot;governance-combines-consistent-budgets-with-automated-policies&quot;&gt;Governance combines consistent budgets with automated policies&lt;/h2&gt;
&lt;p&gt;The second step is governance built on visibility and automation. Visibility starts with budgets at the organization level: all systems report costs in a consistent way and by the same forecasting rules, so they can be compared. The organization also sees which resources it uses, in which versions and where, and breaks the invoice down into smaller items. It reviews data retention and reservations, then decides where to invest time, what to buy and what to drop.&lt;/p&gt;
&lt;p&gt;The first part of automation is policies that enforce good behavior. The rule applies, period, with no reliance on scout’s honor. Where something cannot be checked automatically, dedicated services quickly report the noncompliance. Reporting rules shorten the time between an event and the moment the organization learns about it. If the organization learns about a resource, for example, two months after it started using it, a request to change it will not work. If a day or two passes between use and information, a change becomes more realistic.&lt;/p&gt;
&lt;h2 id=&quot;nudging-gives-teams-ready-made-automation-and-modules&quot;&gt;Nudging gives teams ready-made automation and modules&lt;/h2&gt;
&lt;p&gt;The third step encourages people to follow good practices of their own accord, because a stick without a carrot is not enough. &lt;strong&gt;Nudging is encouraging good behavior and making it easier.&lt;/strong&gt; This stage has two parts: the goals and the way to deliver them. An example goal is good practices and a lower cost of Azure or any other cloud.&lt;/p&gt;
&lt;p&gt;The goals are delivered through, for example, automatic machine start and shutdown, backup automation, and autoscaling of databases, Kubernetes clusters, virtual machines, and PaaS and SaaS services. On top of that come license count monitoring and processes that make it easy to request access. A team tags a resource, and from that moment all the automation covers it. The team can also use a ready-made CI/CD or infrastructure as code module. Governance can add DevOps and security elements to the modules.&lt;/p&gt;
&lt;p&gt;Some organizations, once they have visibility, change how they work internally and with their vendors.&lt;/p&gt;
&lt;h2 id=&quot;monitoring-leaves-room-for-exceptions-and-closes-the-loop&quot;&gt;Monitoring leaves room for exceptions and closes the loop&lt;/h2&gt;
&lt;p&gt;The fourth step is monitoring, because not everything can be automated and not everything is worth automating. It would be easier to impose FinOps on day one with no exceptions, but the business will demand them. Monitoring shows where the exceptions are, protects against regression and lets the process grow. This is the less pleasant part, and it moves slowly: architecture reviews, rules for developing vendor documents, and reporting.&lt;/p&gt;
&lt;p&gt;The conclusions from nudging and monitoring go back to the documents from the first step and show how the organization, its vendors and the existing systems should change. This creates a feedback loop in which the process changes along with the organization.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-finops-process-loop-desktop.DrxVYf-t.webp&quot; alt=&quot;The FinOps process has four steps: vendors and cleanup in the organization itself, governance with budgets and policies, nudging with ready-made automations and modules, and monitoring of exceptions with reviews. The conclusions from nudging and monitoring go back to the documents from the first step, which creates a feedback loop in which the process changes along with the organization.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the FinOps process steps form a feedback loop&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-first-rollout-phase-shows-what-happens-in-the-organization&quot;&gt;The first rollout phase shows what happens in the organization&lt;/h2&gt;
&lt;p&gt;Protopia splits the rollout of the whole process into three main phases, and the first one should show what is actually happening in the organization. It covers the inventory, setting up alerts, a license review and identifying resources through tags. As a rule, Protopia also applies simple standards found in every organization, such as the automations from the third step and backup retention automation. These practices reach beyond FinOps. Along the way, things that do not comply with security policies and compliance requirements often come to light.&lt;/p&gt;
&lt;p&gt;In Protopia’s experience this phase takes roughly 2–3 months, no longer, because Protopia does not want to drag it out. The time depends on the organization.&lt;/p&gt;
&lt;p&gt;In this phase the client mainly provides access and names contact people. They answer questions about naming standards, how the systems are built and which vendor delivers what, in other words tribal knowledge that is often missing from the documentation. As a rule, this adds up to a few person-days over the whole phase.&lt;/p&gt;
&lt;h2 id=&quot;the-second-phase-turns-visibility-into-concrete-changes&quot;&gt;The second phase turns visibility into concrete changes&lt;/h2&gt;
&lt;p&gt;The second phase draws conclusions from visibility and starts conversations about optimizing existing resources. Protopia shows how much the client can save, then asks the vendor how much the change costs to implement. Sometimes the cost of the change may exceed the savings. Sometimes the vendor prices the change high, and Protopia presents counterarguments. This can become a negotiation over where real effort is worth putting in.&lt;/p&gt;
&lt;p&gt;The phase also includes a detailed review of new architectures. Earlier the work was about general things; now it is about new systems and the delivery process in the organization. This review produces a set of good practices that feeds the first step of the process.&lt;/p&gt;
&lt;p&gt;The FinOps process is now tailored to the specific organization, including who the points of contact are. Decisions are made about what becomes a shared resource, what is kept separate and how to manage shared resources. Changes to cloud resources will most often be required, but Protopia tries to keep them as small as possible. Protopia can deliver the shared resources itself, for example as infrastructure as code with a CI/CD pipeline for automatic deployment. The client gets a resource that is manageable and has minimal maintenance cost.&lt;/p&gt;
&lt;p&gt;The client takes part in this work, and the knowledge transfer is meant to let the organization use and manage these tools on its own. The phase usually lasts about six months: it starts with a larger scope and slowly winds down.&lt;/p&gt;
&lt;h2 id=&quot;the-third-phase-maintains-the-finops-culture-indefinitely&quot;&gt;The third phase maintains the FinOps culture indefinitely&lt;/h2&gt;
&lt;p&gt;The third phase never ends, because what is being built is a process. At this stage the organization already has good practices, policies, automations and monitoring. It maintains the FinOps culture on its own, and Protopia can support it. The organization expands this culture, runs the feedback loop and adds small elements that keep the process paying off for the organization. Oversight of old system migrations and reviews of systems that are still being rolled out can continue. The process also expands to the remaining parts of the organization.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-finops-rollout-phases-desktop.9Y6XpF26.webp&quot; alt=&quot;The FinOps rollout has three phases. Phase one, visibility, takes roughly 2–3 months: inventory, alerts, a license review and resource tags, while the client mainly provides access and contact people. Phase two, optimization, usually lasts about half a year: savings are weighed against the cost of the change, new architectures are reviewed and decisions are made about shared resources, while the client takes part and takes over the knowledge. Phase three, maintenance, never ends: the organization maintains the FinOps culture on its own, runs the feedback loop and expands the process to the remaining parts of the organization.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: FinOps rollout phases and the organization&amp;#39;s role&lt;/figcaption&gt;&lt;/figure&gt;</content:encoded><dc:creator>Szymon Warda, Mikołaj Szczerbicki</dc:creator><category>FinOps</category></item><item><title>Taking an AI agent from PoC to production: a documented process, system access and a decision after testing</title><link>https://protopia.tech/en/publications/ai-agent-from-poc-to-production/</link><guid isPermaLink="true">https://protopia.tech/en/publications/ai-agent-from-poc-to-production/</guid><description>You want to test an AI agent before you invest in a full solution. Prepare a PoC: write down the process, check access to systems, set success criteria and decide after testing.</description><pubDate>Wed, 01 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;You have an idea for an AI agent and need to decide whether to build the full solution. A PoC prepares that decision. Before it, the team writes the process down from what employees know and checks access to systems; after testing, it rates the result as scale, iterate or kill against criteria set in advance.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;Before the PoC starts, the person who runs the process writes down in Word what they do and why, or records their screen. This preparation saves a lot of work that should never be left to debugging.&lt;/li&gt;&lt;li&gt;Business people who want to test something should get a playground with the same bare model the agent will be built on, because that model behaves differently from ChatGPT.&lt;/li&gt;&lt;li&gt;In most of Protopia&apos;s PoCs, access to data and systems causes the most problems, and building the agent itself is mostly tedious integration work.&lt;/li&gt;&lt;li&gt;Decide at the start what triggers the agent and at which point a human enters the process. An approval matrix sets what can run automatically. Processes can run without a human when any problem would be very small and would harm neither reputation nor finances.&lt;/li&gt;&lt;li&gt;Success criteria come before testing. Tests stay focused on the goal: the UI is ugly on purpose, and a log records the agent&apos;s run, its duration and its decisions.&lt;/li&gt;&lt;li&gt;After testing, the decision is scale, iterate or kill, and a kill with well-documented findings is a success too, because it keeps the team out of a project that would not have worked anyway.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;a-poc-tests-an-ai-agent-before-you-invest-in-a-full-solution&quot;&gt;A PoC tests an AI agent before you invest in a full solution&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;A proof of concept (PoC) is a feasibility study that checks whether a concept will work and whether it will work consistently.&lt;/strong&gt; Protopia often calls it a proof of value (PoV) as well, because it has to check whether the solution will deliver value and whether it is feasible.&lt;/p&gt;
&lt;p&gt;One client wrote that the prompt is very precise, yet the answers differ in substance. For a PoC to succeed, you sometimes need to approach the process a little differently and use business knowledge differently than when you simply drop a task into GPT, Copilot, Gemini or Claude.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-ai-agent-poc-stages-desktop.Da0ClzNF.webp&quot; alt=&quot;Diagram of an AI agent&apos;s path from idea to decision. Preparation covers writing the process down from what employees know, a playground for the business and a check of access to systems. The actual PoV starts with success criteria, followed by tests with logging. After testing, a decision follows: scale leads to a pilot and production, iterate goes back to new tests, and kill closes the initiative with documented findings.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: from writing the process down to the decision after PoC testing&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-first-step-is-writing-the-process-down-from-employee-knowledge&quot;&gt;The first step is writing the process down from employee knowledge&lt;/h2&gt;
&lt;p&gt;Before the PoC, you need to know exactly what the process to automate looks like: what people do by hand today and how it runs in the systems. Protopia has the person who holds this knowledge write down in Word, step by step, what they do and why, or record their screen. Sometimes two employees or experts can do it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tribal knowledge is the knowledge of the people who run the process.&lt;/strong&gt; Writing it down is the first step, and in Protopia’s projects it is the step that contributes most to success. Often nobody updates the instructions when business conditions, the system configuration or the approach change.&lt;/p&gt;
&lt;p&gt;At one insurance client, an employee wrote down that they first take data from the system, then search the internet to find out who the customer really is and what they could be offered, and finally try to match it all and do an analysis. A Word document like this is a great starting point for planning, because it shows what actually happens and what the person does. A week of preparation and data gathering saves a great deal of work that should never be left to debugging.&lt;/p&gt;
&lt;p&gt;Conway’s law shows that a process mirrors the communication structure of the company, not the logic it should have. Writing the process down can therefore be a chance to check whether it is worth simplifying and cleaning up.&lt;/p&gt;
&lt;h2 id=&quot;a-playground-shows-the-business-how-the-model-really-behaves&quot;&gt;A playground shows the business how the model really behaves&lt;/h2&gt;
&lt;p&gt;After the process is written down, you need to set expectations. Non-technical people should be able to check for themselves how agents behave: on the playgrounds the team later builds the solution from, not in Copilot or ChatGPT. Protopia works with Azure most of the time. &lt;strong&gt;A playground is a chat-like place that gives access to the bare language model, and this model later goes into agents and applications.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A bare model behaves quite differently from ChatGPT in agent mode. For example, in your own OpenAI model hosted on Azure, web search works somewhat differently: the Bing Search underneath does not behave the way it does in the public consumer ChatGPT. On the playground, the business sees how the model actually works, including with a prompt like the one in the client’s message.&lt;/p&gt;
&lt;p&gt;If business people want to test something, they should get access to a playground instead of having the tools blocked. They then form realistic expectations and build trust. These tests happen before the actual PoC, at the stage of play and discovery.&lt;/p&gt;
&lt;h2 id=&quot;access-to-systems-makes-the-analysis-hardest&quot;&gt;Access to systems makes the analysis hardest&lt;/h2&gt;
&lt;p&gt;After the playground tests comes the analysis, the worst phase. In most of Protopia’s PoCs, automation and agent behavior cause no problems. Łukasz Kałużny: “Most of the problems come down to access to systems.” (translated)&lt;/p&gt;
&lt;p&gt;Someone has to do the painstaking work: find out where the needed data comes from, from which system, how to get access to it and whether the system has an API. If there is no API, check whether access can be automated through the browser or whether access to the database is enough. This is an inventory of the systems the automated process uses. Today it is most often about pulling data and acting on it or, much as with RPA (robotic process automation) in the past, writing back to another system or calling an API.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-ai-agent-poc-data-access-desktop.DitMyv3S.webp&quot; alt=&quot;Decision tree for an agent&apos;s access to data. The team finds out which data is needed and which system it comes from, then checks whether that system has an API. If it does, the agent calls the API. If it does not, the team checks whether access can be automated through the browser or whether access to the database is enough.&quot; width=&quot;736&quot; height=&quot;552&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: how to set up an agent&amp;#39;s access to data&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Once the process is written down, most of the remaining technical work is tedious integration. Building agents is very much a programming job, not data science.&lt;/p&gt;
&lt;h2 id=&quot;the-trigger-and-the-humans-role-are-set-at-the-start&quot;&gt;The trigger and the human’s role are set at the start&lt;/h2&gt;
&lt;p&gt;At the start, you need to answer two questions. The first is what triggers the agent. The agent will not figure it out by itself, so something has to trigger its work, for example an email in a mailbox, an event or a human click. The second is human in the loop: at which point a human enters the process, approves the agent’s work or uses it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An approval matrix is a set of clear rules for when something can run automatically and when it cannot.&lt;/strong&gt; It sets the human’s role in the process. Protopia often advises having someone approve at the end, and in many cases that is enough.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-ai-agent-poc-trigger-approval-desktop.CrQ9JQLx.webp&quot; alt=&quot;Diagram of how an AI agent starts and where a human fits into the process. An email in a mailbox, an event or a human click starts the agent. The result of the agent&apos;s work goes to the approval matrix. The matrix sends it to a human who approves the result at the end, or to automatic execution when a possible problem is very small and harms neither reputation nor finances.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the agent trigger and the approval matrix&lt;/figcaption&gt;&lt;/figure&gt;

&lt;div class=&quot;table-scroll&quot; role=&quot;region&quot; tabindex=&quot;0&quot; aria-label=&quot;The human’s role in processes with an agent&quot;&gt;&lt;table&gt;&lt;caption&gt;The human’s role in processes with an agent&lt;/caption&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Process&lt;/th&gt;
&lt;th&gt;What the agent does&lt;/th&gt;
&lt;th&gt;Human role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Customer complaint or inquiry&lt;/td&gt;
&lt;td&gt;Works through the case, prepares the justification and the reply, shortens delivery time&lt;/td&gt;
&lt;td&gt;Checks what goes out to the customer before it is sent; may change nothing and click “send”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PoC that gathers data and prepares an analysis&lt;/td&gt;
&lt;td&gt;Prepares a finished report with all sources&lt;/td&gt;
&lt;td&gt;Checks on their own whether the report is correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Price quote based on web search&lt;/td&gt;
&lt;td&gt;The agent searches the web and proposes a price quote; an extra judge agent named Judge screenshots the source pages and passes the price on when the basic information on those pages checks out&lt;/td&gt;
&lt;td&gt;Approves the result at the end&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;After-sales handling of open training courses under Protopia’s Patoarchitekci brand (sending invitations, adding people to the course)&lt;/td&gt;
&lt;td&gt;Handles the course purchase&lt;/td&gt;
&lt;td&gt;No human: any problem would be very small and would harm neither reputation nor finances&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Protopia always leaves the interpretation of the AI Act to the client, because it is not Protopia’s specialty. When you set the human’s role, what counts is the notion of high-risk systems and how the company’s lawyers and compliance interpret it. The first idea of where the human belongs may turn out wrong, and that comes out during testing. That is one of the things a PoV is for.&lt;/p&gt;
&lt;h2 id=&quot;success-criteria-are-written-down-before-the-first-test&quot;&gt;Success criteria are written down before the first test&lt;/h2&gt;
&lt;p&gt;The actual PoV starts with clear success criteria. In the PoV, Protopia usually tries a first-pass automation of the whole process. You need to agree on what exactly will be tested: for example the most important and most problematic elements, or the whole process without edge cases and corner cases, that is, the happy path with branches that cause few problems. A sample criterion: in the PoC phase, 80% of cases are handled correctly.&lt;/p&gt;
&lt;p&gt;The client’s business expert rates results such as reports. Some quality traits are very hard to measure: whether a letter reads well always depends on human judgment and is highly subjective.&lt;/p&gt;
&lt;h2 id=&quot;tests-stay-focused-on-the-goal-with-an-ugly-ui-and-logging&quot;&gt;Tests stay focused on the goal, with an ugly UI and logging&lt;/h2&gt;
&lt;p&gt;Tests should stay focused on the goal and in most cases should not be about checking technology or playing with a new framework. For testing, Protopia most often provides a piece of UI, usually as ugly as possible, so that nobody is tempted to deploy it to production the next day. This UI is only minimally usable.&lt;/p&gt;
&lt;p&gt;From the start, you need reasonably good logging across the whole process: how the agent worked, how long it took, what decisions it made and what happened next. In a PoC, a text file is enough for this.&lt;/p&gt;
&lt;p&gt;Logs are a catalog of errors and a roadmap: they show what needs fixing and what not to fix. One example is a discovered edge case that must be fixed if the project moves to a pilot or to production, but not at this point. Tests also show run time. At another client, the agent cannot respond within 30 seconds because there is too much data, and there is no way around that. Generating a report can take 5 minutes, because the agent has to go through several dozen pages before it makes a decision. Such limits are fine, because they help set expectations and build the roadmap, so that the decision to go to production brings no surprises.&lt;/p&gt;
&lt;h2 id=&quot;after-testing-one-of-three-decisions-follows-scale-iterate-or-kill&quot;&gt;After testing, one of three decisions follows: scale, iterate or kill&lt;/h2&gt;

&lt;div class=&quot;table-scroll&quot; role=&quot;region&quot; tabindex=&quot;0&quot; aria-label=&quot;Three states after PoC testing&quot;&gt;&lt;table&gt;&lt;caption&gt;Three states after PoC testing&lt;/caption&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;th&gt;What next&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Scale&lt;/td&gt;
&lt;td&gt;The PoC met its goals: it reached the written success criteria, or the business subjectively confirmed that it wants to go further&lt;/td&gt;
&lt;td&gt;Pilot and production; the experiment becomes a real project&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Iterate&lt;/td&gt;
&lt;td&gt;The result is close to success, and the team knows what to improve and that it needs time for it; a gut feeling that it will be fine is not enough&lt;/td&gt;
&lt;td&gt;Another iteration and new tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kill&lt;/td&gt;
&lt;td&gt;The criteria are not being met and the team is going in circles&lt;/td&gt;
&lt;td&gt;Consider closing the initiative; write down the findings&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;In the pricing process, the iteration was to add a judge agent that verifies the result. This change and another full round of tests needed 3–4 more days.&lt;/p&gt;
&lt;p&gt;Kill is the most hated status, because you cannot announce a success. If the findings are well documented, you know why it failed, and that is a success too. Łukasz Kałużny: “The operation was carried out successfully. The patient died.” (translated)&lt;/p&gt;
&lt;p&gt;The reason may be an idea that was not worth automating, or missing access to an API. Unrealistic expectations are another reason, though this happens less and less often. Adding automation to a system may also turn out too expensive when the cost of building the API eats up the potential gains. Closing protects the team from drowning in a project that would not have worked anyway.&lt;/p&gt;
&lt;p&gt;With good analysis, most of these PoCs end in success, even when the cost summary shows that the project makes no sense.&lt;/p&gt;</content:encoded><dc:creator>Łukasz Kałużny, Mikołaj Szczerbicki</dc:creator><category>AI and agents</category></item><item><title>Observability in practice: a Grafana stack rollout plan and demo</title><link>https://protopia.tech/en/publications/observability-in-practice-plan-demo/</link><guid isPermaLink="true">https://protopia.tech/en/publications/observability-in-practice-plan-demo/</guid><description>When monitoring has grown over time as separate islands and per-user licenses discourage broad access, teams look at different systems during an outage. A plan for rolling out a self-hosted Grafana stack in phases.</description><pubDate>Wed, 11 Mar 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;If you run several environments, your monitoring has probably grown over time without a plan, and during an outage admins and developers look at different data. A self-hosted Grafana stack collects metrics, logs and traces from the whole organization in one place and has no per-user licensing. From a small company upward, the stack rolls out in phases alongside the tools you already have, and your team mainly provides access.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;A self-hosted Grafana stack makes sense in a company bigger than a micro business, while a small team with one monolith is better served by an off-the-shelf tool.&lt;/li&gt;&lt;li&gt;The rollout starts with an inventory and a single entry point in Grafana, because a big-bang change in which people do not move properly to the new tools ends in chaos.&lt;/li&gt;&lt;li&gt;Your team gives access to CI/CD and to the Kubernetes cluster (Protopia can also set up the cluster), then points to the systems, and every system handed over ends with a training session.&lt;/li&gt;&lt;li&gt;Modern applications are instrumented with a few libraries, and purchased systems run through an agent with no code changes.&lt;/li&gt;&lt;li&gt;eBPF profiling shows individual method calls without integrating with the application and with lower overhead than other APM systems, but it works only on Linux.&lt;/li&gt;&lt;li&gt;Zabbix or ELK can stay because Grafana works with them, and the rollout can go in parts, although the recommended option is the complete one, with training.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;a-self-hosted-grafana-stack-makes-sense-beyond-a-micro-business&quot;&gt;A self-hosted Grafana stack makes sense beyond a micro business&lt;/h2&gt;
&lt;p&gt;A self-hosted Grafana stack starts to make a lot of sense in a small or medium-sized organization, because at that size the overhead of data collection and the volume of data are already significant. The stack has no limits tied to the scale of use. In the smallest companies, the math is different:&lt;/p&gt;
&lt;p&gt;Szymon Warda: “If you have a team of 10–20 developers and one monolithic system, get something off the shelf. It will simply be cheaper to maintain.” (translated)&lt;/p&gt;
&lt;p&gt;An off-the-shelf tool then gives a good start, a fast rollout and a fast return on investment.&lt;/p&gt;
&lt;p&gt;An example: an organization with Zabbix, with ELK for security, with a hybrid environment (some on-premises, some cloud), with several clouds and different programming languages. For such a company a self-hosted Grafana stack makes sense, and some observability components already run there.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-observability-when-self-hosted-stack-desktop.CbqCfGhX.webp&quot; alt=&quot;The choice depends on the size of the organization. A team of 10–20 developers with one monolithic system is better served by an off-the-shelf tool, because it is cheaper to maintain. For a small or medium-sized organization with a hybrid environment, several clouds and different programming languages, a self-hosted Grafana stack without per-user licenses makes sense.&quot; width=&quot;736&quot; height=&quot;552&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: a self-hosted Grafana stack or an off-the-shelf tool&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;all-monitoring-goes-to-one-place&quot;&gt;All monitoring goes to one place&lt;/h2&gt;
&lt;p&gt;When a company has several environments, its monitoring has most likely grown over time with nobody organizing it. Each system then has its own island: a developer looks at system A, an admin at system B, sometimes even for the same application. The worst case is when admins report an outage and developers answer that everything works on their side, because they do not see the error.&lt;/p&gt;
&lt;p&gt;The first goal is to revisit monitoring and put it in order: collect information from all systems in one place, so that everyone looks at the same data and draws the same conclusions. Clients often say they have legacy systems and not everything runs on Kubernetes. Kubernetes is the easiest platform to monitor, but it is not the only option. The stack also covers virtual machines and applications on physical servers, and it brings monitoring of the cloud, on-premises and hybrid setups with more than one cloud into one view.&lt;/p&gt;
&lt;p&gt;Monitoring also takes in client applications, databases, queues, cache and even DNS servers, because teams and systems affect each other. The cause of an outage can be an overloaded DNS server, a disk array or an identity server, so data has to come from all these places.&lt;/p&gt;
&lt;p&gt;Prometheus probably runs at your company too, but you likely do not use it to the full. The number of systems it can collect metrics from runs into the thousands, and its community has prepared connectors for almost everything. When a team quickly sets up its own system, the result is generally an empty shell: it collects data from 10–20% of what could be collected and gives no picture of the whole. During an outage, nobody then knows what the root cause was.&lt;/p&gt;
&lt;h2 id=&quot;per-user-licensing-discourages-broad-access&quot;&gt;Per-user licensing discourages broad access&lt;/h2&gt;
&lt;p&gt;Off-the-shelf systems such as Datadog or Dynatrace are usually licensed per user, which discourages opening them to the whole organization. They are relatively easy to deploy, but cost is where the problems start, and they are severe. When 100 or 200 people should use the tool, the seat count starts to hurt.&lt;/p&gt;
&lt;p&gt;At that point, the speakers fairly often see organizations limit who has access. During an outage, support teams are then less sure what they can do, and someone who cannot see something sets up yet another monitoring system. The second common licensing model counts processors or monitored systems. The Grafana stack has no per-user costs, so it can reach much further across the organization.&lt;/p&gt;
&lt;h2 id=&quot;the-rollout-starts-with-an-inventory&quot;&gt;The rollout starts with an inventory&lt;/h2&gt;
&lt;p&gt;The rollout runs in several phases and does not need much involvement from your team. The first phase is an inventory: who uses which tool and what actually runs in the organization. A big-bang change looks good on paper. But if you suddenly take their tools away from admins and from first- and second-line support and do not move them over properly, the organization ends up in chaos.&lt;/p&gt;
&lt;p&gt;In the second step, Grafana becomes the single entry point to all metrics and monitoring data in the organization. It replaces the wiki page with a list of URLs that only some people know about. At this stage, forgotten systems come to light, including those the company pays for and does not use. For each one, a decision follows: remove it, replace it or do something else with it.&lt;/p&gt;
&lt;p&gt;Most guides on the internet say to build the platform at this point and send data to it. First, however, comes Grafana Alloy: collectors that gather data from virtual machines, Kubernetes, the cloud and other sources. Alloy improves the quality of this data, checks what it looks like and buffers it. A very common situation in companies: developers set up Kafka to ship logs, although Kafka is not quite suited to that.&lt;/p&gt;
&lt;p&gt;Once the data from system A looks similar to the data from system B, the monitoring platform goes up. Metrics in Prometheus come first: they show the state of the whole environment, its trends and its daily cycles. On these numbers you can then build alerts and forecasts of when something will stop working.&lt;/p&gt;
&lt;p&gt;The base stack has four components: Alloy collects the data, Prometheus handles metrics and is the market standard, Loki collects logs and Tempo collects traces. Loki takes in terabytes of data, is time- and cost-efficient and needs practically no maintenance.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-observability-grafana-base-stack-desktop.CIaDhToo.webp&quot; alt=&quot;Kubernetes, virtual machines, physical servers and the cloud send data to Grafana Alloy collectors. Alloy improves the quality of the data, watches over it and buffers it. It then passes metrics to Prometheus, logs to Loki and traces to Tempo. These three systems form the base stack, and Grafana is the single entry point to all the data.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the base stack from collectors to Grafana&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;your-team-gives-access-to-the-cluster-and-cicd&quot;&gt;Your team gives access to the cluster and CI/CD&lt;/h2&gt;
&lt;p&gt;Protopia sets up the stack. Your company provides access to a Kubernetes cluster (or Protopia sets up the cluster itself) and access to CI/CD, for example GitHub or Bitbucket. The whole stack is built from code to make it resilient. That is where your team’s involvement in this part ends.&lt;/p&gt;
&lt;p&gt;From there, the work goes in phases. You point to the systems, PoCs follow, and Protopia sets up collectors, prepares load tests to check that the system will work, and connects everything to the main Grafana instance.&lt;/p&gt;
&lt;p&gt;Every system handed over for use comes with a training session. People with a full backlog have no time to learn on their own, so it is easier to run a one-day, 4-hour training session and show what they can get out of the tool. This encourages employees to use the new stack instead of the old tools.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-observability-division-of-work-desktop.D2aM0Xp-.webp&quot; alt=&quot;Your team gives access to the Kubernetes cluster, or Protopia sets up the cluster, and access to CI/CD. Protopia builds the stack from code. Then your team points to the systems. For each system, Protopia runs a PoC, sets up collectors, prepares load tests and connects the system to Grafana. The cycle ends with a training session and returns to the next system.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: what your team provides and what Protopia does&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The stack is known and used by the CNCF and by organizations around the world. A new hire is expected to know these tools already, so you do not have to teach them the tools themselves, and the company builds on a standard it can extend.&lt;/p&gt;
&lt;h2 id=&quot;developers-get-instrumentation-standards&quot;&gt;Developers get instrumentation standards&lt;/h2&gt;
&lt;p&gt;Next come the standards the organization requires from application developers. Protopia interviews several groups of developers, learns how they work and where they start from, prepares recommended standards and, if needed, splits the path to them into phases.&lt;/p&gt;
&lt;p&gt;Monitoring works without integration with the applications and without developer support, but more value comes only when the applications take part. For modern technologies such as Java, .NET, Python or JavaScript, the effort is half an hour to 3 hours per system: you add a few libraries and it works. Developers have probably tried already, and the libraries may be under-configured, so small configuration changes make a big difference.&lt;/p&gt;
&lt;p&gt;For systems that you bought, no longer maintain or no longer develop, there is auto-instrumentation. &lt;strong&gt;Auto-instrumentation means starting an application through an agent, with no changes to its code.&lt;/strong&gt; It does not give everything and it is not perfect, but it gives a lot: telemetry, tracing and metrics. It is very safe, and sometimes the only change is the target image the application runs in.&lt;/p&gt;
&lt;h2 id=&quot;team-knowledge-goes-into-alerts&quot;&gt;Team knowledge goes into alerts&lt;/h2&gt;
&lt;p&gt;Standards are also written for admins and DevOps teams. When developers, admins and DevOps look at the same deep data, you can ask when a system is about to stop working and what the main reason for its problems is.&lt;/p&gt;
&lt;p&gt;Tribal knowledge shared in a meeting evaporates quickly, so the knowledge of these people goes into alerts and predictive metrics, and over time the system knows more. Much of this knowledge does not have to be gathered from people. Most popular systems, such as PostgreSQL or Redis, have ready-made or official alert sets, frequently promoted by the vendors themselves. Such sets often have dozens, if not hundreds, of rules. Some Kubernetes deployments load alert rules by default, and you only configure who gets notified and when. This knowledge is available because a great many organizations use the stack.&lt;/p&gt;
&lt;h2 id=&quot;drilldown-in-grafana-shows-errors-traffic-and-response-time&quot;&gt;Drilldown in Grafana shows errors, traffic and response time&lt;/h2&gt;
&lt;p&gt;The most important screen in the demo is Drilldown on tracing data. It shows the number of errors, the number of requests and the distribution of their duration. &lt;strong&gt;RED stands for Request, Error, Duration.&lt;/strong&gt; It is one of the three basic views of any system. The operator does not need to know where the data comes from and how it flows: they get a ready way to look at the system.&lt;/p&gt;
&lt;p&gt;The errors tab separates errors your customers see from errors the system swallows. The latter do not surface, but they are worth dealing with. Next, you see how many errors each system produces, and root cause errors: the statistical source of errors for each service over the last half hour. The cause usually sits lower in the call chain than the system that failed.&lt;/p&gt;
&lt;p&gt;The comparison sets the current behavior of the system against a previous period. Chasing every error probably does not pay off, because something always happens, so what counts is what has changed. Traces show the path of a request through the whole call chain: where something failed, which services it affected and whether the user saw it. The durations tab points to the main source of latency and automatically picks out the requests that really were slow.&lt;/p&gt;
&lt;h2 id=&quot;ebpf-profiling-extends-the-base-stack&quot;&gt;eBPF profiling extends the base stack&lt;/h2&gt;
&lt;p&gt;Use of eBPF-based profiling is growing fast in very large organizations. Monitoring shows days and hours, logs and traces show minutes. Profiling goes below 20 milliseconds, for example when developers analyze performance or when something has failed. It works in real time, only on Linux, and best on Kubernetes. It is possible on virtual machines too, but it is not recommended, because it does not go as smoothly as it should.&lt;/p&gt;

&lt;div class=&quot;table-scroll&quot; role=&quot;region&quot; tabindex=&quot;0&quot; aria-label=&quot;eBPF profiling vs. other APM systems&quot;&gt;&lt;table&gt;&lt;caption&gt;eBPF profiling vs. other APM systems&lt;/caption&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;eBPF profiling&lt;/th&gt;
&lt;th&gt;Other APM systems&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Performance overhead&lt;/td&gt;
&lt;td&gt;1–3%&lt;/td&gt;
&lt;td&gt;usually 5–10%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enabling&lt;/td&gt;
&lt;td&gt;nothing to enable&lt;/td&gt;
&lt;td&gt;often must be enabled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sampling depth&lt;/td&gt;
&lt;td&gt;down to individual method calls&lt;/td&gt;
&lt;td&gt;often much worse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Application awareness&lt;/td&gt;
&lt;td&gt;the application does not know about it&lt;/td&gt;
&lt;td&gt;the application must know about it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;A developer understands a detailed call stack, an operator does not. The Explain flame graph button sends the flame graph data to a language model, in the demo to OpenAI. The model explains what happens in a given piece of code, for example that it fails or uses a lot of memory, and suggests changes based on best practices. The answer is not perfect, but it is very useful: the operator does not need to know the code, and a developer who maintains a system they do not know well gets a concrete answer. The stack does not tie you to one model and works with Anthropic and with any system compatible with the OpenAI API.&lt;/p&gt;
&lt;h2 id=&quot;grafana-connects-to-databases&quot;&gt;Grafana connects to databases&lt;/h2&gt;
&lt;p&gt;With controlled access, databases can be connected to Grafana. A familiar outage scenario: the database throws an exception and someone has to find a DB admin to check what is going on. This feature has to be built and secured. When a trace shows an exception because a command did not execute, one button runs a read-only query on that database. The path from problem to answer then shrinks to seconds instead of hours spent looking for a person with permissions.&lt;/p&gt;
&lt;h2 id=&quot;the-rollout-can-go-in-parts&quot;&gt;The rollout can go in parts&lt;/h2&gt;
&lt;p&gt;Monitoring changes can roll out in phases and in parts: start with cleanup or with better visibility, then add further elements as needed. The recommended option is the complete one: the rollout, real onboarding and a training package, so that people know how to get value out of what has been built.&lt;/p&gt;
&lt;p&gt;Grafana also works with many existing systems, including Datadog, so the move to one central place can be smooth. Tools deeply rooted in the company, such as Zabbix or ELK, do not have to go, because Grafana and Prometheus work with them. The change is evolutionary: the goal is to win people over with value they can see, without forcing anything.&lt;/p&gt;</content:encoded><dc:creator>Szymon Warda, Mikołaj Szczerbicki</dc:creator><category>Observability</category></item><item><title>Observability for managers: how to know the business is running</title><link>https://protopia.tech/en/publications/observability-for-managers/</link><guid isPermaLink="true">https://protopia.tech/en/publications/observability-for-managers/</guid><description>CPU and memory look fine, yet the business reports that the system is down. See how observability differs from monitoring, what it gives a manager and what a rollout requires.</description><pubDate>Wed, 18 Feb 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;If your infrastructure is monitored and you still hear about an outage when the business calls you, CPU and memory monitoring is not enough. Observability gives the business, developers and operations the same data about what happens in the systems, so problems show up earlier and numbers settle disputes. You will learn when optimization and higher availability do not pay off, what to do with an expensive vendor platform and what a rollout on an open standard requires.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;Resource monitoring does not tell you whether the business is running, while observability lets you see a problem before it starts, so the team leaves firefighting mode and can warn the business itself.&lt;/li&gt;&lt;li&gt;Separate tools in each team have to be bought, learned and maintained, and during an outage one side sees it while the other may not.&lt;/li&gt;&lt;li&gt;Numbers show how many users the system is slow for and which machines can be downsized, so you know where to spend money and where to stop.&lt;/li&gt;&lt;li&gt;When the bill for a commercial platform is high, an alternative is OpenTelemetry with the Grafana stack, which your SRE team probably already knows, although the rollout requires changes on the application side.&lt;/li&gt;&lt;li&gt;Tools come first, then processes: alerts, anomaly detection and logging standards, which a platform team can roll out in the background.&lt;/li&gt;&lt;li&gt;Instead of full availability, you agree on a measurable commitment with the business and check which systems can occasionally be down.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;monitoring-sees-resources-observability-sees-system-behavior&quot;&gt;Monitoring sees resources, observability sees system behavior&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Monitoring is&lt;/strong&gt; checking whether the CPU, the memory and maybe also the disk are full. A manager asks something else: whether their business is running. Monitoring reports an outage or, at best, warns before one.&lt;/p&gt;
&lt;p&gt;The speakers have often heard companies say they have some monitoring or observability in place, and yet they almost always find the same situation. The business calls the manager, often quite angrily, to say the system has been down for half an hour. The operations team checks the CPU and RAM, sees that they are fine and replies that the system seems to be working.&lt;/p&gt;
&lt;p&gt;Observability is not a trendy name for monitoring. &lt;strong&gt;Observability is&lt;/strong&gt; a goal: a state in which you can infer from system data how the system behaves and what happens inside it. It covers systems and processes. With it, you see that something is happening, or even see it before it starts.&lt;/p&gt;
&lt;h2 id=&quot;logs-metrics-and-traces-are-successive-levels-of-maturity&quot;&gt;Logs, metrics and traces are successive levels of maturity&lt;/h2&gt;
&lt;p&gt;Everyone has logs, better or worse. Sometimes developers log in to machines and read the scrolling logs like in The Matrix. That is a skill, but at a larger scale not necessarily a useful one.&lt;/p&gt;
&lt;p&gt;Metrics give real numbers, and you can collect a great many of them. &lt;strong&gt;A trace is&lt;/strong&gt; a record of how a whole business process flows through the services and microservices in the organization. Traces show which calls between services work and which do not, their latency, where they fail and how response times are distributed. With them, you know where it pays to invest time in optimization.&lt;/p&gt;
&lt;h2 id=&quot;a-team-in-firefighting-mode-does-not-deliver-its-plan&quot;&gt;A team in firefighting mode does not deliver its plan&lt;/h2&gt;
&lt;p&gt;A team that reacts to outages works in firefighting mode. It has to restore the system immediately, usually at the expense of quality. The problem comes back, because the team does what is needed right away and then usually has no time to fix it properly.&lt;/p&gt;
&lt;p&gt;In this mode, nothing can be planned. The team announces something for the quarter, the quarter ends and the work has not even started. The team did a lot, but it was not valuable work.&lt;/p&gt;
&lt;p&gt;With observability, the team invests in operational processes, has more planned time, and there are simply fewer outages. That leaves time to show developers that something runs slowly or is misused, or that the application performs poorly.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-observability-firefighting-mode-desktop.C0WbauzW.webp&quot; alt=&quot;A comparison of two ways an operations team works. In firefighting mode, an angry call from the business leads to a quick fix right away, at the expense of quality. There is no time for a proper fix, so the problem comes back, the cycle repeats and the plan for the quarter stalls. With observability, the team sees the problem early, warns the business itself and has time for planned work, and there are fewer outages.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: firefighting mode and working with observability&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;When the team knows that an outage is coming, it can call the business itself: something is happening with the system, we are working on it. Even in a serious outage, that is a very different situation from an angry call from the business.&lt;/p&gt;
&lt;h2 id=&quot;separate-tools-in-every-team-are-an-extra-cost&quot;&gt;Separate tools in every team are an extra cost&lt;/h2&gt;
&lt;p&gt;Administrators, developers, ops and SRE teams work on their own pieces of the system. Everyone looks at their own part and everyone says that things are fine. During an outage, a DevOps engineer sits down with a developer, each has their own truth, and they cannot agree. One side sees the outage, the other may not see it.&lt;/p&gt;
&lt;p&gt;Each team has its own tools. You have to pay for them, people have to learn them, and someone has to maintain them. A local debugger is no longer enough for developers. In recent years, they have come to expect numbers on how the system really behaves. If they maintain their tools themselves, they lose time on it, and software development and delivery slow down.&lt;/p&gt;
&lt;p&gt;The direction is centralization. With unified tools, development, operations and monitoring work from the same numbers and speak one language. Few people will argue with numbers that the organization collects itself.&lt;/p&gt;
&lt;h2 id=&quot;numbers-show-what-to-optimize-and-where-to-downsize-machines&quot;&gt;Numbers show what to optimize and where to downsize machines&lt;/h2&gt;
&lt;p&gt;A good level of observability lets you talk about what is happening rather than about impressions, because dashboards and numbers describe every situation. The manager gets an argument to use with every team.&lt;/p&gt;
&lt;p&gt;In a large organization, someone will always complain that the system is slow. With numbers, you see how many people it is slow for, for example 2% of users. Then you decide whether it is worth spending money to improve that result. Szymon Warda: “Not every system is worth optimizing to the limit.” (translated)&lt;/p&gt;
&lt;p&gt;Numbers also bring order to infrastructure decisions. For a new system, a team requests 20 CPUs and 100 GB of RAM, and then it turns out the system uses 50% of that. When a month of charts shows that CPU, memory and disk never go above a certain value, the conclusion is: let’s use smaller machines. That means very large savings, and the more infrastructure you have, the easier they are to get.&lt;/p&gt;
&lt;h2 id=&quot;commercial-observability-platforms-are-a-significant-cost&quot;&gt;Commercial observability platforms are a significant cost&lt;/h2&gt;
&lt;p&gt;A vendor observability system costs 5–6 figures in dollars a year. For a mid-sized team, that is a fairly significant cost, and the amount depends on the size of the organization.&lt;/p&gt;
&lt;p&gt;A large financial-sector client, unnamed here, already pays one of the big vendors a bill of 7 and 8 figures in dollars a year. The client saw what the platform Protopia recommends can do and is considering leaving that vendor. A few years ago, industry publications described organizations that paid 15 and 50 million dollars a year for observability systems.&lt;/p&gt;
&lt;h2 id=&quot;the-grafana-stack-is-the-market-standard&quot;&gt;The Grafana stack is the market standard&lt;/h2&gt;
&lt;p&gt;The platform Protopia recommends is built on the Grafana stack and integrates natively with Kubernetes. When you propose it to your developers, they will probably say they know it. Your SRE team knows which system you mean, because it uses it. If you look, you will find it in your organization, only probably less well maintained and not wired into processes.&lt;/p&gt;
&lt;p&gt;These systems are the market standard in Kubernetes and at the cloud giants, the hyperscalers. The open-source tools in this stack scale very well, are cost-effective and are easy to maintain. You do not depend on a closed tool, so you draw on the knowledge of a huge user community, and you implement some good practices with ready-made components developed in the ecosystem. Open source has its own problems too.&lt;/p&gt;
&lt;h2 id=&quot;a-rollout-requires-work-on-the-application-side&quot;&gt;A rollout requires work on the application side&lt;/h2&gt;
&lt;p&gt;Observability requires some work and changes on the application side, although these changes keep getting smaller. The more the team does, the more value it gets, and this is long-term work. It is worth doing once and doing it right, because OpenTelemetry has become the standard in observability. Big vendors also use it to integrate with applications.&lt;/p&gt;
&lt;p&gt;Do not expect everything to work on its own. One client turned on automatic observability from a very well-known solution on its cluster, and the whole system went down. The mechanism hooked into a place it should not have, and the system stopped working immediately. Finding the cause took two days. “Magic” solutions often work this way: they promise too much and then fail to deliver, because they cannot.&lt;/p&gt;
&lt;p&gt;The goal is a business that runs predictably. You do not want a situation where you upgrade an operations tool and your business is down for a whole day, and nobody knows why. Szymon Warda: “We know full well that rolling out any system has a cost.” (translated)&lt;/p&gt;
&lt;h2 id=&quot;alerts-can-warn-of-an-outage-charts-show-it-after-the-fact&quot;&gt;Alerts can warn of an outage, charts show it after the fact&lt;/h2&gt;
&lt;p&gt;A falling chart shows an outage that has already happened. Organizations sometimes have 4, 5, 6 or more monitors with nice charts. They look like a Christmas tree for show, and nobody uses them.&lt;/p&gt;
&lt;p&gt;Alerts let you act earlier: they can notify you that something will happen in a day, in two days or at some later point. Observability also includes a change in the organization’s habits. First the technical part provides the capabilities, then the team introduces processes, for example alerting and anomaly detection, which catches signals that something is starting to break.&lt;/p&gt;
&lt;p&gt;Systems sometimes produce tens of thousands of metrics, so you have to know how to filter them. Without that, it is drinking from a fire hose, and you can lose your teeth.&lt;/p&gt;
&lt;h2 id=&quot;developers-see-how-their-code-runs-in-production&quot;&gt;Developers see how their code runs in production&lt;/h2&gt;
&lt;p&gt;One unified tool makes shift left possible: developers see how the system behaves in production. The “it works on my machine” argument disappears, because there are numbers from production, including the SQL queries being executed. The data shows that something needs fixing, or that everything is fine and further performance work makes no sense, because the system runs very well there.&lt;/p&gt;
&lt;p&gt;The feedback loop gets shorter. The team deploys a new version and a day later sees whether it works, does not work or runs slowly, instead of learning about it after a quarter. Measurement shows whether the new version is better or worse, so the team detects a regression much faster and knows exactly what happened.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-observability-feedback-loop-desktop.O2ZHsNl-.webp&quot; alt=&quot;The feedback loop after a deployment. The team deploys a new version to production and a day later has numbers from production, including the SQL queries being executed. The measurement shows whether the version is better or worse. On a regression, the team fixes the code and deploys another version. When the system works well, further performance work makes no sense.&quot; width=&quot;736&quot; height=&quot;552&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the feedback loop with numbers from production&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;In critical systems, where it makes sense, you can deploy a new version partially and watch how the traffic behaves. Such deployments do not pay off everywhere.&lt;/p&gt;
&lt;h2 id=&quot;sla-slo-and-sli-define-what-the-team-commits-to&quot;&gt;SLA, SLO and SLI define what the team commits to&lt;/h2&gt;
&lt;p&gt;Without numbers, you cannot measure whether an investment in improving a metric brings value. With numbers, you can set an SLA, SLOs and SLIs and state the service level the team commits to.&lt;/p&gt;
&lt;p&gt;Some companies say they must have 100% uptime. That is unrealistic. To put it jokingly, every additional nine after the decimal point is another zero before the decimal point in cost, sometimes even more, sometimes not.&lt;/p&gt;
&lt;p&gt;With such a commitment, you can measure whether the contract with the business is met. Potential outages and maintenance windows can be translated into the organization’s real finances. Sometimes it turns out that some systems can be down for a certain period without any problem, while others are more critical.&lt;/p&gt;
&lt;p&gt;The same measurements remove friction between development teams that say their system fails because another one does. A team commits that its system runs within the agreed SLA, and development stops being about complaints that something does not work for someone.&lt;/p&gt;
&lt;h2 id=&quot;a-platform-team-rolls-out-standards-in-the-background&quot;&gt;A platform team rolls out standards in the background&lt;/h2&gt;
&lt;p&gt;The organization does not have to decide on its own, before it starts, what to monitor and how. There are universal standards and good practices to adopt instead of reinventing the wheel. Every organization is different to some degree, so the standards need adjusting.&lt;/p&gt;
&lt;p&gt;If you have a platform team, it rolls out further standards in the background: how to log, which metrics to collect, which error levels to report and what goes into the logs. When GDPR came in, teams had a lot of work removing first names, last names and user-identifying data from logs. That was a one-time effort, and now these rules go into the standard.&lt;/p&gt;
&lt;p&gt;Developers do not have to think about it: a new system gets the agreed components and libraries and reports in the agreed way. A costly piece of system development goes away. The standard is current and industry-wide.&lt;/p&gt;
&lt;p&gt;Data from legacy, new, evolving and external systems can be collected into one pipeline and reported in a unified way. Instead of pieces in different tools, you see the whole organization and know how each part behaves and where.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-observability-one-pipeline-desktop.CtSUmaUh.webp&quot; alt=&quot;Data from legacy, new, evolving and external systems flows through one pipeline to a platform on the Grafana stack, which reports it in a unified way. The business, developers and operations see the same numbers and the whole organization instead of pieces in different tools.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: data from every system in one pipeline&lt;/figcaption&gt;&lt;/figure&gt;</content:encoded><dc:creator>Szymon Warda, Mikołaj Szczerbicki</dc:creator><category>Observability</category></item><item><title>AI agent factory: a shared platform in a large organization</title><link>https://protopia.tech/en/publications/ai-agent-factory-platform/</link><guid isPermaLink="true">https://protopia.tech/en/publications/ai-agent-factory-platform/</guid><description>When each AI agent is built separately, you have to set up the infrastructure, open network traffic and certify the environment all over again. A shared platform on Azure is approved for use once, and each new agent is deployed to it through CI/CD.</description><pubDate>Wed, 04 Feb 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;If your company is going to build more and more AI agents, you have to decide whether to build each one separately or to set up a shared platform. An agent factory is that kind of platform, built on Azure AI Foundry, API Management and Logic Apps: separate environments for experiments, tests and production, security approved once, and integrations you connect once and share with later agents. A team with an approved idea gets a secured environment quickly.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;When the company strategy expects more and more agents and APIs, a shared platform is probably the only sensible direction, while the first agent for the board can likely be built without it.&lt;/li&gt;&lt;li&gt;AI Sandbox is a standardized, secured place to test a business hypothesis, and the agent moves from there through the pre-production environment to production; it is deployed from a Git repository through CI/CD like any other software.&lt;/li&gt;&lt;li&gt;In the enterprise variant, thread data sits in resources in the organization&apos;s own subscription, encrypted with its keys.&lt;/li&gt;&lt;li&gt;API Management exposes company systems to agents on one facade, with a description the model understands, so you do not have to write custom MCP gateways or change the source systems.&lt;/li&gt;&lt;li&gt;When a process is truly deterministic, Logic Apps very often turns out more predictable than custom code or agents, and the agent comes in where generative analysis is needed.&lt;/li&gt;&lt;li&gt;Before platform work starts, Protopia tries to hold a workshop that picks, from the client&apos;s business cases, the ones that fit the platform well.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;a-platform-makes-sense-when-agents-and-apis-keep-growing&quot;&gt;A platform makes sense when agents and APIs keep growing&lt;/h2&gt;
&lt;p&gt;If the company strategy expects more and more agents and more and more APIs, a shared platform is probably the only sensible direction. The first agent, meant to show the board a short-term win, can likely be built in a couple of weeks.&lt;/p&gt;
&lt;p&gt;Here, an agent factory means a rollout in a large organization, including one constrained by regulations.&lt;/p&gt;
&lt;p&gt;The team does not set up the whole infrastructure every time, so it starts using agents in a standardized way sooner. Once the platform is approved, there is no need to open network access and recertify the environment for each business case. Security and consistency, for example logs, are solved once.&lt;/p&gt;
&lt;p&gt;A shared platform also keeps knowledge in one place. At several clients, teams test agents instead of yet more open-source frameworks, which keep multiplying and keep changing how they work.&lt;/p&gt;
&lt;h2 id=&quot;an-agent-moves-through-three-environments&quot;&gt;An agent moves through three environments&lt;/h2&gt;
&lt;p&gt;The platform has three environments: AI Sandbox, a pre-production environment and production. AI Sandbox is the place for safe experiments and innovation. In many companies a sandbox means a free-for-all where users do whatever they want. Here it is a secured, shared environment, standardized and subject to corporate rules, possibly with access to production data. It is hard to build agents, and AI in general, on anonymized and synthetic data.&lt;/p&gt;
&lt;p&gt;In the sandbox, the team checks the proof of value. &lt;strong&gt;Proof of value is&lt;/strong&gt; a test of the business hypothesis: whether the scenario will run on agents and work. Writing and testing code takes a back seat.&lt;/p&gt;
&lt;p&gt;A scenario that proves itself moves to the pre-production environment and from then on follows a production lifecycle. There the team checks that all deployments and integrations work and that the agent is fully integrated and ready for production.&lt;/p&gt;
&lt;p&gt;Each environment has four elements: access channels, the agent runtime, the integration part, and the operations and maintenance part.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-agent-factory-architecture-desktop.vRtvQxcv.webp&quot; alt=&quot;Architecture of one agent factory environment in four parts. Access channels, that is a button in a system, an event or queue, and a chat on a website or in Teams, call the agent in an Azure AI Foundry project. The agent uses APIs and MCP through API Management, which exposes backend systems and external MCP on one facade with a description the agent understands. Logic Apps calls APIs in API Management and can also call the agent. The operations and maintenance part covers the Git repository, CI/CD scripts, Application Insights and content filtering.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the elements of one agent factory environment&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-access-channel-is-separate-from-the-agent&quot;&gt;The access channel is separate from the agent&lt;/h2&gt;
&lt;p&gt;The agent does not depend on how you reach it. It can be standalone, part of a mobile app, a chat on a website or a background process. The platform lets you call the agent in any way, and the call itself is an external integration:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;synchronously from a business system, for example with a “generate a report for me” button;&lt;/li&gt;
&lt;li&gt;asynchronously through an event, for example a message, a queue or an email to a monitored mailbox;&lt;/li&gt;
&lt;li&gt;through a chat on a website, in Teams or in another application.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The calling method stays outside the platform so that it does not impose a solution. A proof of value can show that the agent needs to be called in a completely different way. Keeping the channel separate is meant to avoid hard rewiring and a later rebuild.&lt;/p&gt;
&lt;p&gt;One of the recent client cases is an analysis of a counterparty’s financial results. In tests, an analyst pastes the results into a chat window and checks whether the agent answers as expected. The target is a “Run the analysis for me” button in the system that starts the agent. The agent synchronously queries the backends and APIs, collects and analyzes the data, and returns only the final result.&lt;/p&gt;
&lt;h2 id=&quot;agents-run-in-azure-ai-foundry-projects&quot;&gt;Agents run in Azure AI Foundry projects&lt;/h2&gt;
&lt;p&gt;The agent runtime is Azure AI Foundry, in the Microsoft ecosystem.&lt;/p&gt;
&lt;p&gt;The unit of work in AI Foundry is a project, the equivalent of an initiative. Technically it resembles a namespace in Kubernetes or a resource group in Azure. One project can hold many agents built by one team, and agents in a project can talk to each other.&lt;/p&gt;
&lt;p&gt;A project uses models from a defined list: OpenAI and the supported open-source models from the marketplace. Competing vendors’ models are not on the list. Protopia works with clients on OpenAI 99% of the time.&lt;/p&gt;
&lt;h2 id=&quot;an-orchestrator-agent-connects-small-agents&quot;&gt;An orchestrator agent connects small agents&lt;/h2&gt;
&lt;p&gt;Because of how language models work, Protopia usually builds small agents, each for a specific task or part of a business process. Above them runs an orchestrator agent that calls the smaller agents. The orchestrator can get a more expensive reasoning model, and the small agents a cheaper one, for example GPT-4o mini.&lt;/p&gt;
&lt;p&gt;In most cases Protopia uses the connected agents scenario instead of a chain of fully independent agents: a central agent plus small agents, with the steps orchestrated deliberately. In AI Foundry you literally specify which agent connects to which.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-agent-factory-orchestrator-desktop.BKd6rORy.webp&quot; alt=&quot;Connected agents in one Azure AI Foundry project. The access channel passes the thread to the orchestrator agent, which has a more expensive model with reasoning. Through a handoff, the orchestrator passes the thread to small agents with a cheaper model, for example GPT-4o mini, each for a specific task, and waits for their answer.&quot; width=&quot;736&quot; height=&quot;552&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: an orchestrator agent and small agents in one project&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;A thread is&lt;/strong&gt; an agent’s conversation about one matter. &lt;strong&gt;A handoff is&lt;/strong&gt; a temporary transfer of the thread to another agent: the agent waits for the answer and continues processing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A2A (agent to agent) is&lt;/strong&gt; an open-source protocol for communication between agents, started by Google. It is available in AI Foundry, but it is not worth tying yourself closely to it, because agents usually connect within one project.&lt;/p&gt;
&lt;h2 id=&quot;in-the-enterprise-variant-thread-data-stays-in-your-subscription&quot;&gt;In the enterprise variant, thread data stays in your subscription&lt;/h2&gt;
&lt;p&gt;Thread history lets you later resume an asynchronous thread and collect its results. In a less enterprise-grade variant, when you do not need this, the service keeps the data itself. In the enterprise variant, which Protopia usually deploys, the data goes to separate services that you have to deploy:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;storage for files;&lt;/li&gt;
&lt;li&gt;Cosmos DB for conversations and agent data;&lt;/li&gt;
&lt;li&gt;Azure AI Search, with inverted indexes and a vector database for text or semantic search across texts, documents and attachments, the basis of RAG.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Bring your own resources is&lt;/strong&gt; a model in which the organization brings its own storage, Cosmos DB and Azure AI Search, deployed in its own subscription. You see them and control them. The resources are deployed almost entirely on the organization’s terms, so they can match its internal regulatory requirements, for example secure communication and independent logs.&lt;/p&gt;
&lt;p&gt;Conversation history can contain sensitive and protected data, for example in healthcare, pharma or banking. That is why the resources are encrypted with the organization’s own keys (customer-managed keys, CMK), in line with the rules the organization has set.&lt;/p&gt;
&lt;h2 id=&quot;api-management-is-the-central-integration-point&quot;&gt;API Management is the central integration point&lt;/h2&gt;
&lt;p&gt;In this architecture, API Management is the heart of the platform, because an agent without APIs is just a chat. It exposes the organization’s APIs consistently, on one facade, so that agents can consume them easily.&lt;/p&gt;
&lt;p&gt;A business system connected once through an API is available to every agent you connect it to. An agent does not automatically get access to everything: authentication and authorization decide which API a given agent talks to.&lt;/p&gt;
&lt;p&gt;The API has to be exposed differently from the backend system. Systems delivered by outside vendors or by an internal department are often poorly described. An agent does not know from memory what a given method does, and the method description alone does not guarantee it. That is why you add a description in API Management that the agent understands, so it knows which endpoint solves its problem. This works for both OpenAPI Schema and MCP (Model Context Protocol): every API, call and expected response is described in natural language tailored to the model.&lt;/p&gt;
&lt;p&gt;A chat model writes the description, with a system prompt that tells it the text is for its own later use.&lt;/p&gt;
&lt;p&gt;API Management also translates on the fly, for example it exposes a SOAP API as REST. With API Management you do not need to write your own MCP gateway for a system or, worse, modify the source system to expose endpoints for agents.&lt;/p&gt;
&lt;p&gt;API Management has two new capabilities. Any existing API, including one that systems outside the agent world may already use, can be turned into MCP with practically one switch. Recently you can also move external MCP servers onto the facade: third-party ones, or ones someone in the company built as MCP from the start, without an API. The organization controls them like its own, with extra security, policies and orchestration.&lt;/p&gt;
&lt;h2 id=&quot;logic-apps-handles-the-predictable-steps&quot;&gt;Logic Apps handles the predictable steps&lt;/h2&gt;
&lt;p&gt;If a scenario is truly deterministic, meaning you know the process and how to connect it, Logic Apps very often turns out more precise and predictable, and also simpler to build, than custom code or agents.&lt;/p&gt;
&lt;p&gt;Teams often end up with narrowly specialized, low-level APIs and an API Management instance that exposes them, but no glue to orchestrate them, for example: call an API, get the response, send three emails or a text message, wait for approval and confirm on another endpoint. In this architecture the glue is Logic Apps, Microsoft’s no-code tool, which you can also manage from code. Logic Apps together with API Management make up iPaaS (Integration PaaS).&lt;/p&gt;
&lt;p&gt;You can build a Logic Apps flow by clicking, as in Power Apps, but there are more system and business connectors: for Oracle, SAP, CRM or Salesforce. A flow can also call an API in API Management or an agent. There are, roughly, hundreds of ready-made integrations and actions. Control flow covers conditions (IF), loops (FOR), error handling and resuming integrations.&lt;/p&gt;
&lt;p&gt;The agent comes in only where you need generative analysis, something new created, or facts combined on the fly.&lt;/p&gt;
&lt;p&gt;In some scenarios the entry point is a workflow in Logic Apps. Marek Grabarz: “it’s not that the agent triggers the workflow, it’s the workflow that calls the agent.” (translated) Based on the agent’s result, the workflow then does something predictable. The access channel is then the integration in iPaaS.&lt;/p&gt;
&lt;h2 id=&quot;agents-are-deployed-like-code&quot;&gt;Agents are deployed like code&lt;/h2&gt;
&lt;p&gt;An agent is a piece of software like any other, so the operations and maintenance part rests on the DevOps approach. There are no special rules; only some technical details change.&lt;/p&gt;
&lt;p&gt;Agents live in a Git repository: GitHub, Bitbucket, GitLab or another one the organization already has. Prepared CI/CD scripts deploy the agents to the next environments.&lt;/p&gt;
&lt;p&gt;The agent file in the repository holds the agent definition with the system prompt, plus the connected integrations and other agents. You can treat it like an infrastructure as code manifest. In this case these are Python scripts, so the file is also the agent definition in Python code. Once the platform runs, a deployment often comes down to copy and paste, or help from GitHub Copilot or another coding agent.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-agent-factory-deployment-desktop.DnjJ6Uc9.webp&quot; alt=&quot;The agent file in a Git repository holds the agent definition with the system prompt, the connected integrations and the connected other agents. CI/CD scripts deploy it to three environments. In AI Sandbox the team tests the business hypothesis. A scenario that proved itself moves to the pre-production environment, where the team checks deployments and integrations, and then to production.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: an agent from a Git repository to the next environments&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;observability-shows-how-each-thread-went&quot;&gt;Observability shows how each thread went&lt;/h2&gt;
&lt;p&gt;When something goes wrong in an agent, or someone reports a wrong answer, the team checks the thread. AI Foundry stores the full conversation threads of every agent in a database until you delete them: the input, all the tools called, the content of requests to other systems, integrations and agents, and the responses. The thread is stateful. In Azure AI Foundry Studio you can find it, go back to it, view it in the playground and check what happened.&lt;/p&gt;
&lt;p&gt;You turn on observability in a project with almost no deployment work. Application Insights, a standard APM tool from Microsoft, tracks the agent lifecycle: questions, answers and calls, across the whole conversation thread from a chat or the call thread from another system.&lt;/p&gt;
&lt;p&gt;AI Foundry also has a quality evaluation area: another model judges whether the whole flow went well. Response evaluation metrics are starting to appear, but the market is still at an early stage, because one non-deterministic thing is measured by another.&lt;/p&gt;
&lt;p&gt;Content filtering works from the start and protects against users putting something dangerous into the agent. The organization decides which content it allows. Some companies may want to loosen the filters, for example in e-commerce with adult content.&lt;/p&gt;
&lt;h2 id=&quot;usage-and-complaints-measure-the-platform&quot;&gt;Usage and complaints measure the platform&lt;/h2&gt;
&lt;p&gt;The first measure of the platform and the agent use cases is whether anyone uses them after rollout.&lt;/p&gt;
&lt;p&gt;The second measure applies to each agent separately: how many complaints and bug reports there are, for example about answer quality, relative to the number of uses. A good result is usage that does not drop and no complaints about quality.&lt;/p&gt;
&lt;h2 id=&quot;a-workshop-comes-first-where-possible-and-the-sandbox-should-be-ready-the-next-business-day&quot;&gt;A workshop comes first where possible, and the sandbox should be ready the next business day&lt;/h2&gt;
&lt;p&gt;Before work on the platform starts, Protopia tries to organize a workshop. The client prepares business cases for it. Business architects are invited, not necessarily operations people, because it is not that stage yet. At the workshop, the team points out the cases that fit the platform well and will integrate well. With 5, 10 or 15 such examples, the incentive to build the platform is much bigger. Łukasz Kałużny: “it’s an important point, so that we don’t build a platform for the sake of building a platform.” (translated)&lt;/p&gt;
&lt;p&gt;Once an initiative is approved for tests in AI Sandbox, the team should have access the next business day. The platform team can, for example, provide such an environment in a dozen or so minutes. Self-service is possible, but it is not the default here. The time counts from approval and does not include the company’s organizational processes.&lt;/p&gt;</content:encoded><dc:creator>Łukasz Kałużny, Marek Grabarz</dc:creator><category>AI and agents</category></item><item><title>API Management at PZU: a shared gateway to its systems and a plan for AI agents</title><link>https://protopia.tech/en/publications/from-api-management-to-ai-agents-pzu/</link><guid isPermaLink="true">https://protopia.tech/en/publications/from-api-management-to-ai-agents-pzu/</guid><description>PZU operates in a regulated industry, is leaving its SOA platform and is planning for AI agents. It built a shared gateway on Azure API Management that developers use on their own.</description><pubDate>Wed, 07 Jan 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;If your company&apos;s integrations still run on an SOA platform and some APIs bypass the integration platform, you can move to a shared API gateway as PZU did and keep the data flow in your own infrastructure. At PZU that gateway is Azure API Management, and developers publish their APIs through it on their own. The same API portfolio now has to serve AI agents, and PZU expects that without a cloud gateway it will be hard to offer them anything attractive.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;API Management came to PZU as a facade for the new integration platform and process engine, so the company could leave its SOA platform smoothly and gather its APIs in one portfolio.&lt;/li&gt;&lt;li&gt;The self-hosted gateway let PZU use a cloud service while the data flow stayed on its side. That made the rollout easier in a regulated industry.&lt;/li&gt;&lt;li&gt;Shared gateway policies give teams security, observability and data records for accountability, so calling applications do not always need their own logic for them.&lt;/li&gt;&lt;li&gt;The integration team runs the platform and sets the standards, while developers publish APIs on their own and get wider permissions only after they show shared responsibility.&lt;/li&gt;&lt;li&gt;The team is building standards meant to catch departures from good practice while an API is still being built, with automation and AI agents in mind.&lt;/li&gt;&lt;li&gt;AI Gateway is available in the cloud gateway. Without it and without MCP support, agents find it hard to use PZU&apos;s APIs, so at the time of recording the company was speeding up the launch of that gateway.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;api-management-was-meant-to-help-pzu-leave-its-soa-platform&quot;&gt;API Management was meant to help PZU leave its SOA platform&lt;/h2&gt;
&lt;p&gt;Azure API Management was one element of the strategy PZU developed at the turn of 2021 and 2022 to transform its integration area. PZU did not adopt it with AI agents in mind. Before it, PZU built an internal hybrid iPaaS-type platform and deployed a new business process engine. One component was missing: a facade that gives access to these tools. At PZU it is called a “lightweight proxy.” The strategy was meant to allow a smooth exit from the SOA platform, and the combined requirements added up to a recipe for API Management.&lt;/p&gt;
&lt;p&gt;Today API Management sits at the entry point, and the integration platform runs between it and the backends. The SOA platform, from the era of products like BizTalk, used to be the main integration architecture. Its potential is slowly running out, though.&lt;/p&gt;
&lt;p&gt;Krzysztof Radzimowski, Product Delivery for the integration area at PZU: “We need something much more flexible. Something that lets us adapt ourselves to the business, not the other way around.” (translated)&lt;/p&gt;
&lt;p&gt;Another important goal was to attract the APIs that bypassed the integration platform at the time. Some of them still bypass it today. PZU wanted to bring these APIs and their development teams onto the platform to gain economies of scale and build a very attractive API portfolio. That was the main motivation, and it is still what matters most to PZU: the portfolio is meant to feed new solutions, automation and AI agents. Intelligent API orchestration needs as many interesting and useful APIs as possible.&lt;/p&gt;
&lt;h2 id=&quot;the-self-hosted-gateway-kept-the-data-flow-on-pzus-side&quot;&gt;The self-hosted gateway kept the data flow on PZU’s side&lt;/h2&gt;
&lt;p&gt;PZU operates in a regulated industry, and API Management is a cloud service. Under earlier regulatory requirements for the cloud, PZU had a number of doubts and concerns. The answer was the self-hosted gateway, a hybrid architecture: it is a cloud service, but the data flow stays on-premises, on PZU’s side. This made the platform rollout easier. PZU deployed API Management together with Protopia.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-pzu-hybrid-gateway-architecture-desktop.DQd19jlX.webp&quot; alt=&quot;Diagram of the hybrid API Management architecture at PZU. The API Management service runs in the Azure cloud and connects with a dashed line to the self-hosted gateway in PZU&apos;s infrastructure. PZU business units and partners and startups call APIs through the self-hosted gateway. The gateway routes traffic to the iPaaS-type integration platform and to the process engine, and the platform passes it on to PZU&apos;s backends. The whole data flow stays on-premises, on PZU&apos;s side.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the service in the cloud, the data flow on PZU&amp;#39;s side&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Today PZU is a little less afraid of the cloud. In the meantime it did a great deal of work to adapt to new regulatory requirements. API Management is no longer its only cloud solution, because PZU has moved into the cloud very boldly. With other solutions it has already met the regulatory recommendations, and the integration team can build on that. That is why PZU is preparing, with more peace of mind, to launch the cloud gateway, which it has not used so far.&lt;/p&gt;
&lt;h2 id=&quot;business-units-and-partners-treat-apis-as-the-natural-way-to-integrate&quot;&gt;Business units and partners treat APIs as the natural way to integrate&lt;/h2&gt;
&lt;p&gt;The integration team’s main customer is the business side, which represents the external customer: claims handling, policy sales and also the back office. The back office handles customer requests and has to reach data outside the company. The business has largely learned to use APIs and treats them as the natural way to communicate. For PZU’s partners and for startups that offer PZU their solutions, the API is also the entry point.&lt;/p&gt;
&lt;p&gt;PZU tends not to expose APIs that anyone can use. The B2B model prevails: an external company agrees on what information it will receive. An internal business partner comes with a need to work with some organization and contacts the integration team. The team makes that cooperation easier, and sometimes makes it possible in the first place.&lt;/p&gt;
&lt;p&gt;Traffic goes both ways. PZU exposes functions of its sales, policy or claims system to an external organization. It can also bring in external services, which often improve claims handling, for example, and use them in internal processes.&lt;/p&gt;
&lt;h2 id=&quot;an-insurance-comparison-site-gets-quotes-through-api-management&quot;&gt;An insurance comparison site gets quotes through API Management&lt;/h2&gt;
&lt;p&gt;One of the popular insurance comparison sites sends policy parameters through an API, receives a quote and compares it with other offers. PZU built this integration with the external customer in mind. In this channel it exposes the policy system’s API, and the whole process runs through API Management. The integration team is responsible for securing this API.&lt;/p&gt;
&lt;p&gt;The gateway also allows comparison tests: some requests can go to the API and some to another system, to compare how they behave. Rules in API Management can redirect a customer to the call center if the system does not respond fast enough, though not necessarily on the comparison site.&lt;/p&gt;
&lt;h2 id=&quot;gateway-policies-provide-accountability-for-government-registries&quot;&gt;Gateway policies provide accountability for government registries&lt;/h2&gt;
&lt;p&gt;The second example is data from government registries that PZU retrieves from outside. &lt;strong&gt;Accountability is&lt;/strong&gt; the ability to document, at the request of the institution that provides the data, who asked for it and when, and what data they received. The business that uses the registries has to ensure it. API Management has this data anyway, so it stores it in separate databases from which it can be retrieved later. The business does not have to collect it on its own.&lt;/p&gt;
&lt;p&gt;PZU still uses this pattern on the SOA platform today and carried it over to API Management. The response from a registry system that requires accountability can be stored selectively for chosen operations and, if needed, even for individual calls. This logic lives in the policies and in the API itself. Accountability is one of the regulator’s requirements that the integration area has dealt with for more than ten years.&lt;/p&gt;
&lt;p&gt;The API client does not always have to worry about this, because the shared platform does it. If callers integrated with an API directly, every calling application or every API would need its own logic to store the data.&lt;/p&gt;
&lt;h2 id=&quot;developers-get-security-and-observability-out-of-the-box&quot;&gt;Developers get security and observability out of the box&lt;/h2&gt;
&lt;p&gt;PZU expects a calling system to have a very simple way to authorize access and identify itself to the platform. Backends can have their own methods, such as a key or a certificate, depending on whether internal developers, external developers or a vendor designed them. API Management authenticates to each system in the right way. Callers connect to the gateway in one standardized way.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-pzu-shared-gateway-policies-desktop.CvBr5ADu.webp&quot; alt=&quot;Diagram of shared API Management policies at PZU. A calling application and an external partner connect to the gateway in one standardized way. In the gateway, shared policies provide security, observability and accountability. Observability feeds the observability tools with diagnostic data, and accountability stores in separate databases the data on who asked, when and for what. The gateway authenticates to the backends with their own methods: a key or a certificate.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: one entry point, shared policies, different backends&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The security and accountability layer is not enough to attract the developers of domain systems. So PZU also gives them observability built on the platform, out of the box. The developer only publishes an API or connects to one. The integration team takes care of the rest, including security in the built-in policies.&lt;/p&gt;
&lt;h2 id=&quot;developers-publish-apis-themselves-and-the-integration-team-runs-the-platform&quot;&gt;Developers publish APIs themselves, and the integration team runs the platform&lt;/h2&gt;
&lt;p&gt;At PZU, developers have access to the environment and configure their own APIs. The integration team has two roles. It is a Center of Excellence that provides know-how and a governance model. It is also the platform owner: it runs the platform, is responsible for its stability, and introduces the necessary level of standardization and permissions.&lt;/p&gt;
&lt;p&gt;A high degree of self-service was a goal of the rollout, and today it works. A developer who wants to publish an API defines it in the development environment and can set it up by clicking through the interface. Soon they will do it more simply, directly from their own code. Developers should work with API Management as if it were a module of their own application, without getting into the details. Helpers built on the Management API for API Management serve this purpose and are meant to let developers handle things on their own.&lt;/p&gt;
&lt;p&gt;A number of development teams already expose APIs and use the integration team’s code repository with growing confidence. Before anyone gets more permissions on the platform, they have to learn how to work with it and show that they share responsibility for it. Estimates put the number of APIs in the non-production environments at more than 50 today, probably closer to 100.&lt;/p&gt;
&lt;h2 id=&quot;the-integration-team-enforces-one-hidden-standard-and-organizes-apis-in-api-center&quot;&gt;The integration team enforces one hidden standard and organizes APIs in API Center&lt;/h2&gt;
&lt;p&gt;Probably the only standard in place that a developer may not know about is the set of policies that feed observability tools with the data needed for diagnostics. Developers are sometimes not used to following someone else’s guidelines, so the integration team enforces this standard. The standard is well documented, and developers know what to expect in API Management. PZU does not do classic hardening today, such as removing headers, unless someone asks for it. The APIs are used mainly internally.&lt;/p&gt;
&lt;p&gt;In parallel, the team is doing a lot of work on standards that are meant to catch departures from agreed good practices as early as the development stage. As a result, an API that reaches the non-production environments, and later production, should be much more mature. The drivers are automation, AI agents and AI in general.&lt;/p&gt;
&lt;p&gt;Krzysztof Radzimowski: “We need standards that make it possible to understand what an API does directly, without any extra tools, for example for AI mechanisms.” (translated)&lt;/p&gt;
&lt;p&gt;PZU is adopting Azure API Center slowly for now. The service works practically out of the box, but many of its features are still only on the roadmap. It automatically reads the data that is already in API Management and lets the team organize APIs that are not in API Management yet. It is a front door to what PZU already has in the API area. PZU sees how much Microsoft is changing API Management and API Center, and it is optimistic that API Center will mature fairly quickly.&lt;/p&gt;
&lt;h2 id=&quot;without-a-cloud-gateway-pzu-will-find-it-hard-to-serve-ai-agents&quot;&gt;Without a cloud gateway, PZU will find it hard to serve AI agents&lt;/h2&gt;
&lt;p&gt;PZU is investing heavily in the use of AI agent models and is adopting AI tools more and more boldly, just as it is doing with the cloud. The integration team wants to take part in projects that use AI components. Here API Management offers AI Gateway and support for the Model Context Protocol (MCP). Without them, agents, especially external ones, find it hard to communicate with PZU’s APIs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI Gateway is&lt;/strong&gt; the use of API Management in front of AI models, roughly speaking in front of ChatGPT or OpenAI. It shows who is calling, how, how many times and how many tokens they use. All requests and responses can go to internal logs. Another feature is semantic caching.&lt;/p&gt;
&lt;p&gt;At the time of recording, AI Gateway had probably been available for more than a year, but only in the cloud gateway, and PZU was not using it yet. That is why, at the time, PZU needed to move to the cloud gateway sooner. PZU expects that if it cannot offer a cloud gateway, it will find it very hard to give anything attractive to internal AI projects and to external agents that want to use its APIs.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-pzu-cloud-gateway-ai-agents-desktop.CaOzK9KY.webp&quot; alt=&quot;Diagram of the planned API Management cloud gateway at PZU. External agents and internal AI projects connect to the cloud gateway, whose launch PZU is preparing. AI Gateway and support for the MCP protocol run in the gateway. The gateway gives access to PZU&apos;s APIs, and AI Gateway sits in front of AI models such as OpenAI. AI Gateway shows who is calling, how, how many times and how many tokens they use, and requests and responses can go to internal logs.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: a cloud gateway between AI agents and PZU&amp;#39;s APIs&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;API exposure, accountability and access security, for example how many tokens a given process may use, were the topic of workshops planned for September. PZU also hoped to run an internal hackathon later that same year to connect API Management with AI models. It was meant to help people get familiar with what is available in the cloud.&lt;/p&gt;
&lt;p&gt;For PZU, showing the technologies is easy. The harder part is to describe a business need ambitious and broad enough to deliver with these tools. The needs the team runs into can often be met much more simply, without such sophisticated technology. PZU hopes that showing the possibilities will encourage the business to go a step further, and that a project using AI Gateway and Azure Logic Apps will come up in the near term. PZU uses Logic Apps as an orchestrator, and various conversations suggest that Logic Apps are also widely used to build AI-based agents and pipelines.&lt;/p&gt;
&lt;h2 id=&quot;pzu-has-built-a-portfolio-of-integration-technologies&quot;&gt;PZU has built a portfolio of integration technologies&lt;/h2&gt;
&lt;p&gt;As recently as about five years ago, PZU had two large bus-type integration platforms. To a large extent they forced architects to use them and to fit the business need to the platform. Today PZU has a portfolio of integration technologies that let it fit the integration area to the real needs of the business without excessive compromises. The integration team built this portfolio over several years and is looking for more components for it. In the team’s view, the work paid off.&lt;/p&gt;
&lt;p&gt;The portfolio includes a file import mechanism that PZU built together with Protopia. It is a dual-use technology: it handles file transfer and event processing, but it can also carry independent application logic, such as notifications that a file was uploaded.&lt;/p&gt;</content:encoded><dc:creator>Marek Grabarz</dc:creator><category>APIs and integrations</category></item><item><title>API Management as the gateway for AI agents into company systems</title><link>https://protopia.tech/en/publications/api-management-for-ai-agents-mcp/</link><guid isPermaLink="true">https://protopia.tech/en/publications/api-management-for-ai-agents-mcp/</guid><description>When a company plans many AI agents, each team may integrate separately with the same systems. See when a shared API Management gateway makes sense and where its role ends.</description><pubDate>Wed, 17 Dec 2025 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;You are planning AI agents that will use data in company systems and take actions in them. You must decide whether each team connects them to the systems separately or through a shared gateway. Azure API Management is that gateway: it puts the company&apos;s APIs in order, exposes them to agents as MCP servers and keeps security, network and access rules in one place. You will learn when the gateway is needed, what it does not do, how it works with on-premises systems and how to introduce it so the integration team does not become a bottleneck.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;With many agents, a shared API Management platform saves teams from repeating the same integrations.&lt;/li&gt;&lt;li&gt;APIs published in API Management can be exposed to agents as an MCP server, but long processes that keep state and can go back a step belong in an integration platform, because large orchestration in API Management is hard to maintain and debug.&lt;/li&gt;&lt;li&gt;A shared gateway lets you agree on traffic routes and one set of rules for all APIs in advance with the security, network and compliance teams, including APIs built by contractors.&lt;/li&gt;&lt;li&gt;With on-premises systems, a local gateway keeps sensitive data in the private data center, while configuration and monitoring stay in the cloud.&lt;/li&gt;&lt;li&gt;With APIOps, trained developers submit changes themselves through a pull request with automatic linting, and the integration team only reviews them.&lt;/li&gt;&lt;li&gt;Workshops with a list of processes quickly show that there will be many agents, and a shared platform may cost a bit more at the start.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;with-many-agents-api-management-becomes-essential&quot;&gt;With many agents, API Management becomes essential&lt;/h2&gt;
&lt;p&gt;API Management becomes essential when a company builds many agents or a whole agent factory for the organization. For an agent to reach the company’s information sources and automate further actions, it has to talk to internal systems, and the preferred form of that communication is a standardized API.&lt;/p&gt;
&lt;p&gt;A single agent does not need an API if the data it needs can be extracted from the systems and delivered in static form. The API becomes critical when the underlying data changes dynamically and the agent’s actions are dynamic too.&lt;/p&gt;
&lt;p&gt;With many agents and no shared platform, each group of agent builders may do the same tedious work. In e-commerce, for example, each agent would integrate separately with the product data API. API Management delivers the same APIs to agents in a standardized way:&lt;/p&gt;
&lt;p&gt;Marek Grabarz: “The API systematizes this, and API Management lets us reuse that API in different places.” (translated)&lt;/p&gt;
&lt;h2 id=&quot;azure-api-management-has-three-parts&quot;&gt;Azure API Management has three parts&lt;/h2&gt;
&lt;p&gt;Put very simply, Azure API Management has three main components: the Control Plane, the Developer Portal and the API endpoint itself.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Control Plane is&lt;/strong&gt; the management portal. The APIs already exist in the company, so you import and configure them here, and you describe the rules with policies. A policy can, for example, admit only members of an Active Directory group, define security and monitoring, or set rate limiting: the limit on calls a given system can make per minute or per second.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Developer Portal is&lt;/strong&gt; the visual version of what the Control Plane defines. A developer or a team member who integrates with an API does discovery here: they find APIs and see their descriptions, endpoints, business value, how to authenticate and call them, and the allowed number of calls.&lt;/p&gt;
&lt;p&gt;The third part is the API endpoint itself, exposed to systems. Agents, mobile apps, internal line-of-business apps, portals and even desktop apps use it.&lt;/p&gt;
&lt;h2 id=&quot;existing-apis-reach-agents-as-mcp-servers&quot;&gt;Existing APIs reach agents as MCP servers&lt;/h2&gt;
&lt;p&gt;API Management can expose the same API as REST, gRPC or as an MCP server that agents understand. For an API that already runs as REST or in another form, API Management now has, put simply, one checkbox: “Also expose as MCP.” API Management does the translation itself. To a large extent, this lets you connect old systems to agents without modifying them.&lt;/p&gt;
&lt;p&gt;A company that already exposed its APIs well for applications can use the same APIs for agents. For example, a wholesaler’s agent lists the products in a chosen category, adds a product to the cart and places the order through APIs that are already exposed.&lt;/p&gt;
&lt;p&gt;This works only if the operations have descriptions. Names like GetCustomer or GetProduct sound clear, but for more complex business operations the endpoint name alone does not say what the endpoint does. That is why each operation gets a description in API Management, so the agent knows which endpoint to use and how. In Protopia projects, the team tries to have an LLM write these descriptions. The model gets an instruction to describe the endpoints in a way that it understands itself, which means in a way the agent understands.&lt;/p&gt;
&lt;h2 id=&quot;multi-step-processes-need-an-integration-platform&quot;&gt;Multi-step processes need an integration platform&lt;/h2&gt;
&lt;p&gt;API Management exists to expose APIs, so processes with many steps and business logic need an extra layer underneath. &lt;strong&gt;Integration Platform as a Service (iPaaS) is&lt;/strong&gt; the layer of flows, workflows and orchestration between API Management and the company’s systems. API Management does not always reach deep into SAP or a database, and it does not send emails, notifications or SMS messages itself unless the company has an API for that.&lt;/p&gt;
&lt;p&gt;For example, an agent calls the product order API exposed in API Management. Instead of calling the system’s API right away, API Management starts a workflow that sends the customer an email about the order or about a change in its status. An integration with 5 or 10 steps, where the process can go back a step and holds state, has to store that state somewhere, restore it and behave predictably.&lt;/p&gt;
&lt;p&gt;API Management offers simple orchestration. You can build large orchestration too, but it is hard to maintain and debug, and it is hard to find what fails in it. It is also not stateful.&lt;/p&gt;
&lt;p&gt;Marek Grabarz: “I always advise clients not to build large orchestration in API Management.” (translated)&lt;/p&gt;
&lt;p&gt;Simple orchestration in API Management helps when there is no other option, for example when you need to add SMS messaging to an old system without modifying it.&lt;/p&gt;
&lt;p&gt;If an old system, SAP for example, exposes an API, that API goes straight into API Management. If the system’s only interface is a SELECT query against its database, you need an intermediate layer that exposes an API in the form of a workflow. First check whether someone has already exposed this data through an HTTP service: if a SELECT query against the database exists, someone may have had this need before.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-apim-ipaas-orchestration-desktop.RHuaUFgC.webp&quot; alt=&quot;An agent calls the product order API in API Management. If a system, SAP for example, exposes an API, API Management passes the call straight to it. A process with many steps goes to a workflow in the integration platform (iPaaS), which stores the state, sends the customer an email about the order and reaches the database when a select is its only interface.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: API Management exposes the APIs, and iPaaS handles stateful processes&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;one-gateway-organizes-traffic-and-security-rules&quot;&gt;One gateway organizes traffic and security rules&lt;/h2&gt;
&lt;p&gt;When systems connect through API Management, there is one gateway instead of point-to-point connections between every system. In large companies, it is a shared security and network facade: APIs are not exposed to the public internet piecemeal. Agents and applications get one consistent interface, even though the systems underneath use different protocols, SOAP for example, and different authentication methods. You can strengthen authentication at the gateway and leave older systems alone. A developer who wants to integrate has to identify themselves and describe the integration, so you can see who uses which system, also inside the company.&lt;/p&gt;
&lt;p&gt;The speakers see a recurring pattern at clients: the companies they work with, often in regulated industries, have fairly tight security and network rules. If the systems are in one place and the APIs in another, it is hard to agree with the security, network and compliance teams on who can talk to whom. A shared gateway lets you set up network connectivity, approvals and security contracts in advance for two defined routes: from agents to API Management and from API Management to the individual APIs.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-apim-single-gateway-desktop.CxmYWrUA.webp&quot; alt=&quot;On the left, agents, mobile apps and portals connect separately to each system: the product API, a SOAP system and a purchased SaaS. On the right, all of them connect through API Management, which holds the policies and monitoring. Traffic then runs on two defined routes: from the consumers to the gateway and from the gateway to the individual APIs.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: point-to-point connections versus one API Management gateway&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;In a sense, this is allowlisting of internal APIs: you know who can consume which data. Compliance, usage accounting, visibility, monitoring and the security approval of how an API is exposed sit in one place.&lt;/p&gt;
&lt;p&gt;Shared security rules at the gateway are another reason not to integrate systems directly one-to-one. The speakers see APIs as the main place where someone misuses the company’s systems, for example to rack up points or to try to get access.&lt;/p&gt;
&lt;p&gt;The reference point is the OWASP API Security Top 10. Microsoft’s documentation makes it easy to find how API Management and other tools, for example Microsoft Defender for APIs or DDoS protection, address the individual threats on this list.&lt;/p&gt;
&lt;p&gt;You cannot fully trust even an internal client. In a company with, say, 10 departments, different groups build APIs: employees, external contractors and vendors that built and delivered something. On top of that comes purchased SaaS that works as a black box and cannot be modified. Each of these groups has different experience and a different approach to security, so each can make a mistake. Since everyone integrates through API Management, the platform team and the security team can set shared rules for all exposed APIs and in this way patch potential holes.&lt;/p&gt;
&lt;h2 id=&quot;partners-get-a-chosen-subset-of-apis-through-products&quot;&gt;Partners get a chosen subset of APIs through products&lt;/h2&gt;
&lt;p&gt;You can expose some APIs externally, and the products mechanism defines which subset a given consumer sees. The consumers are business partners, other companies or startups. They meet the agreed contracts, approvals and access conditions and consume only the data the company deliberately exposes to them. This way, the company can open new communication and sales channels, perhaps also new products.&lt;/p&gt;
&lt;p&gt;This is part of the API Economy concept, which has been around for roughly 10 years. &lt;strong&gt;API Economy is&lt;/strong&gt; the management term for the value that a company’s APIs provide. It is mainly about internal monetization: the company builds further systems and add-ons on the exposed data and creates value inside the company and outside it. It is not about selling APIs as SaaS with a pay-per-use fee.&lt;/p&gt;
&lt;p&gt;A product in API Management groups APIs, for example as an integration with a certain type of system. The name sounds like SaaS, but you do not buy a product: you assign it to a consumer, a company or a project. One API can belong to many products. Through products, consumers authenticate and get authorized in a standardized way.&lt;/p&gt;
&lt;h2 id=&quot;a-self-hosted-gateway-keeps-data-in-the-private-data-center&quot;&gt;A self-hosted gateway keeps data in the private data center&lt;/h2&gt;
&lt;p&gt;A self-hosted gateway lets you serve on-premises systems without sending traffic to the cloud when the application and the API are in a private data center. In Protopia projects, roughly half of API Management deployments run in a hybrid model. In this model, the Developer Portal and the Control Plane, where you configure APIs, stay in the cloud. The gateway that exposes the APIs can run in two places.&lt;/p&gt;

&lt;div class=&quot;table-scroll&quot; role=&quot;region&quot; tabindex=&quot;0&quot; aria-label=&quot;Two places for the API Management gateway&quot;&gt;&lt;table&gt;&lt;caption&gt;Two places for the API Management gateway&lt;/caption&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Gateway&lt;/th&gt;
&lt;th&gt;Where it runs&lt;/th&gt;
&lt;th&gt;Call when the application and the API are in a private data center&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cloud gateway&lt;/td&gt;
&lt;td&gt;In the cloud&lt;/td&gt;
&lt;td&gt;Application → cloud → API → cloud → application&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-hosted gateway&lt;/td&gt;
&lt;td&gt;A container with an image from Microsoft, deployed by the client, usually on its own infrastructure: on-premises or in another cloud&lt;/td&gt;
&lt;td&gt;Traffic stays local&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The local gateway shortens the call path, which helps with latency and cuts the cost of data transfer to the cloud and back. The third gain is compliance: banking, medical or insurance data does not leave the private data center in any way, and the cloud holds only monitoring and possibly observability.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-apim-hybrid-self-hosted-desktop.BGOnhkWv.webp&quot; alt=&quot;In the cloud run the Control Plane, where you configure APIs, the Developer Portal and monitoring. In the private data center, the client runs a self-hosted gateway as a container with an image from Microsoft. The application calls the API through the local gateway, so banking, medical or insurance data does not leave the data center, and only monitoring and observability go to the cloud.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: hybrid model with a self-hosted gateway&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;A gateway hosted locally on Kubernetes or on a virtual machine can also expose APIs externally, through the company’s own endpoint and firewall, under its internal rules. The APIs are then not exposed from a cloud region, such as West Europe in Amsterdam. So you can expose on-premises systems externally and at the same time bring them into the agent platform.&lt;/p&gt;
&lt;h2 id=&quot;api-center-and-linting-put-apis-in-order-before-exposure&quot;&gt;API Center and linting put APIs in order before exposure&lt;/h2&gt;
&lt;p&gt;The company finally knows which APIs it has, and linting in CI/CD lets only APIs described according to the standard onto the gateway. In the past, API Management deployments usually started from a pent-up need: the company had a large number of APIs built over the years and wanted to systematize them and expose them centrally. Often an integration team that wants to simplify its work is behind this, or people in a CTO role who see API Economy as central to the company’s future success.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Azure API Center is&lt;/strong&gt; an API catalog that you can call a CMDB for APIs. Protopia often deploys it alongside API Management. APIs in the catalog do not have to be exposed through API Management: they can run directly or elsewhere. API Center holds a full description of each API, including machine-readable ones such as Swagger or OpenAPI, and replaces the outdated wiki pages that describe APIs.&lt;/p&gt;
&lt;p&gt;API linting works like code analysis that checks security, naming, structure and vulnerabilities.&lt;/p&gt;
&lt;p&gt;Many APIs already run today, so new standards create tension: the departments that expose APIs have to bend to them at some point. Protopia mainly deploys the platform, the processes and linting. The integration teams and the groups that expose the APIs are responsible for adapting the APIs to the standards.&lt;/p&gt;
&lt;h2 id=&quot;apiops-moves-api-configuration-into-a-repository&quot;&gt;APIOps moves API configuration into a repository&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;APIOps is&lt;/strong&gt; an approach from the Microsoft API Management world, similar to GitOps in Kubernetes: policies, integrations and entire APIs are configured in a code repository, and the team only promotes this configuration to the next environments. Protopia has deployed it several times already.&lt;/p&gt;
&lt;p&gt;An API and its revisions, versions or changes move from the development environment through integration, UAT and pre-production to production. There can be as many environments as the company needs. Reconfiguring an API by hand in each of them means more work and a risk of human error: someone clicks the wrong thing, moves something incorrectly or misses a setting in a policy or an orchestration. Git, on the other hand, keeps a change history: who changed what, when, where and why.&lt;/p&gt;
&lt;p&gt;Without APIOps, an integration team of, say, 5–10 people would at some point become a bottleneck, because the whole company would come to it asking to integrate their APIs. One of the speakers once worked at companies where exposing an API required signing up for an architecture review board meeting, with a date 2 weeks out. The board returned a long list of fixes, and the whole process took 2–4 months, even though it was only a simple iteration.&lt;/p&gt;
&lt;p&gt;With APIOps, developers create a branch themselves, code the integration and submit a pull request. Someone from the integration team reviews the changes and comments on what to fix. The pull request gives approval to deploy to the next environments and triggers automatic steps, such as linting for compliance with the standards. Developers need basic knowledge for this, so introducing APIOps comes with training in the company and awareness building. Knowledge spreads out, and the central team reviews changes instead of being a bottleneck.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/4-apiops-branch-to-production-desktop.CuMOJ4IP.webp&quot; alt=&quot;A developer creates a branch, codes the integration and submits a pull request. Automatic linting checks compliance with the standards, and someone from the integration team reviews the changes and comments on fixes. The pull request gives approval to deploy, and the configuration from the repository moves from the development environment through integration, UAT and preproduction to production.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: an API change in APIOps from branch to production&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-rollout-starts-with-workshops-and-a-list-of-processes&quot;&gt;The rollout starts with workshops and a list of processes&lt;/h2&gt;
&lt;p&gt;Before Protopia starts technical work, it usually runs workshops with the client. It brings together business architects and the people responsible for areas of the company and finds out from them which processes agents could improve. From such a long list, clients quickly see that they will end up with 10 or 30 agents or more, not one, two or three.&lt;/p&gt;
&lt;p&gt;With a real platform in the organization, you pay the cost of API integration anyway. Protopia usually builds the platform so that it can be reused. At the start, this can cost a bit more: you have to build the foundations, connect the elements and pass internal certification in the security or procurement department.&lt;/p&gt;
&lt;p&gt;Many of Protopia’s API Management deployments were built before anyone talked about LLMs, and putting the APIs in order was then a goal in itself for the organization.&lt;/p&gt;</content:encoded><dc:creator>Szymon Warda, Marek Grabarz</dc:creator><category>APIs and integrations</category></item><item><title>The AI agent factory: how to start and reach production</title><link>https://protopia.tech/en/publications/ai-agent-fundamentals/</link><guid isPermaLink="true">https://protopia.tech/en/publications/ai-agent-fundamentals/</guid><description>Once the first AI agents are written in Python, the next question is the platform: a ready-made service or your own, how to connect company systems, and how to move an agent to production safely.</description><pubDate>Thu, 27 Nov 2025 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;TL;DR&lt;/h2&gt;&lt;p&gt;You want to roll out AI agents in your company, and you must decide what to build them on and how to give them safe access to data and systems. An agent factory consists of a ready-made platform (if you already have a cloud, your provider&apos;s service), one shared gateway to company systems and a closed test space from which an agent goes to production like any other application. You start with a simple case where you know what the agent receives and what it must return.&lt;/p&gt;&lt;h2&gt;Key takeaways&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;An AI agent is the same LLM-based process as a chatbot, except that it works longer on data and makes its own decisions.&lt;/li&gt;&lt;li&gt;At the start, one orchestrator agent with small, specialized agents usually handles a business process, because a long context is expensive and the agent then follows instructions less well.&lt;/li&gt;&lt;li&gt;Integrations exposed through one gateway serve every agent you build later and cost less to maintain.&lt;/li&gt;&lt;li&gt;When you already have a cloud, you use your provider&apos;s ready-made service, and you can later swap the model and the platform without rebuilding the integrations.&lt;/li&gt;&lt;li&gt;The sandbox enforces a secure configuration. From there, an agent kept as code in a repository goes through CI/CD, first to a non-production environment and, after tests, to production.&lt;/li&gt;&lt;li&gt;The first stage is usually the hardest and the longest: choosing the case and working out which systems the agent must integrate with.&lt;/li&gt;&lt;/ul&gt;&lt;h2 id=&quot;an-ai-agent-decides-on-its-own-within-the-limits-you-set&quot;&gt;An AI agent decides on its own, within the limits you set&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;An AI agent is a process that uses a large language model (LLM), works longer on data and makes its own decisions within the limits you set.&lt;/strong&gt; Some definitions describe an agent as an independent, living entity. In practice, it is the same chatbot or language model plugged into a process, for example in an application, just in a different form.&lt;/p&gt;
&lt;p&gt;A runtime framework runs the agent. It is either a ready-made cloud service or code in Python or another language with a framework that connects to an LLM, for example one from OpenAI. Under the hood, an agent is a loop that invokes itself based on the input data.&lt;/p&gt;
&lt;h2 id=&quot;agents-at-protopias-clients-handle-documents-reports-and-processes-with-no-fixed-path&quot;&gt;Agents at Protopia’s clients handle documents, reports and processes with no fixed path&lt;/h2&gt;
&lt;p&gt;Document processing keeps coming up with Protopia’s clients. The agent receives a document after OCR, for example from an email in a mailbox or from a system where an employee uploaded it. It recognizes the PESEL (the Polish personal ID number), the customer number and other identifying data, tries to gather information from various company systems and supports the employee. A typical document is a complaint, probably the most common case of this kind. An example from one client: someone disputes an invoice, the agent finds that the discount was calculated incorrectly and prepares a draft email for the sales rep that explains it.&lt;/p&gt;
&lt;p&gt;The second case is reports built from data spread across systems. A few times, clients needed a merged view of a customer across several systems (CRM, ERP, help desk or another domain-specific system), to stop jumping between them or to find inconsistencies in the data.&lt;/p&gt;
&lt;p&gt;The third case is processes whose path is hard to set in advance. The agent can then, depending on the situation, collect data or take actions in systems. If something is missing, it fills in the data in several places.&lt;/p&gt;
&lt;h2 id=&quot;at-the-start-one-orchestrator-and-small-agents-usually-handle-a-business-process&quot;&gt;At the start, one orchestrator and small agents usually handle a business process&lt;/h2&gt;
&lt;p&gt;The starting setup is one orchestrator agent per business process, with small agents attached to it that take actions in individual systems. Technically, this scenario is called connected agents. The orchestrator acts like a conductor or a team leader: it distributes the work and combines the results. Other scenarios exist too, but they are not recommended at the start, because they can be a nightmare to test and the results can be very disappointing.&lt;/p&gt;
&lt;p&gt;Clients who build agents on their own first try to make a super agent and connect everything to it. The opposite trend resembles microservices: a very large number of small agents.&lt;/p&gt;
&lt;p&gt;The reason is the context window. &lt;strong&gt;A context window is the limit on how much text an LLM accepts as input.&lt;/strong&gt; In automation, a long context turns out to be expensive, and the agent does not follow instructions. That is why the context is kept to a minimum and tasks are split among smaller, specialized agents.&lt;/p&gt;
&lt;p&gt;Besides the budget, the technology itself is a limit for now. An agent does not always follow its instructions. Anthropic says its goal for the coming months and years is for its language model to follow the instructions it gets and the execution order exactly.&lt;/p&gt;
&lt;h2 id=&quot;one-integration-gateway-exposes-systems-to-all-agents&quot;&gt;One integration gateway exposes systems to all agents&lt;/h2&gt;
&lt;p&gt;A system connected once to the integration gateway is then available to every agent the company builds. Protopia recommends Azure API Management to clients as the central gateway that exposes various systems to agents in a uniform way. The language model and the runtime code are universal, and the company must standardize how it integrates with its own systems.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/1-ai-agents-integration-gateway-desktop.C2Ow7OWK.webp&quot; alt=&quot;An orchestrator agent handles a business process and distributes work among small agents. The small agents call company systems through one integration gateway, Azure API Management, which holds the call descriptions. Behind the gateway are CRM, ERP, help desk and an HR system. A next agent uses the same integrations, connected once.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: agents reach company systems through one gateway&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Łukasz Kałużny: “Because the heart of it, contrary to appearances, is not the agent, but the way you deliver knowledge to it and the ability to take those actions through integration with other systems.” (translated)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An agent’s integration with a system is a natural-language description of the calls, optimized for the LLM.&lt;/strong&gt; You write the description the way you would explain the call to a person: what data the method reads or writes and what you must pass to it. Example: this call returns customer data by PESEL number or by customer ID.&lt;/p&gt;
&lt;p&gt;A team that writes agent code without a ready-made platform can try to build the integration itself. At a larger scale, source systems may need changes before they can be connected. A larger organization also wants to reuse an integration it built once in other places, as cheaply as possible. It is possible without a gateway, but in Protopia’s experience it is more expensive and slower.&lt;/p&gt;
&lt;p&gt;The gateway also holds the central documentation of the descriptions, for example: in the HR system, employee data is here and their leave is there. Other people who build and test agents later use these descriptions. The gateway makes integrations cheaper to maintain later.&lt;/p&gt;
&lt;h2 id=&quot;the-agent-factory-works-like-platform-engineering-for-agents&quot;&gt;The agent factory works like platform engineering for agents&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;An agent factory is a platform that lets you move quickly from testing an agent to running it in production.&lt;/strong&gt; You can look at it as platform engineering: the path goes through a proof of concept (PoC) and a proof of value and, if everything is fine, ends with the agent released for use.&lt;/p&gt;
&lt;p&gt;At a small scale, a data scientist or a developer writes an agent in Python with frameworks from Microsoft, OpenAI or Google. That is how a PoC starts and the first agents appear. Protopia’s clients, often companies in finance, in regulated industries or larger organizations, want to bring order to this process.&lt;/p&gt;
&lt;p&gt;Tests usually run on data as close as possible to the data in production systems, to see the agent’s value.&lt;/p&gt;
&lt;h2 id=&quot;with-a-cloud-in-place-choose-your-providers-ready-made-paas-service&quot;&gt;With a cloud in place, choose your provider’s ready-made PaaS service&lt;/h2&gt;
&lt;p&gt;If your organization already has a cloud in place, your provider’s ready-made platform as a service (PaaS) is usually the best and most cost-effective choice, especially when you want to show value and start testing. Every hyperscaler has an equivalent of Azure AI Foundry. From the point of view of Protopia, which works with Azure every day, AI Foundry is “good enough”.&lt;/p&gt;

&lt;div class=&quot;table-scroll&quot; role=&quot;region&quot; tabindex=&quot;0&quot; aria-label=&quot;Platform options for agents&quot;&gt;&lt;table&gt;&lt;caption&gt;Platform options for agents&lt;/caption&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What it means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A cloud provider’s PaaS service, e.g. Azure AI Foundry&lt;/td&gt;
&lt;td&gt;For organizations with a cloud in place. It has some limits, but you can use it right away once the platform is rolled out.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A ready-made open-source agent platform, e.g. Dify&lt;/td&gt;
&lt;td&gt;An alternative when you have no cloud. You still have to supply the language model somehow.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Your own platform built on open source&lt;/td&gt;
&lt;td&gt;Probably several months before it fully works. That time does not go into building agents.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;A ready-made service limits the amount of code and the number of frameworks. The team focuses on two things: designing agents, and connecting systems and data sources along with testing. The work shifts from writing code to writing prompts. Do not bury yourself in open-source frameworks and gorgeous diagrams from LinkedIn.&lt;/p&gt;
&lt;p&gt;There is no major vendor lock-in when you move away from Azure AI Foundry. You can swap the language model. System prompts, the agent’s instructions, need a check and an adjustment after a model change, because every model has its nuances, but this is more like polishing. Every framework, open source included, consumes the integrations in Azure API Management in exactly the same way. The exit strategy can therefore be described as copy-paste plus tests.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/2-ai-foundry-exit-strategy-desktop.DAG9uubF.webp&quot; alt=&quot;After leaving Azure AI Foundry for another framework, open source included, you can swap the language model. You check the system prompts and adjust them to the new model, which is more like polishing. Every framework consumes the integrations exposed in Azure API Management the same way. The exit strategy is copy-paste plus tests.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: what changes when you leave Azure AI Foundry&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;ai-sandbox-enforces-a-secure-configuration-on-every-experiment&quot;&gt;AI Sandbox enforces a secure configuration on every experiment&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;An AI Sandbox is a prepared space, for example in Azure, where the team experiments with agents safely.&lt;/strong&gt; That is Protopia’s name for this concept in its work with clients. The preparation covers security controls, the platform configuration, access management and control of inbound and outbound network traffic.&lt;/p&gt;
&lt;p&gt;The cloud platform’s mechanisms let you enforce a configuration in which neither a developer nor a data scientist can click an “expose to the internet” button.&lt;/p&gt;
&lt;p&gt;The sandbox also limits configuration changes, so the team can focus on testing the agent. At one client, the team spent time working out how to connect to an AI service from an open-source library and sign in securely.&lt;/p&gt;
&lt;p&gt;Łukasz Kałużny: “So the sandbox has two goals. One is security and giving room to experiment, but on the other hand also forcing a focus on business value, not on playing with technology.” (translated)&lt;/p&gt;
&lt;h2 id=&quot;an-agent-goes-to-production-like-any-other-application&quot;&gt;An agent goes to production like any other application&lt;/h2&gt;
&lt;p&gt;An agent moves from the sandbox to production just like another application on Kubernetes. When it rolls out the platform, Protopia prepares a set of deployment scripts for the whole process: the agent goes from the sandbox to a non-production environment and, after tests, to production. The platform is built in this order: first a secure AI Sandbox from PaaS or ready-made services, then the process of moving to the environments where the agent runs.&lt;/p&gt;
&lt;p&gt;The AI Sandbox is a shared space, divided much like namespaces in Kubernetes. The same development practices apply in it as for regular applications. The agent lives as code in a repository and is deployed through CI/CD, with no magic and no special clicking. The sandbox has the same configuration that the agent later gets in production.&lt;/p&gt;
&lt;figure class=&quot;diagram&quot;&gt;&lt;picture&gt;&lt;img src=&quot;https://protopia.tech/_astro/3-ai-agent-path-to-production-desktop.C_JRhJQl.webp&quot; alt=&quot;The agent is built in the AI Sandbox. Its code goes to a repository as agent as code, and CI/CD deploys it to a non-production environment. After tests, the agent goes to production. All environments have the same configuration, as with a regular application.&quot; width=&quot;736&quot; height=&quot;414&quot; class=&quot;astro-zd3urm2l&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/picture&gt;&lt;figcaption class=&quot;text-small text-secondary&quot;&gt;Diagram: the agent&amp;#39;s path from AI Sandbox to production&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;the-first-stage-is-a-simple-case-with-a-defined-input-and-output&quot;&gt;The first stage is a simple case with a defined input and output&lt;/h2&gt;
&lt;p&gt;To start, choose simple cases for testing agents. Next, define what goes into the agent, for example a piece of information, a request or a document, and what the result must be. Then work out which systems you must integrate with the agent. Only with this knowledge do you move on to implementation. This stage is usually the hardest and the longest.&lt;/p&gt;</content:encoded><dc:creator>Łukasz Kałużny, Mikołaj Szczerbicki</dc:creator><category>AI and agents</category></item></channel></rss>