Rolling out AI in the SDLC: process standard, measurement and pilot

For:
CTOs, CIOs, IT managers and development team leads who plan to roll out AI in the software development life cycle
Reading time:
10 min
Episode:
28 min

This episode is in Polish. Full Polish version with transcript

After you click, the video loads from YouTube (Google). Google may store data on your device and process it in the USA as well. More (PDF, in Polish) · Watch on YouTube

In brief

Your teams already use AI tools and you see no effect beyond the cost, or you are deciding which processes get AI first and how to measure success. The gain from AI shows only after you redesign the software development life cycle: a specification before code, a shared standard for working with the agent and automatic feedback, all tested in a pilot on a real application. Measure the current state before the pilot, because only the comparison with it shows whether to scale or to fix the foundations first.

Key takeaways

  • Speeding up only one stage clogs the process somewhere else, so you standardize the whole development cycle: the specification before code, the work with the agent, skills, documentation in the repository and an automatic feedback loop, and you choose the tools last.
  • The rollout starts with an assessment of process readiness: CI/CD, release cadence, testability and deployment automation.
  • The pilot works best with a team that volunteered, on an application the team really develops and maintains.
  • The measurement before the pilot is a diagnostic with no set targets: you measure delivery, flow and quality, not tokens, lines of code or the share of AI-generated code.
  • The pilot can reveal a bottleneck outside AI, such as a queue at acceptance, and that is also a result: the process gets measured and improved.
  • Scaling does not have to cover the whole organization: depending on company size and criticality, part of it or a few development teams are sometimes enough.

Handing out AI tools alone brings no visible speedup

Handing teams tools such as GitHub Copilot, OpenAI Codex or Anthropic’s Claude Code does not by itself speed up the work in any real way, and the only visible effect is the cost. The speakers hear this in the market and when they exchange experience with peers.

One of the main topics clients raise is AI wired into processes that are neither measured nor standardized. In conversations with the management of several organizations, two situations come up. In the first, the company is still working out which processes get AI first and how to measure success. In the second, developers start using AI on their own and the manager agrees. Paid, more expensive and supposedly more secure versions of chat apps then appear, next to custom interfaces wired to the API. None of it is standardized, and at some point the organization sees neither an effect nor a gain.

Speeding up one stage clogs the process elsewhere

The SDLC (software development life cycle) is the whole path from gathering a business need to working, deployed software.

Reports from companies that focused on writing code alone speak of gains of around 4–5%, sometimes less, depending on whom you ask. In some very small organizations you hear of much higher numbers, because they have no organizational overhead. In a small product, such as a small SaaS, a developer can write a change, test it and deploy it within a day. In the large organizations Protopia works with, that is very often not possible.

Łukasz Kałużny: “if we speed up just one stage, then realistically it will get clogged somewhere” (translated)

That is why, when you roll out AI in the SDLC, you standardize the whole process.

The gain from AI comes only after a process redesign

In 1987 the Nobel laureate Solow observed that you could see computers everywhere, but nowhere could you see a rise in value from them. Organizations waited about 15 years for a return on their investment in computers. With the electric motor in factories it took about 40 years, because processes had to adapt to the new technology.

AI follows the same path: the gain will show only after the process and the approach are redesigned and the organization matures. The payback cycle will probably be much faster than with computers or the electric motor.

If you want to control this return, you have to measure it, that is, assess the state before and after you improve the process. You plan what to measure and at which points before the rollout starts, whether it runs all at once or in stages.

The rollout starts from the current process and evolves it gradually: it adapts what exists to working with AI tools. The goal is a repeatable, scalable effect for the whole organization. The success of one team may sometimes be enough, but then you have to check whether you can repeat it later.

The 2025 DevOps Research and Assessment (DORA) report describes AI as an amplifier that lifts mature teams and intensifies the dysfunctions of weak ones.

The standard covers the specification, agent work and the feedback loop

Spec Driven Development is a method in which the team gathers requirements in a repeatable way, generates a specification from them, then a plan for the agent’s work on a given feature, and only at the end the code. It is the opposite of vibe coding. It builds on open standards such as GitHub Spec Kit or OpenSpec, or on a lighter version prepared for a specific organization.

The goal is one consistent standard for working with the agent across the organization’s repositories. The team gets a base it can reuse in a new project or in a project that the standard covers.

Skills are sets of commands and prompts that tell the agent how to do something. They can be company-wide, for example “review the requirements against the global policy”, or project-specific. One of the popular skills that Protopia rolls out at clients analyzes a bug report from production to locate the bug faster.

Documentation goes into the code repository, stays up to date and is readable by both the agent and a human.

Harness and guardrails are the whole environment around the agent that gives it a short feedback loop. Just as a developer uses unit tests, the agent gets automatic formatting, linting and testing, blocking checks and quality gates. In some places the gate is test coverage or cognitive complexity (cognitive load), a measure of how hard the code is to understand and how complex it is. You choose the metrics so that the code is easier to remove or maintain in the future, and you wire them in automatically. This way the agent gets feedback as fast as possible and works in the direction you set.

You choose the tools last, because many of these practices carry over between models and tools. Tools currently change every quarter, and some say the trends change even every month. That is why you choose them together with the models to fit your tech stack and your compliance requirements.

The standard covers the whole SDLC. Spec Driven Development handles requirements. The agent’s tools, together with Spec Driven Development, handle code, tests and review. Maintenance relies on integration and unit tests that run locally and in CI/CD.

Layer diagram. The top row shows three stages of the software development life cycle: requirements, code with tests and review, and maintenance. Spec Driven Development sits under requirements, agent tools together with Spec Driven Development sit under code, tests and review, and integration and unit tests that run locally and in CI/CD sit under maintenance. Under all stages lies a shared standard for working with the agent: skills, documentation in the repository, and harness and guardrails that give the agent a short feedback loop.
Diagram: the standard for working with the agent covers the whole development cycle, not only writing code

The rollout starts with a readiness assessment and the choice of a pilot team

The first step is an assessment of the current state and of readiness. It checks whether the process is ready for AI and assumes from the start that it will need adjusting. It shows what your SDLC standard looks like and whether it has bottlenecks already on paper: how CI/CD and the build servers work, what the release cadence is and whether the company has standardized it. It also covers testability hygiene and deployment frequency. If frequent deployments are not needed, it checks whether deployment is at least automated. The assessment draws on the DevOps approach and its feedback loop.

The second step is choosing the pilot: one or two teams, systems, modules or applications. Teams that volunteer are the best choice, because the pilot needs the commitment of the whole team. The pilot team should be active in testing and later carry the solution further in the organization as its evangelist.

Flow diagram in five steps. The rollout starts with an assessment of process readiness: CI/CD, release cadence and testability. Next comes the pilot choice, ideally teams that volunteer, and measuring the state before the pilot as a diagnostic without KPIs. Then follows the pilot on a real application, and at the end the comparison with the state before. A dashed arrow labeled baseline links the measurement of the state before directly to the comparison.
Diagram: the steps of rolling out AI in the SDLC, from the readiness assessment to the comparison with the baseline

Measuring the state before the pilot is a diagnostic with no set targets

The third step is measuring the state before the pilot, without KPIs. KPIs bring the risk of window dressing and gaming the metric. The measurement is a diagnostic: it shows the bottleneck and the place to start, so you speed up wisely.

You do not measure vanity metrics: tokens used, lines of code generated, the number of pull requests or the share of code generated by AI.

Łukasz Kałużny: “This is not productivity, and we will not reduce it all to one number; it will be a multidimensional assessment.” (translated)

The measurement has 3 levels, or 2.5, depending on how you count. The goal of the whole measurement is faster or more stable delivery.

Three levels of measurement before and after the pilot
Level What it shows How you measure it
Delivery Whether you deliver faster and with more stability 4 DORA metrics, among them how often you deploy, the share of failed deployments and the time to restore after a failure
Flow The time of the whole process From the arrival of a request for a feature or a change, such as a new screen, to its release in production
Quality Whether the agentic approach made quality worse, made it better or kept it stable (in a mature team) You set the method with Protopia for your case, for example the number of reported bugs

The pilot runs on a real application and ends with a comparison to the state before

The fourth step is the pilot with the chosen team or teams on a real application under active development and maintenance, not on made-up cases tested on the side. At the start, Protopia supports it with mentoring and consultations for 1–2 production release cycles, and fits the length to the sprint, so the team goes through and tests the whole path end to end. Once everything is configured, the approach can start working after 1–2 sprints.

After the pilot, you compare its data with the state you collected earlier (the baseline) and check whether anything moved forward. You have to be ready for nothing to move. Some research found, for example, that very mature teams saw no significant speedup.

The pilot can also reveal areas to improve that do not depend on AI, such as acceptance processes. For example, the team builds and tests changes faster on its side, and the queue clogs at acceptance because someone cannot keep up with testing. At the start, this is a very welcome scenario: you then consider how to speed up or improve that stage. If the gain turns out small where AI was supposed to help, you may in the end decide against AI in that place. Success then means that the process got measured and improved in another part.

The pilot result decides on scaling and its reach

After the comparison, you have two paths. If you do not accept the improvement, you do not scale: you diagnose the constraints and perhaps go back to the foundations. You check which processes in the organization block the change, where the bottlenecks are, whether the problem lies in the pilot setup, and how review, testing and observability looked in this area. If you accept the improvement, you turn what worked into standards, templates and a definition of done.

Decision tree after the pilot. Two paths lead from the comparison with the state before. If the improvement is not accepted, you do not scale: you diagnose the constraints, such as bottlenecks or the pilot setup, and perhaps go back to the foundations. If the improvement is accepted, you scale: you turn what worked into standards, templates and a definition of done, prepare an adoption plan for part of the organization or all of it, and ambassadors and mentoring support the teams.
Diagram: the pilot result decides whether you scale or go back to the foundations

A rollout to the whole organization may not be needed. Depending on company size and criticality, part of the organization or a few development teams are sometimes enough. You update the standard and prepare an adoption plan for the projects that are ready and willing to adopt it.

One way to scale is ambassadors: selected people from the teams attend dedicated ambassador workshops, then show their teams day to day how the process should work. Teams that adopt the standard get expert support: ad hoc consultations and mentoring.

The standard sets the general direction and the main guidelines. Every project still needs final polishing, so the team always has some work to do before it gets value from the standard.

Speakers

  • Łukasz Kałużny

    Łukasz Kałużny

    Founder, Managing Partner, Technology Advisor.

    Łukasz Kałużny co-founded Protopia and is a Microsoft MVP in the Microsoft Foundry category. He has co-hosted Patoarchitekci since 2019, talking about IT architecture, GenAI and AI agents without the marketing spin.

    All posts by this author
  • Mikołaj Szczerbicki

    Mikołaj Szczerbicki

    Head of Sales & Business Development.

    Mikołaj Szczerbicki is Head of Sales & Business Development at Protopia and co-hosts the Powered by Protopia podcast. He scopes and prices projects, so he asks about the cost, risk and timeline of AI, Azure and Kubernetes work.

    All posts by this author

FAQ

Is the rollout a set of one-off workshops and demos for show?

No.

Do DORA metrics also help find where a speedup is not needed?

Grupa Pracuj rolled out DORA metrics as a diagnostic tool, among other things to know where a speedup may be entirely unnecessary. Episode 121 of the Patoarchitekci podcast covers this.

What about organizational resistance when you scale?

The speakers have seen more than once how organizational resistance appears after something is built inside a company. The standard goes to the projects that are ready to adopt it and want it. The team has an ambassador who represents the process day to day.