Going full agentic
Yontrack development has a long history. It started with version V1 around 2010, and then went through many versions:
- V2 - new UI, where Yontrack acquires its distinctive look and API
- V3 - modern stack, switching to PostgreSQL, GraphQL
- V4 - going enterprise level, auto-versioning, auto-promotions
- V5 - Next UI, modern authentication, environments & deployments,
During all this time, the development of Yontrack stayed pretty classical:
- creating issues, discussing with users, prioritising items
- starting the development, introducing a TDD approach very early on
- no pull requests, push to
mainand checking the CI pipeline - and of course, monitoring the Yontrack delivery with Yontrack itself π
This was working, quality stayed up thanks to the early Yontrack integration.
But there was already a big issue: procrastination. Some issues or initiatives just seemed too big. Even splitting them in small chunks proved to be too difficult. Why? Because, well, Yontrack was still a one-man affair. And some initiatives just became left on the side: security tracking, stack upgrades, UI modernisation, etc. I was mostly focusing on small things, keeping the boat afloat, focusing on users' immediate needs.
End of 2026, V6 is now around the corner, and it'll likely be the biggest release ever of Yontrack:
- Scorecard to quickly assess the software delivery maturity of projects or groups of projects
- Audit trail: build trail, endorsements, evidences on validation runs, non repudiation
- Security findings: findings model, neutral/SARIF/Trivy inputs, tracking of their fixes
- Delivery map views & graph views: seeing builds from different angles
- Full GitLab & Bitbucket Cloud support
- Mobile UI
- Upgrade of stack - Spring Boot 4, NextJS 16.x, JDK 25 upgrade
- Search on PostgresSQL: ElasticSearch dropped from search, βK palette, modern results page
- Ant Design 5 β 6
- and the list is not complete yet
The development of V6 only started on September 11th, we're now mid October. So many features and improvements in only one month. What happened? What did I change? Did I recruit 10 engineers in a couple of days?
The answer would come as being obvious nowadays: AI agentic development. What is vibe coding π± ? Was I still able to deliver with confidence (1)?
In the rest of this post, I intend to describe how I integrated AI in my development process, how it improved the delivery, but also the quality and many other factors.
How do I work nowadays?
Vibe coding π No, I'm joking, don't go away. The goal was to integrate AI agentic development, spec-driven processes, while keeping a very strong CI attitude (very frequent commits, main always green, always releasable, released every night in a demo environment).
I had already used Claude Code for quite some time (and before just some Chat AI to help me with the code) but even if it did improve my life, it stayed very superficial, just helping me in some tricky corners.
What changed everything for me was a presentation at the start of the 2026 summer where somebody mentioned the Matt Pocock's skills. Well, if you still don't use them in any way, please have a look, I hope it changes the way you see agentic development the same way it has for me. So in one word: grilling. In a few more words: code is not the issue any longer, focus on the specs, leave the agents deliver based on the specs. After a few days (hours?) of experimentation, my life as a developer changed completely: I was no longer a developer (yeah π - see note 2) and became a full-time product owner (yeah π₯³).
Preparations
A preliminary to my incoming new way of working was also to migrate from my good old self-hosted Jenkins instance to fully managed GitHub workflows. I knew beforehand that starting to work with agents would increase the load on the CI engine by a big factor. And I mean, it has, by 40 fold, as Yontrack could show me one month later:

Moving to GitHub allowed me to reduce the costs: free build agents vs. on-demand costly VMs on the Cloud.
On-par with switching from one CI engine to another, I also worked a lot on:
- hunting down any instability, any flakiness - any time a flaky test was detected, the number 1 priority was to stabilise it (4). No discussion. Really top priority. This became super important when the agents started to pile up the work (more on that later).
- reducing the duration of the whole "first feedback pipeline", going from ~ 1 hour down to less than 25 minutes (5): parallelisation, optimisation, etc.
Finally, adapting my local development stack so that several instances of it could run in parallel. Yontrack typically uses a Docker Compose file for local development and it had to be adapted to run on completely isolated ports from each other (9). This will allow local agents to work in parallel on the same machine.
Spec first, code & CI (like in Continuous integration)
The preparations were actually done in parallel of already starting to work with agents.
The main drive: spec first and continuous integration. Oh, and some code in between that I seldom see any longer (6).
Grilling and specifications
First round: specifications. Requirements, requests, etc. are born in Yontrack usually following discussions around the use of the application. Customers come with ideas, I come up with ideas, we discuss, etc. In the end, an issue is created in GitHub, with human-specified requirements (7).
As a product owner, I decide what needs to be done and in which release (major, minor or patch). I then enter a grilling phase where the goal is to make an issue ready for agent.
I typically start a grilling session for an existing issue like:
Grill me on issue 1234. At the end, update the issue and mark it ready for agent for milestone 6.0.The grilling session starts, the agent asks me aggressively about the limits of what was in my head or in the ticket. It can be sometimes one or two hours of work, sometimes with long pauses for me to think about it or to do additional research (sometimes just to understand what the agent meant).
In the end, the issue's description is updated and its labels mark it as ready for an agent.
Side note about defects - normally, a bug does not need all this grilling, I just go through the "diagnosing bug" skill, or pure Claude Code, and straight to implementation. The CI feedback loop is still respected though (see later), because it's enforced at the repository level.
Initiatives
Now is a good time to talk about initiatives. An initiative is a group of issues on the same topic that are ready to be picked by an agent:
- they can span several milestones (6.0, 6.1, etc.)
- they are all ready to be picked by an agent
- each issue is a piece of work ready to be delivered (very very important)
- issues can depend on other ones
Such an initiative is typically the result of an intense grilling session, typically done in several phases.
I elaborate a brief, manually, with my own hands on my own keyboard. The outcome is a docs/grilling/cves/brief.md document in Yontrack detailing what I want to do, like "I want Yontrack to be able to track CVEs and link this info with the time it takes to fix them", following by several ideas I have about this initiative. I can even commit and push this. No big deal, just some text, [skip ci], etc.
When I'm ready for this to start, I go and do something like (8):
Grill me on the briefing at docs/grilling/xxxx/brief.md.
The outcome is a series of issues, all in the same initiative, ready for agents.Heavy (and sometimes painful) grilling session ensues. At the end, I have a list of issues (from 5 to 20 sometimes), all labelled with something like initiative: cves, ready-for-agent and status: todo.
I'm now ready to launch the initiative and go to sleep (more on this later).
Implementing an issue
Back to a lonely issue that is ready for an agent (I'll come back to the initiatives, because they are so cool).
When I'm ready for an issue, I just run:
Implement issue xxxxThe "implement" here refers implicitly to the implement skill of Matt Pocock. Driven by some additional instructions in my CLAUDE.md file, it does:
- start a Git worktree (branches are so 2010s)
- implementation in TDD mode, based on the detailed specs in the ticket
- code & spec review - the spec reviews in particular are important for the consistency (does the implementation match what has been specified and all of it?)
- the execution all the tests locally, from unit tests to full blown UI tests, launching the local dev stack (Docker Compose)
- a direct push to
main(I hear some readers gasping - no worries, no build was harmed in the process, on the contrary) - the agent is not finished - it monitors the status of the CI on the
mainpipeline- if the CI is OK, the issue is closed and marked as ready -
status: ready - if the CI is not OK, and that's the beauty of it, the agent will keep trying to fix the issue until it passes (10)
- if the CI is OK, the issue is closed and marked as ready -
That's it, issue done. This can go from 5' to sometimes more than one hour, depending on the issue. Add to this the duration of the CI feedback and the number of times it had to be tried again (which is seldom).
Back to the initiatives and going to sleep (or not)
That is well and good for a lonely issue. In an initiative, imagine now 5 to 20 or more issues that you want to tackle, you'd have to go issue by issue, wait for them to complete and then go with the next issue, following their dependencies manually. No fun, repetitive, boring, ideal for an orchestrator agent π
I've created a run-initiative skill in Yontrack (11) and the only thing I do before going to sleep is to run the following prompt in a Claude Code session (13):
Run initiative <name>where the name is the name of the initiative in the GitHub issues (like cves in the initiative:cves).

This skills does the following things:
- it collects the list of issues for this initiative, orders them using their dependencies (leaves first, end nodes last)
- it asks me for a final approval - like "do you agree with this list of issues for immediate implementation outside of any control?" (more or less)
- I say yes and go to bed (more on this below, very important step)
- the root agent (or orchestrator, or majordomo, whatever you call it) will loop over the issues in sequence and for each of them (12):
- Start a worktree, etc. See the work done for a single issue in the section above. Told you, it becomes very repetitive and boring (which the the ultimate goal). At the end, the agent has pushed to
main, the CI is green and the issue is closed. - The worktree is always started from the latest origin
mainso that it piggies back on the previous issue
- Start a worktree, etc. See the work done for a single issue in the section above. Told you, it becomes very repetitive and boring (which the the ultimate goal). At the end, the agent has pushed to
Several giant benefits:
- Each issue being done results of the
mainbeing deliverable, even if the initiative is not finished (14) - Each agent, being run as subagent, is fully autonomous - it cannot ask for help. This is not a constraint - this is the foundation on which the whole setup is built.
- I can go to bed. I sleep. In the morning, I check the result: the whole initiative is done and done, deployed in my demo environment (deployment is done automatically every morning).
βοΈ That's this loop that has allowed me to deliver the V6 in just under a month and I've not sacrificed any good practice:
- continuous integration at its best - small chunks of work, always deliverable
- in my pipeline I trust - this is the keystone of the whole process, your
mainpipeline must deliver... trust! If you have to stop all the time, check the results or the flaky tests, don't even go this way! Fix your delivery first.
ποΈ One important factor of this story was able to go to bed. I'm not joking (sleep is important to me). But then, your computer where your initiative runs has to stay awake. I'm mostly using the native feature in Claude Desktop: "Keep computer awake". When the initiative completes (or stalls), my computer hibernates and I save a bit of energy. I also activate the "Remote Control" feature, handy to check the status from your phone and answer any problem if any.
π° I'm using Claude Max, that was the only way not to exceed the budget when using only a Pro subscription and have stalled sessions all the time. The bare minimum was actually pretty OK in the context of Yontrack, even for giant initiatives. Most of the time is spent on waiting for the pipelines (see note 5 again) and this does not consume tokens. I'm also always using Opus as a model (latest version) and even Fable (for tricky requirements, only for the grilling phase, never for the coding). Generally, most of my initiatives stayed under a 400K context (15).
What did not change?
I really loved refocusing on the what instead of the how. This allows for more quality time discussing with users, collecting feedback, thinking about new features or working on marketing stories. By the way, did you check the new website at https://yontrack.com ?
What did not change was the emphasis I've always put on the continuous integration (small chunks, fast feedback, trunk-based development, always working) and continuous delivery (always ready to be delivered, deployed in demo environments every night). On the contrary, this has gone even further than what I was doing before: faster, more stable feedback.
What does not work?
Still, nothing is perfect. Today, when I specify an existing issue, I lose the initial description, which is a pity. For issues that were created by users, I usually put the whole spec as a comment and it seems that the process is still happy with this. So that's something to explore for sure and made systematic.
The CI feedback is key! And today, it's way too long (20 to 25 minutes). I need to work on this. So far, parallelisation has worked pretty fine. Every job running on distinct agents is fine as long as you don't pay for them; I wonder how to do this in the case of private repositories where you'd have to pay for your agents (hosted or not).
What improved?
Quality and reliability
All in all, and I have Yontrack for Yontrack to prove it, quality and stability have improved. It's not only about the agents doing more work and faster than me, it's all about having a mature CI/CD ecosystem: you get speed and you get quality.
Effort in eliminating any kind of flakiness pays a lot! Just do it, and no excuse. Flaky tests had never any place in any software delivery chain, but even less now with agents.
Especially with recent models, agents understand the code way better than me, even my very own code: they always check everything, see everything. Sometimes, a grilling session comes with issues with the requirements that I would never have seen myself, especially on a codebase like Yontrack.
In the end, I have way more confidence in my delivery system now than I had two months ago, and that's a very good feeling (and that allows for good sleep).
Focusing on specifications
Starting from V6, I focus way more on the specifications than I ever did. I work more deeply in elaborating, designing new features, discussing them. I do mock-ups, I see more of the product that is Yontrack than its code (which I seldom see any longer).
Keeping a strong CI culture
I have already said that, but I want to put some emphasis on this: you cannot work at scale (with humans as with AI agents) if you don't have a strong CI culture! If you work only with humans (me with me), not having a strong and reliable feedback, you can live with that because time is on your side (coding takes so much time). But as soon as you unleash an army of autonomous agents, any flaky test, any unreliable CI system will make the whole thing collapse and drive you into agentic hell.
You may have noticed that I don't work with pull requests. That's a deliberate choice, that I had already done way before I switched to agents. Tests fully run at the agent side and once again on the main pipeline, reviews are done by the agent - I don't see the point in doing that again. After all, a PR is only about code and sharing with humans - the spec review is what matters at the end and there is no human in the loop in my case. Go continuous integration on the main branch π !
Reducing the implementation barrier
One great benefit of adopting agents was to reduce what I call the "implementation barrier" - what I described at the beginning of this article, as being overwhelmed by the complexity of a feature, a fix or a technology. No longer - almost nothing is impossible nowadays. I've done stuff in the last month that I was only dreaming of before. I could even tackle technologies that I knew nothing about.
Lost in the code?
You could argue that I will get lost in the Yontrack code. Yes, that's kind of true, but it's still not vibe coding either. There is a strong architecture foundation in Yontrack. Rules, and skills, and technical development guides are also part of Yontrack, constantly maintained, by me or by agents, followed by agents, enforcing how Yontrack is built.
If I have a question about Yontrack's code or its architecture, I have now the tendency to ask Claude directly instead of opening Intellij.
Agent autonomy
Going from chatting with AI about code to delegating the implementation of a complete initiative in full CI/CD mode has literally changed my job. It's like some blinds have finally been opened and I see the light again, I see my product, Yontrack, as it must be - helping engineers & release managers deliver software better, with confidence - and not as a bunch of code.
Future
While Yontrack V6 is almost out, while V7 starts already to have some briefs in place, my journey in agentic software delivery is not finished. This is a (very) fast moving world and what I will do in six months will likely be very different than what I'm doing at the present.
- I'm still experimenting with the frontier of what can be done with spec-only driven development, where the specs are the core of everything you do, and less and less the code.
- Build pipeline feedback being more important than ever, I really have to improve on their speed. This is not a nice-to-have any longer.
- My current process has proven to work very well on complete initiatives, or on a few parallel issues. What about complete teams working on a unique product? Software delivery, together with a strong CI/CD culture, has evolved so much that my current feeling is that we'd need less and less software developers to tackle product development, but more engineering, architecture & product management profiles. Future will tell.
- My agents are currently on my MacOS laptop (and I take care of not closing the lid at night) - what about making agents work on the cloud instead? From what I've researched, there is no immediate advantage for me - it would only add the cost of running agents on the cloud instead of on my laptop. It would however gain in isolation. I don't have any strong incentive, right now, to change this.
- What about AI in Yontrack itself? That's a complete different story of course and I don't rush into doing this. I don't want to add some AI features just for the sake of adding AI in my product - this happens too often around us. If a customer expresses some need where AI would definitely add some value, then yes, it'll be definitely considered: discussing, collecting feedback, briefing, grilling, launching the initiative, going to bed. Business as usual π΄
Notes
(1) well, besides using Yontrack of course.
(2) I've been in IT since 1998, I've gone through so many languages, frameworks and architectures (3), so I've known for quite some time that yes: the code is not what matters the most and this was never (for me) the most interesting part.
(3) I started coding before Internet... That's how you know you get older π
(4) And let's face it, fixing a flaky test with Claude Code: really easy. Sometimes on the fly, without any issue, using ci: commits. No excuse.
(5) I know... Still too long, I'm working on it.
(6) The only code I ask to see, and only from time to time, is some code (or script) touching somehow production. But not always, especially if the code/script in question has already been tested in CI.
(7) Not always, customers can create issues using AI, or the whole process will come up with additional issues.
(8) I do that so often that I really need to create a skill... Putting a knot on my handkerchief to remind me to do that...
(9) I'm currently exploring other options, like a local K8S cluster with Docker Desktop or K3S. Isolation may be easier to achieve - if you have ideas, I'm all ears.
(10) So far π€, I did not have infinite loops but ideally, I should improve the process by telling the agent to not try more than N times. After that many failures, drop the effort, revert the commits (and yes, yes, create a branch) and contact a human (me).
(11) Today, I'm using it only for Yontrack but this could be ported to any project and any issue system I guess.
(12) I tried running stuff in parallel. It was a Bad Ideaβ’.
(13) Note that there is an optional argument to specify the milestone when an initiative is spread among several ones.
(14) When you grill the initiative, you can even bring some feature flags. I don't necessarily do that: it's not because an issue is ready to be delivered that it's actually delivered. So until I do a release cutover, nothing is visible yet as an actual public release.
(15) This just gave me an idea for Yontrack π‘ V6 will also come with the tracking of the work per agent - it would be great to track the actual cost per agent contributing to your release. What do you think?
Final notes:
- I'm using a lot of
-as a punctation mark. I'm not an AI agent (so I say) and I'll keep using them. I'm a big reader and anybody who has read John Irving, among others, will know that carets are not the watermark of evil AI agents but have always been used. - I'm also a big fan of Terry Pratchett, hence the abundance of notes. Sorry about that.