AI Can Build Anything. Learn What Not to Ship
Organisations with successful products need to rethink their SDLC
Here’s how things look like at most startups I’ve worked with. Ideas, discussions, more discussions, reviews, consensus building and constant reprioritizations.
While leadership expects AI to drive up productivity through the roof, engineering teams still find themselves gasping for time if not running on fumes. In fact, the McKinsey Technology Trends Outlook 2026 cites that 1 in 3 companies that adopted agentic coding tools see a drop in productivity.
This is because while AI helps every individual involved in the SDLC, it is still not helping the organisation holistically.
It’s time to rethink the SDLC.
Instead of each idea going through engineering, engineering builds the systems that let ideas prove themselves.
Just like in factories, engineers don’t build cars themselves - they build what builds cars.
The new SDLC has 3 pillars
- Ship: A reliable platform that gets code, configuration, feature flags reliably & hermetically to production.
- Guard: A harness that ensures that existing features always work & core features are always performing per defined SLAs & SLOs.
- Measure: A worthwhile experimentation engine that allows wiring up of reliable A/B tests and accurately measure important user & business metrics.
Now, anybody with an idea can ship an MVP with AI and prove its utility on a small slice of real users.
The harness ensures AI has not messed up existing features. The deployment platform & flag framework brings it reliably to production. The A/B test framework reliably “turns on” the feature for a slice of live users. If any core experience is degraded, the experiment simply turns off.
Experiments don’t need scale, so it should rarely need significant infrastructure tweaks.
By principle and convention, this keeps everyone honest to what an “MVP” of a proposed new feature should be (Better get the CTO’s approval if you want a new data store, a “hot cache” or 3 new services for an experiment)
If the ideas are too over the top, they need to be scaled down into a series of smaller ideas. You necessarily convert a one way door into a series of two way door decisions. In fact, a series of smaller ideas can help prove a larger point - this is what agile was supposed to be!
The MVP is always kept simple. This is a necessary barrier to entry.
Once an experiment is live, we get 3 possibilities:
- It hurts key metrics. The proposer can iterate on it with AI or abandon the idea altogether & AI cleans up everything.
- The metrics are “Flat” or neutral. Neutral is good. It proves there isn’t a detectable harm. It’s also in “not proven yet” territory. Either iterate, or ship if that’s what was expected.
- The metrics look good. Now, onto the next small experiment to build on top of this OR we have a working, proven feature that is now a PRD - bring in the cavalry of developers who can debate, discuss, agonise over the best approach and implement it for scale.
Bringing successful experiments to 100% production scale will always be hard. But the key element of the new SDLC is that the collective engineering team always picks up, prioritises & works on proven ideas.
I can hear the “Buts”
- But you’re oversimplifying, our stack is complex. Maybe. But, in my experience, complexity rarely comes from product requirements - it stems from years of bolting on code. In fact, building the 3 pillars to wrap an existing stack allows you the opportunity to simplify it safely in the future. I’m writing an open spec and a set of AI skills to do exactly this; reach out if you want to talk through how this would work at your company.
- But it’s AI. It may break prod. It may. Just like prod may break today. You just tighten the harness. Here’s another way to put it - the team lead needs to think like a CTO. Today, a CTO doesn’t know every line of code written. They’ve setup the right harness - the developers & the culture. A team lead needs to think the same way about the harness. The harness is responsible for guarding everything - from functionality to user data.
- But a bad “MVP” may kill a good idea. True. Better scoped MVPs do evolve from iterations, and the traditional SDLC may shape ideas better when people talk. This can still be overcome by iterating with AI (remember - a flat result isn’t necessarily bad), or even in smaller focussed groups where experiment design & results are discussed.
Wins
- Measure on real users - in days. Testing ideas becomes an IC activity, and anyone with a great idea can test it out to harden their thesis. The fact that an idea is always backed by numbers means you’re settling “what if” debates with data. Forever. Won’t that be nice?
- Small, measured bets beat big, confident ones. I can’t count the number of times massive effort across an engineering org has been spent on a confident big idea that never panned out. Small, measured bets mean that an organisation’s time is spent on ideas with a higher probability of working. Speaking of time & money, a failed experiment costs one person a few days and a few dollars of tokens; a failed traditional project costs a team a quarter.
- A repository of things that did not work: Each failed experiment carries with it the context of things that didn’t work - and the exact metrics they hurt. This context shapes every future experiment. It’s a queryable decision log of what worked, what didn’t work and why it matters.
I understand that an experimentation-first approach is most valuable for consumer-facing products, where I’ve spent most of my career; but I think the core idea of tweaking the SDLC to figure out what won’t work still holds value for a lot of software companies.
In fact, while writing this piece, I remembered how KoBold Metals - a mineral exploration company - treats a Dry Hole as a goldmine of information. Their goal is to “Predict Everything” and “Quantify the uncertainty”, and each drill - whether it yields minerals or not - is used to refine their models. This actually helps reduce false positives in the future too.
In the software world, where building is now cheap, the hard part will be knowing what not to ship.
About Me: I am a software developer who started coding before there was a cloud. I’ve spent about a decade with startups of every stage and ~4 years at Google, where nearly every user-facing change shipped as an experiment first. Presently, I run my own software consultancy specialised in helping organisations build the right bespoke platforms to speed up their engineering (more about me here).

