Custom Operations Software for SMEs Is Now Trivial to Build. With the Right Harness.
Keep finance and inventory in a packaged ERP. Build the operations layer custom. AI coding models inside the right harness made that affordable for an SME, and here is the method and the evidence from my own repository.
For most of my career the answer to “should we build this ourselves?” was no. I implement ERP systems for a living, and I have told many SME owners to bend the process to the package instead of paying for custom code. For finance and inventory I still say that. For the operations layer, the planning, dispatching and exception handling that is specific to your industry, my answer has changed.
What changed is that AI coding models, working inside a written set of rules and checks that I call a harness, made a first working version of that layer affordable. Below is where I draw the line, how I work, and what it has built.
Custom operations software for SMEs starts where the 80% ends
I use “the 80%” as shorthand for the part of a business that looks like every other business: buying, selling, stock, invoicing, the ledger. It matches my own years of ERP work, and it is an industry rule of thumb. Paul Saunders, Head of Product Strategy at SAP, put it in print: “approximately 80 percent of business capabilities and processes are non-differentiating”, and the other 20 percent is where a company can differentiate. It is a rule of thumb and not a measurement.
Gartner’s pace layering framework draws the same line without a percentage: systems of record on one side, systems of differentiation on the other.
Microsoft publishes no such number, but its guidance points the same way. Its implementation guidance says Dynamics 365 is designed to meet standard business processes, and it tells customers to “adopt wherever possible, adapt only where justified”. Specialised and industry-specific processes are where it says partner solutions tend to focus.
To be fair to that guidance, its order is: configure the standard, then buy a partner add-on, and only then extend. A Dynamics partner will tell you an add-on exists for your problem, and often that is true. My point is about what is left when the add-on does not fit, or forces your process into its shape. That remainder usually ends up in spreadsheets that bridge systems.
Finance, VAT, e-invoicing and inventory stay in a packaged ERP or accounting system. The reason is not that they are hard to code. The rules change every year and somebody has to keep up. Belgium made Peppol e-invoicing mandatory on January 1, 2026 and has a draft law for near real-time e-reporting from 2028. Germany requires businesses with more than €800,000 in prior-year turnover to issue e-invoices from January 1, 2027. A vendor absorbs those changes for all its customers at once. A custom ledger makes each one your project.
The custom layer talks to that core through its API, the documented door the vendor provides for other software. The ERP itself stays unmodified, which is the route with the least upgrade risk.
Germany’s 2027 e-invoicing deadline·lands in 5 days
Dynamics NAV End of Support: what custom code inside the ERP costs you later·lands in 8 days
Why the operations layer used to be out of reach, and what changed
Start with what is measured. In Eurostat’s 2025 survey, 41.08% of small EU enterprises used an ERP and 24.69% a CRM, against 88.71% and 65.43% of large ones. Eurostat measures use, not the reason for it.
Licences are not the barrier. On October 1, 2026 Microsoft listed Business Central Essentials at $80 per user per month, and Odoo listed its Standard plan at €24.90. The expensive part was always people: implementation, and above all bespoke development that then had to be maintained. I have no independent figure for that cost and will not invent one. The closest is Panorama’s 2026 ERP report, where the most common reason for a budget overrun was an unexpected need for additional technology, custom builds included, after a misfit was discovered late. Its respondents have a median revenue of $200.5 million, so read it as direction, not as SME data.
The arithmetic on the other side is what changed. One Business Central Essentials seat is $80 a month. For $200 a month I have a Claude Max subscription at its 20x tier, which includes Claude Code, and that takes one builder a long way. Hosting is no longer a project either. On Heroku, an app with a server, a web front end and a database runs for roughly $75–100 a month: on its list prices of October 1, 2026, a Standard dyno at $25 or $50 plus a Standard-0 Postgres database at $50.
Security comes with it to a large degree. Heroku patches and replaces the underlying systems for you, and it runs on infrastructure accredited to ISO 27001 and SOC 2. Securing the application itself stays your job, which is what the standards further down are for.
Neither figure includes the time of the person who specifies and owns the system. But the tooling and the hosting together now cost less than four ERP seats.
It also changes who can do the work. You can do this in-house far more easily than before, which means no more waiting, or doing things by hand for weeks or months, until a more custom solution finally arrives. And if you would rather outsource it, that has become cheap too. Our sister company, Swiftly Developed, is a software studio that builds native apps, web applications and the backend systems behind them.
So what do I mean by trivial? Something narrow: a first working version of one specific operations workflow, built by someone who can specify it precisely, inside a harness.
Independent studies do not support “trivial” in general. METR’s 2025 randomised trial found that 16 experienced developers took 19% longer with AI, and METR calls its own 2026 follow-up an unreliable signal. Anthropic’s study of its own staff found most could fully delegate only 0–20% of their work. Google’s DORA 2025 report links AI to more throughput and less stability. The evidence in this article is my own repository.
Two sessions: one to think, one to build
A session is one working conversation with the model. I keep two kinds apart. The first is a sparring session about the specification: what the workflow must do, for whom, with which exceptions, and what is out of scope. I run it with a top-tier model: Claude Fable 5.1 or Opus 5.5 on the Anthropic side, GPT-6 Astra or GPT-6.1 Sol on the OpenAI side. This is the only interactive part of the whole method.
I put the model in a challenging mentality. Its job is not to agree with me. It challenges the solution I propose, looks for the gaps in my own knowledge, and keeps asking until the specification is as complete as we can make it.
I ratify the specification before execution starts. The build session refuses to begin on a specification that is not ratified or still has open questions. It does not guess what I meant.
The main session orchestrates, sub-agents execute
In the build session, my main session only orchestrates. Sub-agents do the executing: fresh AI workers that each get one job and a clean context of their own. The goal is long-running sessions, and several workflows in parallel. This is the loop:
Phase split: the work is cut into phases.
Phase validation: an independent sub-agent checks the split.
Phase correction, if needed.
A plan document for each phase.
Plan validation: an independent sub-agent checks each plan.
Plan correction, if needed.
Plan execution.
Validation of the executed phase by an independent sub-agent.
Correction, if needed.
Testing per workflow: unit, integration and UI.
Lessons written back to the standards.
Archival of the working documents.
An interactive diagram of the specification session and the twelve-step build loop. Before the loop, the specification session is the only interactive part: the owner spars with an agent told to challenge the proposed solution and look for gaps. The twelve steps are all done by agents, in four rows of three. Phases: split the work, an independent check of the split, a correction if needed. Plans: a plan per phase, an independent check of each plan, a correction if needed. Build: execution, an independent check of the result, a correction if needed. Close: testing, lessons written back to the standards, archiving. Steps 2, 5 and 8 are checks by a separate agent. A few questions can still reach the owner during steps 1 and 4. After the loop, the owner approves every merge into a protected branch and every deploy.
Phases
Plans
Build
Close
done by an agent
checked by a separate agent
only if the check finds something
you are involved
All twelve steps are done by agents. Select a step to see who does it and where a question can still reach you, or press Walk through.
The build loop. Select any step, or walk through it: the specification is the only interactive part, the twelve steps are done by agents, and the owner approves the merge and the deploy.
All twelve steps are done by agents. I still get questions during the phase split and the phase plans, but they are minimal, and not every sub-agent session asks one. Everything I had to say, I said in the specification.
Two things keep it controllable. Within one workflow the sub-agents run one at a time, and the parallelism comes from running several workflows side by side, not from agents racing inside one. And I approve every merge into a protected branch and every deploy myself. The harness reduces effort and error. It does not remove the owner.
One honest footnote on step 8. Checking each phase after it is executed is my practice. The written workflow spells out the independent checks before execution and at the test stage, and not yet this one.
Archival is housekeeping with a purpose. Finished working documents move to an archive folder, so the queue of work in flight and the repository stay readable, for me and for the next session.
Where the checks actually catch things
The plan checks earn their place. On one feature, measured on September 8, 2026, the five plan checks found 26 defects. Every one was in how the plan proposed to prove the work was done. None was in the design. The models reason well about what to build and poorly about how to prove it, so the checkers now run each success criterion instead of reading it.
The second example is humbling. One bug passed 19 phases of green checks and showed up only when the real server was started. Since then, starting the real server and calling it is a step of its own.
Let the AI write the tests, then make it run them
It is worth letting the AI spend real time writing tests, in whatever framework and language your project already uses, because tests are how it checks its own code without waiting for you. My standard has five layers: unit tests for a single function, integration tests for the full round trip to the database, UI tests that drive the actual screens, migration tests against production-shaped data, and performance tests. Each behaviour is tested once, at the cheapest layer that can prove it.
Writing is not running. One piece of work was reported complete with 12 test flows written and compiled, and zero executed. I found out only because I asked. Now a suite counts only when the run reports how many tests executed, and a blocked run means not done.
Paul Hudson made a related point on July 17, 2026. He describes asking the model for an adversarial sweep of unit tests in which every test has to earn its place. His caveat is that you must “make sure your tests are pulling their weight” and check that each one tests what you think it does. His post is about unit tests. Extending the idea to integration and UI tests is my step, not his.
Paul Hudson's post on X from July 17, 2026, shown as a card with a button that loads the original post from X. In the post he shares a prompt for a coding model: an adversarial sweep of the unit tests that targets edge cases, malformed input, race conditions and boundary values, keeps only tests that earn their place, fixes failures, adds a regression test for each bug found, and repeats until the suite passes cleanly.
Paul Hudson@twostraws on X · July 17, 2026
A prompt to hand a coding model: run an adversarial sweep of your unit tests that goes after edge cases, malformed input, race conditions and boundary values, keeps only the tests that earn their place, fixes what fails and adds a regression test for each bug it finds, and repeats until the suite passes cleanly.
The limit is real. Tests written by the same model can encode what the code does and not what was intended. That is why the specification comes first and why the validation is done by an independent sub-agent.
One agent codes, another one tests
Someone needs to write the code, and someone else needs to test it. Grading your own test does not work in real life, and it does not work for a model either. An agent that just built something reads its own work with the same assumptions it built it with, so it tends to confirm what it already believes.
So in my loop the agent that writes never signs off on what it wrote. The split, the plan and the result are each checked by a separate agent that starts with a clean context and only the specification and the plan to judge against. The same goes for tests: one agent works out which tests are needed, a separate one checks that judgement, and only then are the tests written and run.
Finance teams will recognise this as the four-eyes principle. The person who enters a payment is not the person who approves it. It is the cheapest control there is, and with agents it costs one extra step.
The standards: how the AI remembers your preferences
A model starts every session knowing nothing about the last one. The standards are how it keeps track of my preferences and learns from its mistakes: 28 plain-text files, 17,469 lines, read at the start of the kind of session they apply to. For a business reader, these are the ones that matter:
Tenant isolation: one company’s data is never visible to another.
Permissions: who may see and change what, down to the field.
Audit trail: a tamper-evident history of who changed what.
Safe database changes: how the data structure may change without an outage.
Security and compliance: a security baseline, and which capabilities are allowed in which country.
E-invoicing: the Peppol rules.
Testing and user interface: the five layers above, and how screens should look and behave.
After each piece of work, every sub-agent notes what it learned, and those lessons are written back into the standards. The repository shows 54 such write-backs so far.
If you want to work this way, start here, before the first feature: outline your standards. Do not copy mine. Research your own, according to your own needs and preferences: your industry, your stack, what you will not compromise on. Then keep building them along the way. Mine grew that way too, as the next section shows.
Rules that exist because something broke
Two database outages, on May 9 and May 11, 2026, became the rulebook for database changes. A change that works on the live database can fail on a fresh one, so every change is now tested on a fresh database and on production-shaped data.
A light load test, 20 requests per second on one heavy function, ran a 512 MB beta server out of memory. It did not recover by itself. Load tests against a shared server that testers rely on are now forbidden.
A lessons write-back shrank a plan from 668 lines to 119 and lost the test agreement a later phase depended on. Lessons are now appended, never rewritten.
A sub-agent was cut off by a usage limit with roughly a third of a session’s work unsaved. Work is now saved after every completed unit.
None of these rules came from a best-practice list. Each one cost me something first.
Keeping the instruction file short is recurring work
The instruction file (AGENTS.md, or CLAUDE.md in Claude’s tooling) is what the model reads first. My rule is that it should be a bare skeleton that only says what to find where, under 100 lines if possible. Anthropic’s own guidance is more lenient: under 200 lines per file.
It does not stay short by itself. Every lesson wants to add a line and nothing removes one, so the file grows until the model is reading a rulebook before it reads the task. Compaction is recurring work: move the detail into the standards, leave a pointer behind, and do it again a few weeks later.
The structural answer is a retirement agent, whose job after each piece of work is to remove the rules that work made false. Its brief says its success is measured by what it removed, not by what it added.
What this built
Two products came out of this way of working. FeedbackKit, a feedback platform for app developers, is built with the same method in its own repository with its own rules. I make no measured claim about it here. Swiftly Workspace is the repository every method figure in this article comes from: 27,197 commits between October 2, 2025 and September 26, 2026, of which 20,636 carry a Claude co-author line.
A total redesign of the Workspace website is on its way. It is perhaps a bit early to share the link, so if what you see still looks like the old site, come back in a week or so.
Both are my own software products, not operations systems built for a client. That matters for how far you can lean on them.
The bottom line
Keep the book of record in a package and let the vendor carry the statutory change. For the operations workflow that no package fits, a custom build is now within reach of an SME, provided someone can specify it, a harness checks the work, and a named person owns it afterwards. Trivial is the right word for the first working version. It is the wrong word for everything around it.
Dynamics NAV End of Support: What to Decide Before January 2028
Extended support for Dynamics NAV 2017 ends in January 2027 and for NAV 2018 in January 2028. What that means in practice, the three realistic options, and a plan that works back from the date.
Comments
Loading comments…