andregoepel.dev
Start a conversation
All posts
Software Architecture 5 min read

Changing Architecture in Production: Lessons from a Long-Running Platform Project

Working on a platform for several years means living with the consequences of earlier architecture decisions. Some hold up longer than expected. Others create work under real load that was barely visible at the start. In one long-running project, I helped analyse these limits and drive corrections. Implementation was a team effort. Here is what those experiences can teach us about the next change.

.NETSoftware ArchitectureModernizationLegacy Systems

Production brings new evidence

In a platform project spanning several years, part of my work was following technical problems over time. The application ingested data from different sources, processed it, and made it available for further use.

As usage grew, limits emerged that were difficult to assess in a controlled environment: extra deployment effort, troubleshooting across several components, and ingestion that was sensitive to connection problems.

These observations are initially clues. One slow request does not justify an architecture project. It becomes more interesting when the same kind of problem recurs and local fixes merely move the effort elsewhere.

First correction: bringing services back together

The platform had been split into microservices early on. In operation, that introduced additional deployment coordination, latency between services, and diagnostic work. In that implementation, the benefits of independent scaling did not sufficiently justify those costs.

We consolidated the platform into a modular monolith. Deployments became simpler, troubleshooting more direct, and runtime behaviour more predictable.

That was a correction for this platform. I would not turn it into a general rejection of microservices. What mattered was the independence the existing split actually provided and what it cost in everyday work.

My lesson: an existing boundary should have an explainable purpose. If two components must be changed, tested, and released together anyway, it is worth asking what their separate deployment still provides.

Business boundaries remain important after consolidation. Fewer processes do not replace clear responsibility for code and data.

Second correction: decoupling ingestion from processing

Another problem area was ingestion from different environments. Production load exposed connection interruptions and timing problems that were hard to reproduce.

We changed the flow to processing through a persisted queue. Reliability improved. The important difference was that accepting work and processing it were no longer coupled in the same way in time.

A queue can buffer incoming work and absorb temporary demand peaks. However, if arrivals consistently exceed processing capacity, the backlog grows. Microsoft explains this relationship in its Queue-Based Load Leveling pattern.

For this kind of change, I would explicitly examine repetition alongside the normal path: what happens if the same job is processed again? How do permanently failed jobs become visible? At what backlog level should someone intervene?

The queue is a technical component. Whether it makes the business process more reliable also depends on those answers.

Another lesson: every layer needs a purpose

At the time documented in the project account, simplifying heavily used import paths was another area of work. Additional API layers introduced detours and made data ownership harder to follow. I am deliberately describing this as work under way at that point, rather than claiming a third completed success today.

The underlying question is useful beyond this project: what responsibility does a layer fulfil now?

It might control access, centralise business rules, or provide a stable contract for other components. Those responsibilities do not disappear when the layer is removed. More direct access only improves the system if the necessary rules and ownership remain clearly established.

How I would structure a change like this today

The following steps are my recommendations for comparable projects. They are not a reconstructed delivery log of the project described above.

1. Start with an observable failure

“The architecture is too complicated” is insufficient. A useful starting point might be: after a connection interruption, import jobs remain stuck and require manual attention.

That makes it possible to describe an objective. For example, a specific failure should allow controlled recovery. I would choose the technical means after establishing that objective.

2. Change the smallest meaningful area

One import path, module boundary, or deployment step may be enough for the first correction. The change should be large enough to solve the problem and small enough for its effects to remain understandable.

For larger replacements, the Strangler Fig pattern describes a gradual transition in which old and new functionality temporarily coexist. It is one possible approach; the case described here does not establish that this particular pattern was used.

3. Decide in advance what happens if it fails

Starting the old application version again is not always sufficient. If the new path has already changed data, that data must remain usable by the previous path or be recoverable through an explicit correction.

Before switching over, I would document who can stop the change, which states must be preserved, and how failure will be detected. These questions belong in the architecture decision, rather than waiting for the final deployment meeting.

4. Check the effect and finish the transition

After the change, return to the original problem: are fewer manual interventions needed? Can failures be traced more quickly? Has the bottleneck merely moved somewhere else?

Once the new path proves itself, removing the old one needs a date too. Otherwise, the team remains responsible for understanding and maintaining two alternatives indefinitely.

What becomes valuable over several years

My most valuable contribution was not a single architecture diagram. It was knowing the system long enough to recognise recurring problems and develop workable corrections with the team.

A decision can be changed later. What matters is that the new evidence is understandable and the transition remains manageable for the people who continue developing and operating the platform.

Does your platform need to keep running while you change it?

I help companies analyse established .NET applications and break modernisation into manageable steps. In a free introductory call, we can discuss where your platform creates work today and which first intervention might help.

Get new posts by email

One mail per article. No newsletter theatre, no tracking pixels.

Unsubscribe with one click. Never shared.