operations · career · reliability · design
What Operating Mission-Critical Systems Taught Me About Software
Five years running systems where failure has a cost changed how I write code — and what I think 'reliable' means. Four principles I now apply to everything I build.
People imagine the IT inside a big organization as slow and old. After five years operating the systems at the heart of a busy public department, I learned something more useful: reliability is a human property, not just a technical one.
Nothing I learned in that job came from a textbook. It came from watching what happens when a system misbehaves at the wrong moment, and who pays for it.
The system is only as calm as its operator
A tool that confuses the person using it will produce bad data no matter how elegant the backend is. I started designing for the tired operator at 2 a.m., not the ideal user in a spec.
This sounds like a UX platitude until you sit on the other side of it. The person operating a live system is handling real cases with real deadlines while the software misbehaves. Every unclear label, every ambiguous error, every action that can’t be undone adds to a pile of stress that eventually shows up as a mistake in the data.
The operator’s calm is an input to system quality. Treat it that way and a lot of design decisions make themselves.
Audit trails beat cleverness
When something goes wrong — and it will — the question is never “did it fail?” It’s “what exactly changed, and by whom?”
That’s not a nice-to-have. In an environment where records have to be defensible, an entry that changed without leaving a trace is a problem bigger than the original error, because now you can’t reconstruct what happened. The timeline is the feature.
That’s why change tracking is a core feature of the CMS Extension I built, not an afterthought. When a record’s classification changes, the system records the before, the after, and the timestamp. Accountability isn’t documentation you add later — it’s something the schema has to make possible from day one.
The failures that actually cost you
In operations, the failures that hurt are almost never the dramatic ones. Nobody loses sleep over a hard crash that everyone notices immediately. The expensive failures are quiet:
- The alert that didn’t fire. A deadline passed and nobody was told.
- The duplicate record. Two entries, one truth, and now every report is wrong in a way that’s hard to spot.
- The timestamp that doesn’t match. A record says one thing, the log says another, and now you have an audit question.
- The form that submitted but appeared not to. The user retries. Now there are two records.
Each one is individually minor. Together they erode trust in the system faster than any outage — and a system people don’t trust gets worked around, and workarounds are where bad data is born.
This is the single biggest lesson I carried into my own development work: software’s job is to reduce surprise. Everything else is decoration.
What I carry into my own projects
Quiet interfaces. Less animation, more clarity. A system that moves for decoration spends attention that belongs on the data.
Reversible actions. Every destructive step should be traceable, and ideally undoable. Confirm before you delete, not after.
Graceful degradation. The system should fail small, not fail loud. A banner that says “you’re offline” and lets you keep working beats an error dialog that blocks you.
Assume the network will vanish. The same instinct, applied to delivery. If a user can lose work because a request timed out, the design is wrong regardless of how it behaves when everything is healthy.
Why this shapes what I build now
When I started building my own products, I expected the hard part to be the technical work. It isn’t. Building an internal system for operators who already know the job and will tell you exactly what broke is, in some ways, easier than building something new from scratch.
That’s what I wanted. Aafiyat, the Nankana Home Care tools, the CMS Extension — they’re all shaped by the same question I learned to ask on the job: what happens when this fails, and who deals with the consequences?
Abdullah Tayyab — Full-Stack Developer, Punjab, Pakistan.
Enjoyed this? Say hello →
Comments
Loading comments…
Comments are reviewed before they appear. No account needed.