A few days ago I recalled a story from the old pre-LLM era of IT. It actually happened only a few years ago, but nowadays it feels like a completely different age. So instead of sharing yet another set of AI thoughts with you, I decided to slow down for a bit and tell you this warm, cozy old story.
So, there was an internal tool in our company that was used for some business processes. In general, "internal tool" doesn`t mean a poor-quality or buggy app, but this particular one was exactly that. And the most annoying thing was that roughly every 10th call to the database triggered an intermittent bug that crashed the whole app. It simply threw an IndexOutOfRangeException during a call to the DB through a DataRow, and there was no crystal-clear reason for it. It was old software, but what`s even worse, there was no particular team or developer who owned it. Different people modified it from time to time according to their needs, with no architectural requirements, no overall vision, nothing. A few developers had also tried to fix the IndexOutOfRangeException bug, but their attempts had been unsuccessful. And then it was my turn to try.
After some debugging I assumed that the bug was caused by incorrect multithreaded access to the DB. But calling the AcceptChanges method (which commits the current changes as the new state) didn`t help. So the next option I saw was to wrap the methods that threw the exceptions in a lock statement. That didn`t help either. After many more hours of debugging I finally found out that there were actually many DataRow objects: the business logic spawned a lot of threads, and each of them owned its own DataRow (omg).
And that is where the real problem was hiding. A DataRow doesn`t live on its own - it belongs to a DataTable, and that table was shared by everyone. So while one thread was reading its row by index, another one was modifying the table`s row collection underneath it. The internal storage got rearranged, the index the first thread was holding no longer pointed to anything, and the app crashed with an IndexOutOfRangeException. Nothing looked wrong in any single place of the code, because nothing was wrong in any single place of the code - the bug only existed in the way objects were spread across threads at runtime.
So the fix was to make all the threads go through one single shared synchronization object instead of each having its own. I created a separate service that owned that object and registered it in the DI container with a singleton lifetime, which is exactly what a singleton registration is good for. To be fair, the container itself wasn`t strictly necessary here: a plain static readonly object would have solved the problem just as well. Anyway, it did help - the bug was finally gone.