Housekeeping of test cases

Invisibly but firmly, quality of tests make or break software projects. The production code brings money today. But it is the quality of its accompanying test suite that ensures tomorrow’s production code keeps bringing money going forward.

Let’s talk about some routines that people should perform from time to time on their automatic test suites.

The most valuable thing that can be done to a piece of code, production or test, is to be deleted.

Consider a test case for removal if nobody uses the production feature it exercises

Tests are also a liability. Their price: time spent on regularly running them, but more importantly it is time spent on maintaining them.

If a test case causes your team to spent more on it than it brings in value, then the test should be removed, simple as that. Precisely measuring the “effectiveness” of any piece of code is not well defined, but there are certainly indicators of low value ones.

Is there a test case that exercises a feature of your production code that is no longer worthy to maintain? It could be a module that no customers use any more, or a host platform/bitness/OS/environment/configuration that none or very few customers continue to use? If that is the case, removing tests for that obsolete feature will release resources currently spent on it for something else.

One immediate pleasant result of stopping testing a subsystem is that the subsystem in question does not have to work correctly anymore. Moreover, it can be deleted, outright or gradually, from the code base! That is even less code to maintain for your team.

How to tell which features are not used by your customers? Well, add counting how often entrypoints to that feature are invoked by customers’ workloads, to your telemetry system. You do collect telemetry for your product, right?

Check that test cases can still fail

I have stopped thinking that that tests can fail. Fine, surely they can error out as any piece of software, but that is not why we keep the tests around.

Test cases are small programs that react on behavior of another program — your production code. Their dual reaction includes reporting an assertion failure or reporting nothing (except for the return code zero) for success.

We are usually get notified whenever a test case reacts in its assertion. We are not explicitly notified when the same test passes. A passing test creates a sense of satisfaction. It is usually means less work for us to do.

But is this sense of safety warranted? That test case that has just passed — will it react properly when the time comes to do the opposite? A test that can never fail is less than useless, it is strictly a waste: it wastes time, and it fools its users.

Find a way to look through execution history for a given test case. Ideally that history should include not exclusively test runs done in the CI environment, just before a pull request is about to be approved and merged. Pull requests that are sent to CI are usually biased towards including more stable production code that fewer tests would react on. No, it is testing done in the development environments, by developers — that is where we want our test cases to really serve us, being close to the changes. You want statistics over all kinds of states of the production code, especially on corrupted, messed up and work-in-progress variants.

Suppose it has been months or years since the given test case has last failed. It is time to take a deeper look at it. By looking at its code, can you devise a regressive change to the production code that could cause the test to fail again? If you cannot, then you have a infallible test case.

Maybe somebody has accidentally disarmed the test case some time ago by forgetting an early return in it, or commenting out its assertions? Maybe it was disarmed right from the start, because nobody has ever checked that this test case can and will fail?

A buggy test that cannot fail should be fixed. An infallible test that cannot possibly fail, should be deleted.

Check that, when the test reacts, its message can be understood

We all know those test cases in our code base. They fail, oh yes they do. Everyone dreads the moment when they do fail, because what it means is that you or some other poor soul will have to dig through mountains of logs, find a way to attach a debugger in a hostile environment not meant to be used for interactive debugging, add sprinkle debug prints all over the place. To do a lot of manual work to understand what happened in the production system to cause the test to react.

Would it be nicer if, whenever it reacts, a test would tell you precisely what went wrong, and why it went wrong? Who writes the error messages that our tests leave? Why, it is us. Would it be nice if the message contained all that information we so much crave? Well, it is in our hands.

When you meet a test case, which error message you cannot understand from a glance, after making sense of it, spend five more minutes and improve its diagnostics so that the next time it fails, it would be easier to understand the context of the failure.

If a value is expected to be true but it is not, explain what that value means in the message (not in a comment, because it is not visible until someone opens the test source code). An error message “val is not True” leaves the reader wishing for more context; and that reader is usually a future you. If two values in expect(a, b) are not equal, give descriptive names to the compared values. If the compared values are compounds (lists, dictionaries, trees, etc.), use a more sophisticated comparator/printer so that divergence in their subcomponents are easier to notice.

Maybe unsurprisingly, small, simple and focused tests are easier to improve in this manner. They usually check one thing. Usually it is possible to make their assertion message to express that thing.

I call test cases with error messages as unhelpful as “something went wrong, human, go debug me” — arrogant tests. They certainly know more details, but they won’t tell you them upfront.

Remove redundant prints on success path

Who reads logs of an automatic test that has passed? Nobody (unless it is currently under suspicion that it is infallible, and so you aim to make it react). Who reads logs for a reacting test, wishing it would just tell you what went wrong? It is you!

Do you need all those prints that trace “normal” execution of a test case? Not really: when it passes, nobody even opens those logs. When it fails, everybody just jumps to the end of the log file, because that is where a problem usually laid out.

What if a warning about an imminent test failure appears somewhere earlier in the test’s flow? Then it has high chances to be lost in the sheer volume of non-essential messages, if the test is not shy to tell every little minute detail.

Alarm fatigue is a real thing, even for programmers reading logs. Do not show a message in a context where you do not expect a human to read it and take action based on it.

If you see a redundant message left in a test case, remove it. If you do not have a heart to remove it outright, lower its logging priority so that the message is not shown by default. When the test passes, nobody will care about it disappearing. When if reacts, it should be trivial to activate the extra logging again. But your failure messages are made sensible and self-explanatory, right? So again, all the now concealed verbose logging is not really needed, so good riddance of not having to filter it.

But if you find yourself reading the same log file of the same test case over and over again and often needing those intermediate logs, you have a bigger problem on your hands: a manual test that pretends to be automatic. Human-in-the-loop manual tests are needed sometimes, but they are much more expensive in terms of time, and they do not scale as well as automatic tests.

Remove debugging cruft and copy-paste

Prints in tests are a form of debugging scaffolding. Other stuff that may reside in our tests without carrying enough value are copy-pasted sections taken from earlier tests, which were essential there but here, they are preserved just-in-case (i.e., because you do not really understand what they are for). Conditional blocks for code paths that are never visited in the course of test suite execution should go away: you are the only customer of your test suite, nobody will run the suite in a way that would invoke those “just-in-case” code paths.

If you can remove a line of code from a test case, and the case still passes and fails the same way on the same changes to the production code, then the removed line is truly redundant.

Check for early bailouts

Related to the “can this test case ever fail?” question, there is an issue of test cases that are allowed to bail out early from their execution. Usually they contain early returns before some (or all) assertions in their body have been reached.

Whenever such test case passes, you are left wondering: did it pass because all of the assertions were true, or because it bailed out early?

Ask yourself: are there any environments in your CI when the bailout condition for the test case is false? I.e., does the test regularly (or ever) exercise the system-under-test, or does it always terminate early?

Tests should not decide their own applicability for themselves; that has to be determined on a higher-level of the test suite management framework. A test case should only do its job, and do it without leaving uncertainty of whether or not the system-under-test has truly been subjected to it. If you must conditionally skip some of the test cases, you should have a monitoring system that allows you to timely discover whenever a test case becomes infallible because of that.

Check adherence to the AAA rule

Do you have multiple test cases collected in a single test file? When it reacts, and a human tries to comment out an earlier section of the test file to isolate the failing case, and suddenly nothing inside the whole file works? Well, this file is suffering from so-called assertion roulette. That, in turn, is a consequence of another problem: assertions are interleaved with acts and arrange phases. It means that later test cases implicitly depend on the state left after their predecessors.

Why does this happen? I found that, when I create a series of related test cases, it is very convenient to add them one-by-one into the same file. As soon as the latest case passes, I add a new one that fails, then make it pass by changing production code, and repeat the the loop until done. No need to open a new file, give it a name, copy-paste the arrange steps etc. The system-under-test in the existing file is already in a known good state, so it would not hurt if I just attach one more test case that makes use of that known good starting state, right?

While this process follows the rules of test-driven-development, it is optimized for the writer of tests. But when a test case fails, we want it to be optimized for a reader of tests, the person who debugs them and tries to make sense of them. The reader wants test cases to be independent from each other, each one starting from a clean slate and doing its own arrange independently from any other test case. A reader benefits from having each test case isolated in its own file, with phases in a strict order: a single arrange followed by a single act and arriving to a single assert.

We should optimize test suites for readers, because tests are read more often than they are created. If you find yourself dealing with such a writer-optimized test file, consider splitting it into independent test cases. It is not always trivial to do, but if done once, it will help everyone else who returns to these test cases later.

Split into smaller, more focused sub-cases

Is there a test case that causes your team a lot of maintenance headache? Have you tried to delete it? You cannot because it tracks an important customer use case? Well, have you tried to do the next best thing after that — making that test case simpler by splitting it into many focused test cases?

Try to make it faster

Do you find yourself not wanting to run a test because it would force you to sit idly until it finishes? Or, more realistically, to make a context switch to something else, or to continue working on the same thing while only hoping that you will not have broken that long test?

Chances are, you are facing a slow test. It is the worst. Usually such a test accumulates a lot of other sins within itself: it spews a lot of logs nobody reads, its error messages are cryptic and are lost among all the unwanted output it produces, it has bailouts for some of its checks and alternative code paths depending on environmental factors, it iterates over a series of almost identical subcases, but some of them depend on the exact order in which everything is invoked. Beyond that, its subcases may depend on slow subsystems you do not fully control (nor really care for) in the particular test context, such as network, filesystems, authorisation, databases, etc.

If such monstrosities can be tamed and made to run faster, it is often achieved by splitting them into smaller, more focused, independent files that can be run in parallel. If the test does the same arrange phase over and over, maybe its earlier execution results can be somehow cached, constructed synthetically, replaced by test doubles, or bypassed. There is no one easy solution to the problem of slow tests.

It will not be easy to speed up a slow test. The payoff of achieving that is increased team productivity and higher satisfaction from the process, and these are goals worth pursuing.

Pay attention to issues pointed out by your tools

To repeat myself, test suites make or break software projects. We should treat code that constitutes our test suites with as much attention as we treat our production code. In particular, we can and should test our tests. I wrote about how to test the tests before and their asymmetric relation to production code.

For example, static code checkers can reveal unreachable code branches of forgotten debug returns, tautological comparisons in copy-pasted assert(value, value), overly complex setups (but you usually know about those places already, since they also create fragility and require constant maintenance).

Even running a spellchecker against sources may reveal semantic issues, such as an assignment to a mistyped variable, which leaves the original variable unchanged, etc.

Dynamic analysis can also be of great use to identify issues in the test suite. For example, code coverage analysis for test code (not production code!) quite effectively reveals test files that are never run for some reason and tests cases that always bail out. Even doing such relatively simple thing as counting how many of assertions are reached and how many are missed will help to find permanently disabled tests.

Let’s play a game

The next time you find yourself sitting and waiting for a long-running test suite to complete, do one of the following things.

  1. Grep your test suite for early returns: they usually correlate with bail outs. In Python, look for regular expression of return$; in C/C++, look for void functions having return;
  2. Find run history for individual test cases. Are there any that have never failed? Are there any that fail disproportionally often?
  3. Sort log files coming from your test cases by size. Open the largest log file. How much of the output you see do you dare to remove?

References

Kent Beck is always worth reading: https://medium.com/@kentbeck_7670/test-desiderata-94150638a4b3


Written by Grigory Rechistov in Uncategorized on 05.10.2026. Tags: tests, maintenance,


Copyright © 2026 Grigory Rechistov