Some of my test smells

I’ve been using test smells from Gerard Meszaros’ xUnit Patterns when reviewing pull requests for a while now. It is a good addition to the regular code smells (see e.g. Marcel Jerzyk’s catalog, although I am usually using a different set).

While code smells are usually applied to the production code, the test smells try to catch specifics and issues pertinent to automatic test cases.

Over the years, I’ve discovered and name-given a few test smells on my own. Often, my code smells are specializations of a pre-existing code or test smell, but accompanied with additional context to make it worthwhile to have a separate name for them.

Each entry below first describes symptoms of a smell, then it explains why it is a bad sign, and finally, gives suggestions about how it could be addressed.

Chatty test

A test case generates a huge number of artifacts (e.g., log messages) for all runs, even when it passes.

When it fails, the actual error message is lost among all the irrelevant stuff (i.e., it is a subset of the Obscure Test smell). Copious output that needs to be written to disk or sent over network may also slow down this and other concurrently run tests (contributing to the Slow Test smell).

With the idea that logging is rethrowing an exception to a human, excessive logging is also an indicator that the test is not really designed to be automatic but smells as Manual Intervention-one. Automatic tests strictly return a binary result (fail/pass), not a wall of text for a human to process and judge.

How to avoid

Keep it to the Unix principles for command line utilities. A successful execution should generate no diagnostic output, just a return code indicating success. A failing test should be allowed to be verbose in its description of the detected problem, but every character it spews out should be relevant to a human looking for the context of a fix.

Do not enable verbose logging in test cases by default, unless a test cases exercises the logging operation itself.

Test cries “wolf”

This is a variation of Chatty Test. A log file for a passing test contains a lot of text lines with words “warning”, “error”, “fatal” or other markers with similar meaning of an anomaly happening. This is typical for Long (== Slow) and Obscure tests.

Such alarming lines often indicate ignored or corrected (but still logged) problems. Sometimes they are just a result of the test or the system-under-test being overly dramatic. When an actual test failure happens and it gets logged, its presence is lost in the sea of similarly looking alarms.

How to avoid

Having warnings is unprofessional. They must be fixed or at least kept to minimum, especially if they are coming from components that do not play any role in the specific test. As before, enabling verbose logging on everything for test runs does not automatically make them better. It should be possible and easy to activate helpful tracing for manual debugging, but it should not be a default.

Log messages that show up in test logs will certainly sneak into production environments, where they will be even more misleading. The customer observing the logs will be confused at first, at will start ignoring them soon after.

Catch-all assertion

The try-catch-all statements in production code is a known code smell. The usual recommendation is to explicitly write catch sections for expected problems, and let the runtime deal with all the unexpected ones.

The same rule applies to assertions in tests.

Tests must be specific, their messages must be valuable. Having a catch-all leads to printing a generic or misleading error message, worsens failure localization and causes the test case to have Frequent Debugging smell. This is again an indication that an automatic test isn’t; rather, it is a manual one since it often invites a human.

Suppose an exception caught by a catch-all section is ignored in the test, because other expected classes of exceptions were deemed harmless. At some point, a previously-never-seen exception will be thrown inside the try block, but it will be silently swallowed up by the catch-all; this is not what was meant to happen here originally! It will lead to a test case that cannot fail, i.e., the Buggy Test smell.

How to avoid

  • Don’t use catch-all; catch specific narrow exception classes instead. If something truly unexpected is thrown in the test, it will be handled on a higher level, that is why the exceptions are structured.
  • Don’t use equivalents to catch-all. For example, waiting for a timeout and then reporting “something is wrong, go figure” should not be a normal situation in your tests, and it should not be the only assertion in the test.

Speaking of number of assertions/expectations in a single test case…

Test case has no assertions

A case without assertions ensures only that no exceptions are thrown. It ensures very little about the behavior of the code under test.

This could be treated as a variation of the Catch-All Assertion smell. Not expecting anything in a test to happen is as if we expect its failure action to throw an unspecified exception, then a catch-all behavior of the outer test runner framework to catch it and to report it as a generic error. Alternatively, we do not expect the test to ever fail, in which case there is no point to run it at all.

Such a test fails to localize the error. It forces a human to dig in and debug the case every time it fails, tying the smelly behavior back to the Debugging and Manual Intervention smells.

Sometimes a test’s expectation is that a certain piece of production code will not be called during the test scenario. This is usually represented by a test spy double object, which unconditionally fails whenever it is called. So, if it is not called, the test succeeds after its last statement is over. If the spy gets called, that will be the failure reaction.

Such design of a test still smells with Obscure test, unless it is obvious to a reader what the intention here is. It is not trivial to test for things to not happen, but it is out of scope for this discussion.

See a similar rule of SonarQube.

How to avoid

Be explicit of what kinds of erroneous situations or exceptions the test case monitors for. Put a try{ ... } catch SpyWasCalledException around the Act phase of your test, where SpyWasCalledException indicates that the spy object was indeed called. In the catch block, rethrow explicitly with a message that explains the “an interaction should not have happened” situation.

Before, it was Arrange-Act, with the Assert phase existing but misplaced. Afterwards, the linear Arrange-Act-Assert structure that good tests should have is restored.

Test case logic contains an early conditional bailout

This is another face of a code smell for Robert C Martin’s book “An Ignored Test Is a Question about an Ambiguity”. See also the rule about separating the logic from policy of applying this logic: https://wiki.c2.com/?MechanismRichPolicyFree. It is also about not mixing up decisions of different levels of abstraction.

Consider this example of a Python test for a hypothetical optional Foobar subsystem:

def test_when_foobar_is_active_then_barbaz_is_returned(fixture):
    if not subsystem_is_present(fixture, "foobar"):
        return
    # Actual testing of a scenario involving Foobar follows
    # ...

The first if-block provides an early bailout for a situation when a particular system-under-test does not include the feature we were about to exercise.

Here, we have a test case that gives strictly less than one bit of information (but more than zero bits) after its completion.

Only when such a test case fails, it offers useful information about Foobar implementation. However, when it passes, there is no way to tell whether it actually has stressed the subsystem at all or has simply bailed.

How to avoid

Do not allow tests to decide for themselves whether or not they should be skipped. If such a decision must be taken, it should be taken by the test runner application, whose job is to discover test cases, schedule them and run them.

Even then, the number of skipped test cases in the test suite that are not regularly run even once a while should be tracked and dealt with. After all, it is the developer team who is the only customer of their test suite. If you do not run a test case, nobody does, and it can be deleted thus saving the team both compute time and maintenance effort.

Several test cases react together on a single change in production code

An ideal test suite contains test cases that provide 100% cumulative coverage, while simultaneously offering minimal overlap of coverage between the test cases it is comprised of. With full coverage, any regression is reported by at least one test case. With no overlap, any regression in production code is reported by exactly one test case.

“At least one test reacts” is certainly much better than “zero tests react on a regression”, but it is not as good as “exactly one test reacts”.

Having the same or tightly coupled expectations in many test cases means that all of them will start failing at the same time if corresponding production behavior gets changed. The test cases become coupled through that expectation. It worsens localization of eventually reported regressions.

How to avoid

Do not add assertions to tests “just in case”. If you need to test an additional piece of state or a fact about your system-under-test, consider creating a new separate test case for it, instead of burdening existing test cases with more responsibilities. If you still need/want to add a new assertion to an existing test case, do not add it to more than one test at a time.

Appendix: overview of test smells from xUnit Patterns

  • Code Smells
    • Obscure Test
    • Conditional Test Logic
    • Hard-to-Test Code
    • Test Code Duplication
    • Test Logic in Production
  • Behavior Smells
    • Assertion Roulette
    • Erratic Test
    • Fragile Test
    • Frequent Debugging
    • Manual Intervention
    • Slow Tests
      • Slow Component Usage
      • Asynchronous Test
      • Too Many Tests
  • Project Smells
    • Buggy Tests
    • Developers Not Writing Tests
    • High Test Maintenance Cost
    • Production Bugs

Written by Grigory Rechistov in Uncategorized on 03.10.2026. Tags: tests, smells,


Copyright © 2026 Grigory Rechistov