Blog

On the Importance of Logging and Testing

Yuriy Melnikov

A note on why it's crucial to verify the entire application workflow, not just the part you worked on.

There's a saying—"The shoemaker's children go barefoot." Today, that's somewhat about me. Considering I develop a service for the delivery, verification, and auditing of webhooks, the situation is doubly comical. But first, some background.

Today, a new release of Adal CLI v0.1.2 was launched, along with some server-side updates.

One of these changes was the modification of how webhooks are stored within Adal. For the end user, this migration went unnoticed, but for Adal's infrastructure, it significantly improved stability and performance.

One of the upcoming improvements will be limiting the storage of the request body. This could be important for certain user categories who forward data through webhooks that are, for example, confidential and prohibited from being stored by third parties.

After making the necessary code changes, I tested its functionality. Webhooks are received, processed, storage flags are set, everything reaches the CLI, and is sent to the final destination.

For final destinations, I have two addresses set: one points to a definitely working host, and the other to a definitely non-working one. This is quite understandable: to see how different delivery scenarios are handled.

But I overlooked one thing: what exactly does the final host receive? It was receiving requests without the original body and headers. Instead of the expected payload, only internal service data, not intended for the final recipient, was arriving.

The situation is unpleasant but not critical. Fortunately, the error did not affect the reception and storage of webhooks. Requests were correctly received, stored, and processed through the entire pipeline. The problem only arose at the payload formation stage for delivery.

The error was that when the request was saved, the storage type flag was set but not checked by the delivery service to the user. The service that forms the payload for user delivery received an unknown storage type and in such cases returned an empty result.

The bug itself was, of course, unpleasant. But I liked the system's behavior in this situation. The system preferred not to substitute default values and not to try to "guess" what was meant.

One could have made a condition "if we don't know the type, then consider it a default type," but I thought it better to consider "if we don't know what it is, we don't process it." Because "considering by default" is something from a black box, as you then have to guess "why did this request get this particular response."

This case reminded me once again of a simple thing: you need to test not a separate service or function, but the entire user scenario as a whole. You can ensure the webhook is received. You can ensure it's saved. You can ensure it's queued and delivered. But until you see what exactly the final endpoint received, the check isn't complete. That's why good logs and observability are as important as the business logic itself.

So I created another endpoint for observation and changed the host address that definitely works to the address of the new endpoint. Now I will clearly see what exactly arrives in the request. And once again, I realized: nothing finds bugs better than using your own product every day. Especially when this product is designed to solve your own tasks.