Blog

Why I Built Emlyzer

Why I Built Emlyzer

Why I Built Emlyzer: Searching Log Files Wasn't Enough

The original idea for Emlyzer came to me while running.

At the time, I was already spending a significant amount of my work analyzing log files. They came from very different environments: customer systems, embedded systems, and even the process of building Linux derivatives.

The technologies were different, but the basic task was often surprisingly similar:

Something wasn't working as expected, and somewhere inside a large text log was the information needed to understand why.

Searching Wasn't the Problem

I mainly used Notepad++ for working with plain-text logs. Depending on the system, I also worked with specialized tools such as Gurock's SmartInspect Console and Microsoft's Service Trace Viewer.

These tools could search logs perfectly well.

But over time, I realized that searching itself wasn't really the problem.

Imagine finding an interesting identifier in a log file:

DeviceId=123

Finding the next occurrence is easy.

But that's often not what I actually want to know.

I want to know:

  • Where does this ID occur throughout the entire log?
  • Does it occur regularly or only in certain parts?
  • What other events happen around it?
  • Is there another error or message that follows the same pattern?
  • Do two seemingly unrelated events actually occur in the same regions of the log?

A normal text search can answer where the next match is.

It doesn't necessarily help me understand the pattern.

And with large, complex logs, that distinction becomes important.

What If I Could Just Mark What Matters?

That was the basic idea that eventually became Emlyzer.

Instead of repeatedly searching for something, I wanted to be able to simply select an interesting piece of text and say:

This matters. Show me where it occurs.

That selection becomes what Emlyzer calls a token.

Once a token is defined, its occurrences are highlighted throughout the log. More importantly, their positions become visible across the entire file.

Instead of navigating from search result to search result, I can get an overview of the pattern.

Then I can select something else.

An error code. A thread ID. A device identifier. A function name. A state change. A communication message.

Each additional token adds another piece to the picture.

This was the core idea of Emlyzer from the very beginning.

From Finding Text to Recognizing Patterns

There is a subtle but important difference between searching a log and analyzing it.

Searching asks:

“Where is this text?”

Analysis asks:

“What is happening here?”

When investigating a difficult problem, I rarely care about just one line. I'm trying to reconstruct a sequence of events.

Perhaps a device disconnects. Before every disconnect, a certain warning appears. Somewhere else, a state changes. A few lines later, the system reconnects.

Individually, these messages might not look particularly interesting.

Once their occurrences are highlighted across the log, relationships can become much easier to spot.

This is the workflow I wanted Emlyzer to support:

Select. Highlight. Observe. Compare. Understand.

Without having to repeatedly configure searches or mentally remember where previous matches occurred.

Analysis Is Knowledge – So Why Throw It Away?

Once I started working with tokens, another question became obvious.

Why should I have to define the same tokens again the next time I analyze a similar log?

And even more importantly: why should someone else have to rediscover the same things?

That's why Emlyzer tokens can be stored in token files.

A token file isn't just a collection of colors and search terms. In many cases, it represents knowledge about a system.

Someone who understands a particular application might know that certain messages, IDs, states, warnings and errors are important when diagnosing a specific problem.

That knowledge can be captured once and reused.

The next time I receive a similar log, I can load the token file and immediately start with the relevant highlights.

And I can give the same token file to a colleague, customer or support engineer.

Instead of explaining:

“Search for this message first. Then look for this ID. Highlight this state as well. And if you see this error, check whether that other message occurs nearby…”

I can provide the analysis setup itself.

The next person starts with the same visual information I had.

I think this is one of the more interesting aspects of Emlyzer: log-analysis knowledge becomes reusable and shareable.

Filtering Without Throwing Away the Interesting Part

Of course, highlighting alone isn't always enough.

Sometimes a log contains so much information that filtering is necessary.

But filtering creates another problem I've encountered many times: the matching line often isn't the most important line.

The cause might be five lines before it.

The consequence might be three lines after it.

If a filter removes everything except the matching lines, exactly that context disappears.

That's why Emlyzer can include a configurable number of predecessor and successor lines for filtered results.

For example, instead of showing only:

Connection lost

I might want to see five lines before and ten lines after every occurrence.

This reduces the amount of information without completely removing the sequence of events I'm trying to understand.

One Problem, Multiple Log Files

There is another problem that becomes especially relevant when debugging systems consisting of multiple components.

An issue rarely respects the boundaries of a single application.

Imagine an embedded device communicating with a PC application. The device writes its own log file, while the PC software writes another one.

An error observed on the PC might actually have been caused by something that happened on the embedded device a few milliseconds earlier.

Looking at the two files separately makes reconstructing that sequence unnecessarily difficult.

With the upcoming version of Emlyzer, multiple text log files can be opened together. Their log lines are ordered chronologically and combined into a single view.

Instead of switching between files and manually comparing timestamps, you can follow the sequence of events across components as one continuous timeline.

For example:

Embedded device → communication → PC application → response → embedded device

becomes visible in the order in which it actually happened.

The same concept also applies to systems consisting of several applications, services or devices, as long as their logs provide timestamps that can be correlated.

And because the combined result behaves like a log inside Emlyzer, the same token-based analysis can be applied across all of those sources.

A token may reveal an event from one component while another token highlights the corresponding reaction in another.

This changes the question once again.

Instead of asking:

“What happened in this log file?”

you can start asking:

“What happened in the system?”

Emlyzer Isn't Tied to One Kind of Log

My own work has involved very different kinds of logs: customer systems, embedded systems and Linux environments.

That's also why I don't see Emlyzer as a tool for one specific technology or development stack.

At its core, it works with text.

If a system produces complex textual log files and someone needs to understand what happened inside them, that's the problem Emlyzer is intended to solve.

That might be a software developer investigating a bug, a support engineer analyzing a customer issue, an embedded developer debugging a device, or a system engineer trying to understand a sequence of events across several components.

The source of the log isn't particularly important.

The challenge is the same:

Turn a large amount of textual information into something a human can understand efficiently.

Search Less. See More.

Looking back, Emlyzer didn't start because I needed another text editor.

It started because searching wasn't enough.

I wanted to mark something interesting and immediately understand where it occurred throughout a log. Then mark something else and see how the two relate.

Today, that idea extends beyond a single file. When a problem involves multiple components, their logs can become part of the same timeline and the same analysis.

That simple idea – which came to me while running – became the foundation of Emlyzer.

The product has grown since then, but the principle is still the same:

Don't just search through logs. Make their patterns – and the sequence of events behind them – visible.

If you regularly analyze complex text logs and this problem sounds familiar, you can learn more about Emlyzer on the website or see how it works in the documentation.

An unhandled error has occurred. Reload 🗙