Monday, August 24, 2026

Chasing the Cheese

Homelab · Networking · Troubleshooting

Chasing the Cheese


The first thing to do today is to chase the cheese, just as a mouse does in a maze.

After going in a circle over a solution which, at first, appeared rather simple.

I had not experienced any serious network problems for more than two months. One morning, while I was quietly having my breakfast, I noticed that the connection had gone down. I then finished my breakfast, had my coffee and went to look at the monitor.

All services were down.

Perfect.

The cheese was located somewhere in the maze.

To begin with, I believed that the internet connection had simply gone down. I looked at the admin computer and found that it wasn't connected to the network. I then restarted the OPNsense server, hoping that it was one of those issues which can be fixed by a restart and some patience.

It didn't work. At that moment I understood that the problem was likely more complicated than it appeared.

There is an emergency procedure specifically for cases like this. I carried out the bypass and, for now, kept the admin computer connected directly to the ISP. The concept was straightforward: ensure that at least the basic functions continued while I worked on discovering what was going on with the main network.

It took about five or six minutes for the bypass to stabilise, but even then the connection was still not behaving properly. There's where the labyrinth started.

Wrong Turn One

The First Problem.

To begin troubleshooting I connected directly to the server.

First mistake.

When I needed to be working on the LAN, I plugged the cable into the WAN port since I was looking for a DHCP problem and so I began to check what was going on in OPNsense.

The console began to provide me with clues, even though none of them appeared to point directly to the issue.

DHCP was down.

When I checked the processes, I didn't see either dhcpd or kea-dhcp4 running; the only one there was the WAN DHCP client.

I used some commands that I had remembered from earlier installations of pfSense.

clog.

It didn't exist.

I attempted the command service dhcpd onestart.

Nothing.

Then service kea-dhcp4 onestart.

Nothing.

Finally, service kea onestart.

The system began Kea.

However, the process stopped right away.

By that stage I began to consider the possibility that I was encountering a known bug in the version of OPNsense that I was using, one that might have something to do with a socket or with a previous process which had become blocked.

I was searching for a solution on the server.

The problem did not lie there.

The Turning Point

The Piece of Information That Changed Everything.

I attempted to log in using SSH.

Timeout.

Not a connection rejection.

A timeout.

I then looked at the network configuration from a different machine and found something interesting: my Wi-Fi was on a different subnet, namely 192.168.1.x, since it was connected directly to the ISP's router through the bypass.

But then I asked a much more basic question:

Did ping ever work?

The answer was no.

Not even directly connected.

The fact about that matter completely altered the diagnosis.

I looked at the server's physical interfaces.

ifconfig igb0

And there it was:

status: no carrier.

There was never any physical connection.

I was just as puzzled by wg0, but it was a virtual interface and had no connection with the physical port.

So I checked igb1.

Here was the problem.

The cable was plugged into the incorrect physical port on the dual NIC.

I repositioned the cable to igb0.

The interface became active.

Ping.

Working.

I reconnected the WAN and checked the connection using 10 out of 10 packets with 8.8.8.8.

Problem solved.

Or so I thought.

The cheese was still dripping.

Wrong Turn Two

The Second Problem.

The network was restored.

It then began to drop every ten or fifteen minutes.

The switch would freeze, and the only fix was to turn it off and then on again.

Once.

Twice.

Three times.

Always the same.

The switch in question was a Linksys SE3008v2, a simple unmanaged switch. It had no interface available, no logs that could be examined, and no straightforward method of finding out what it was doing.

Therefore I began to look at the topology.

I had two switches.

The first held the connected critical infrastructure — OPNsense, the NUC, the NAS, and other equipment.

The second held the office computer and several access points.

You can see a map of the full network layout in this earlier post.

The first idea was simple: what would happen if I take out the first switch?

I could simplify the network temporarily and carry on with just one switch until I received a replacement.

However, there was a problem.

I couldn't remove OPNsense.

I then had to discover what was causing the switch to overload.

Then I saw it.

The lights on the switch were flashing rapidly.

All of them.

This wasn't normal traffic. It seemed as if a tiny green disco had been set up in my office.

All the evidence indicated that there was a broadcast storm.

The Hunt

Isolate the Problem.

I cut the uplink that connected Switch 1 to Switch 2.

Switch 1 settled down immediately.

The loop lived downstream. Now I needed to find exactly where.

I disconnected everything from Switch 2. Only the uplink stayed in place.

Sixty seconds. Nothing.

Ninety seconds. Still nothing.

Switch 2 itself wasn't the problem.

One by one, I started bringing devices back online.

Archer first. Twenty minutes. Clean. It was even serving internet on its own.

Then Linksys02101.

Within minutes, the uplink and Archer's port both lit up — frantic, sustained, the same green disco from before.

Linksys itself stayed almost calm.

I disconnected it.

The storm stopped in seconds.

The cause had a name now: Linksys02101.

I checked its configuration, expecting to find Wireless Bridge mode — the classic mistake, an access point talking to another one over the air while it's still wired into the same switch. Two paths. One loop.

It wasn't that.

Plain Bridge Mode. Correctly set. Exactly as it should have been.

I never found out why.

The Fix

The Solution.

Getting the Linksys off the shared network was the goal — not necessarily figuring out what was actually wrong inside it. I set it up in full router mode: its own subnet, its own DHCP running locally instead of trusting OPNsense's, a unique SSID separate from the shared mesh name, hung off the Tenda instead of the main switch. With NAT in between, it wouldn't matter what mechanism had been causing the loop — there'd be no path left for it to travel back into the rest of the network.

It froze mid-configuration. Not later, not under load — while I was still setting it up.

That settled it. Whatever was actually wrong with that router, reconfiguring around it wasn't going to fix it. It came out of service entirely, replaced for now with a simple wireless range extender — no Ethernet port at all, nothing for it to loop through even if it tried. A real mesh system is already planned for October. Until then, this holds.

The Point

A whole day, and I still don't know why.

Problems like this eat the whole day. Not almost — the whole day; this one didn't wrap up until 6 PM.

It's the kind of thing no tutorial prepares you for. I don't know if formal education ever really does, either. But I trust my contingency plans enough now that leaving the property without internet — even partially — doesn't stress me out anymore, as long as the services that actually matter keep running. Someone telling me "we have no internet" doesn't rattle me. I just tell them to wait, that I'm working on it.

I've gotten good at following a process when something breaks. But some hardware problems look simple at first and turn out to be genuinely hard to pin down — this was one of them. One of the access points, the newest one in the mesh, was quietly bottlenecking the entire network. Incredible, honestly. These are things I don't see coming. Halfway through the day, I was convinced the real problem was the OPNsense server's motherboard.

Realizing that failures like this can happen at the level of any single component means thinking about newer hardware eventually — not urgent, just something to keep moving toward. The goal isn't making sure this never happens again. It's making sure that when it does, the property keeps running while I'm away.

This post contains Amazon affiliate links. If you purchase through them, I may earn a small commission at no extra cost to you. Every product listed here is something I personally own and use daily.

0 comments:

Post a Comment