4 min read

The Monero miner attack that finally made me tackle my todo list: egress control

On 11 September at 10:51, a 7 MB XMRig build appeared inside one of my pods, in a directory named to look like an X11 socket: /tmp/.ICEi-unix/javae. By 11:13 it had written a config naming a mining pool and a Monero wallet, and it was mining.

By 11:55 it was contained. Not because an alert fired — nothing alerted. It was contained because eleven minutes after it landed, and knowing nothing about it, I had committed a network policy that cut its route to the pool.

The rollout and the miner ran concurrently. Containment came before detection.
The rollout and the miner ran concurrently and how I accidentally discovered and contained it.

The morning

The cluster had 27 NetworkPolicies. Exactly one of them had an egress section. Around ninety pods sat on a flat network where any of them could reach anything — the NFS servers, SSH on every host, the hypervisor’s admin port, the Kubernetes API, each other’s databases, and the whole internet.

It was an ordinary maintenance day. Fleet update check at 07:14, a round of image bumps at 07:53, and then the thing I had been meaning to get to: writing egress policies. umami went in the first batch at 11:02, because it is internet-facing and sorts near the top of any list ordered by exposure. The batch reached the cluster at about 11:44.

The miner’s pool callbacks started failing.

I found it by accident, six minutes later. What surfaced was an egress error in the logs that looked like a false positive — the rollout had been producing those all morning, each one a legitimate app whose real dependency I had missed. I was working through them in order. This one’s destination turned out not to be a dependency at all.

By 11:55 umami was scoped to DNS and its own database, the pool was verified unreachable, and an IOC scan across every running pod came back with exactly one infection.

What the policy did not do

It did not stop the break-in.

The entry vector was a remote code execution flaw in the framework underneath the app. umami is built on Next.js, and v2.19.0 shipped an unpatched one: CVE-2025-66478, a CVSS 10.0 deserialization flaw in React Server Components that lets an unauthenticated attacker run code with a single crafted HTTP request. The app’s own logs contain the exploitation in plain sight:

Command failed: echo <base64> | base64 -d | sh

umami is internet-facing on purpose — its analytics collection endpoint is deliberately open, because a tracking beacon that requires authentication does not work. That surface was exposed before, during and after this.

What the policy removed was the payoff. A cryptominer that cannot reach a pool is a process burning CPU for nobody. The same block cut the second callback host, and it would equally have cut an attempt on the NFS exports, the hypervisor, or another application’s database.

The miner itself was crafty too. Digging into the logs + container turned up a watchdog: a base64-encoded script that kill -9s everything on the box except the miner and a small allowlist, to clear out competitors and respawn itself. Props where props are due.

The todo list

The reasoning is blast radius: given that something will eventually get in, how far does it travel? On a flat pod network the answer is “everywhere”. A compromised container does not need to break out to be useful — it already has a route to the storage tier and a route to the internet. Those two facts turn a foothold into a payload.

The break-in is identical in both cases. Only the reach differs.
The break-in is identical in both cases. Only the reach differs.

Every pod gets a policy naming only the peers it needs, and a namespace-wide default-deny closes the rest:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-egress
spec:
  podSelector: {}          # every pod in the namespace
  policyTypes: [Egress]    # …gains the Egress type, with no allow rules

NetworkPolicy is additive-allow, so that empty policy is the whole trick: once a pod has any egress policy attached, it reaches only what is explicitly permitted. A pod deployed tomorrow with no policy of its own gets nothing, DNS included, until someone writes one.

The order is what cost me. That capstone can only go on once every pod already has a policy of its own, or switching it on breaks whatever you missed. So it went last, after 161 per-pod policies had brought uncovered pods to zero. Its value is for the future: the next service I deploy and forget about fails closed instead of open.

The rest went in that afternoon — media apps, messaging stack, long tail, healers, data tier, twenty Postgres clusters, and at 15:43 the twenty-one CronJobs, which are the easiest thing in a cluster to forget because most of the time they are not running. The capstone landed at 15:50.

It then took my 1-2 weeks to find now failing (but actually valid) request within the services (now blocked by default) by scrolling through logs + fixing them.

Cleanup

The pod was rebuilt from its image — the miner lived in ephemeral /tmp and did not survive. The image went v2.19.0 → v2.20.2. Credentials that had been in the pod’s environment were treated as compromised and rotated: the app secret and the Postgres password. The open collection endpoint now carries a rate limit, 120 requests per minute per IP.

Sources