Every visitor to Spluur looked like the same person.
Nothing crashed. No alert fired. The API kept answering requests like nothing was wrong. But when I looked at the IP address it believed each request came from, it was the same address every time, and that address wasn't a user.
That is my favourite kind of bug to write about, because it is silent. It also taught me more about my own platform than a month of things working.
Here are three of them.
Bug 1: Everyone had the same address
A request to an app on Spluur doesn't go straight to the app. It goes through a reverse proxy first:
Browser → Traefik → Fastify APIFrom the API's point of view, the network connection comes from Traefik, not from the browser. So the address on the socket is Traefik's address on the Docker network. Every user, same address.
Traefik does pass the real client address along, in a header called X-Forwarded-For. But the API can't just believe that header. Anyone can send a fake one. Fastify only reads it when you tell it which proxies to trust, through the trustProxy option.
I had that setting wrong. So req.ip quietly fell back to the proxy's address, and the real one was never used.
The fix is one line, but the choice inside it matters:
import Fastify from "fastify";
// Trust exactly one hop: Traefik, directly in front of the API.
const app = Fastify({ trustProxy: 1 });Setting trustProxy: true also makes the error go away, and plenty of tutorials suggest it. But it tells the app to trust every hop in the chain. Depending on how your proxy handles incoming headers, that can let a client pick their own IP address.
Anything that depends on a visitor's IP breaks quietly when this is wrong. Rate limits treat the whole internet as one client. Logs can't tell two people apart. Abuse detection goes blind.
A trick worth keeping: a debug route that prints three values side by side.
app.get("/debug/ip", async (req) => ({
ip: req.ip,
forwarded: req.headers["x-forwarded-for"],
socket: req.socket.remoteAddress,
}));If ip and socket match while forwarded shows a different address, the proxy is passing the truth along and the app is ignoring it. That takes ten seconds to see, once you know to look.
Bug 2: The domain verified, and nothing routed
Spluur lets you attach your own domain to a deployment. The flow works. You add the domain, the platform verifies you own it, and it should start serving your app.
Verification succeeded. Traffic didn't follow.
Traefik can learn its routes from different places. Two of them are Docker labels on containers, and YAML files on disk. Spluur's real routing comes from the YAML files that the deploy worker writes. But the custom-domain code, after verifying a domain, updated the labels on the container.
The labels were correct. Traefik just wasn't reading them.
A route in the file-based setup looks like this (names are illustrative):
http:
routers:
my-app-custom:
rule: Host(`app.example.com`)
service: my-app
tls:
certResolver: letsencrypt
services:
my-app:
loadBalancer:
servers:
- url: http://my-app:3000I fixed it by sending custom-domain changes through the same file writer the deploy worker already uses. One writer, one place Traefik looks.
The lesson was not really about Traefik. When a system has two ways of being configured, the bug lives in the gap between them. The code that wrote the labels never failed. It succeeded at writing to a place nobody was reading, and that is worse than failing, because nothing tells you.
Now I ask one question before I trust a configuration change: what does the running system actually read?
Bug 3: A migration history that no longer told the truth
Early in a project it is tempting to use prisma db push. You change the schema, run one command, and the database matches. It is fast, and for prototyping it is great.
It also leaves no history. There is no migration file saying what changed or when.
I kept using it for too long. By the time I switched to proper migrations, the migration history no longer matched what the database actually looked like. Prisma could see the mismatch, and the usual advice for a mismatch like that is to reset the database. That is fine on a laptop and impossible when real data lives there.
The cleanup is tedious. You compare what the database really contains against what the migrations claim, then mark the migrations that are already applied so the history lines up again. Prisma's baselining flow exists for exactly this.
It changed one rule for me: migrations are the source of truth from day one, even when the project is small and nobody else is touching the database. Speed from db push is a loan, and the interest shows up later.
The pattern
All three bugs have the same shape. In each one, I trusted a layer I hadn't looked at closely:
- The proxy would hand over the real IP.
- The routing layer would read the thing I was updating.
- The migration history would stay in sync by itself.
Owning the whole stack means each of those assumptions belongs to you. There is no vendor support page to blame, and no platform team to ask. That is the cost. It is also why I keep building the platform myself: I now know what every layer does, because every layer has had a turn at being wrong.