Home โ€บ Guides โ€บ Hosting & Speed

Protect Your Tube Server: Disk Space, Databases and the 18-Hour Outage We Caused

Tube sites move a lot of video through their servers: downloads, conversions, clips, thumbnails, backups. All of it lands on the same disk as your database. When that disk fills, the database can’t write and the site goes down. It happened to us, and this guide is the honest version of how, with the safeguards we put in place the same day.

What happened

We added a feature that downloaded only a few selected videos out of large multi-video torrents. Two things went wrong at once:

The disk hit 100% in the evening. PostgreSQL could no longer write, crashed, and kept restarting into recovery. The sites that depended on it returned errors, background jobs failed, the nightly backups failed, and some login tokens for third-party platforms were lost because their renewals couldn’t be saved. It was 18 hours before we found it.

The safeguards, in layers

No single rule protects you; each of these works even if the others fail.

1. Alert early, not at the edge

We used to alert at 25 GB free. That is too late on a server that can write tens of gigabytes an hour. Now we warn below 100 GB and raise a critical, emailed alert below 60 GB. Pick thresholds based on how fast your server can fill, not on a percentage.

2. Stop starting new work well before the danger zone

The downloader refuses to start new downloads below 80 GB free. Work already in progress continues, but nothing new piles on.

3. An automatic brake

Below 40 GB free, a check that runs every minute pauses every download at once and records a critical alert. It only resumes once there is comfortable room again. This is the layer that saves the database when everything else has gone wrong.

4. Keep a reserve you can delete instantly

A 20 GB placeholder file sits on the disk doing nothing. If the disk ever fills anyway, deleting it gives the database room to recover in seconds while you find the real cause. Write the one-line instruction next to it so anyone can use it.

5. Hard caps on anything big

6. Never orphan a process

If a database row controls an external process (a download, an encode, an upload), never delete the row first. Mark it finished or dead, stop the process, then delete its files. And check regularly for processes that no record controls; we now flag any untracked download over 2 GB every hour.

Separate what can fill from what must not

The strongest protection is putting downloads, renders and backups on a different volume from the database. If that volume fills, the site stays up. If you can’t, the layers above are the minimum.

After an outage: the recovery checklist

  1. Free space first (the reserve file, then the cause).
  2. Wait for the database to finish recovery and confirm it accepts connections.
  3. Restart the application servers and background workers so they reconnect cleanly.
  4. Clear failed scheduled jobs so they run again on their next schedule.
  5. Re-run the backups that failed. Don’t wait for tonight.
  6. Check anything that stores credentials or tokens: renewals that failed during the outage may have invalidated logins.
  7. Write down what happened and which safeguard would have stopped it.

Back up off the same disk

Our nightly database dumps were on the same disk that filled, and they failed that night. Backups are only as good as the place they live. Keep at least one copy somewhere else.

Spotted something outdated, or run a service we should look at? Email [email protected]. Listings are never paid.