Protect Your Tube Server: Disk Space, Databases and the 18-Hour Outage We Caused
Tube sites move a lot of video through their servers: downloads, conversions, clips, thumbnails, backups. All of it lands on the same disk as your database. When that disk fills, the database can’t write and the site goes down. It happened to us, and this guide is the honest version of how, with the safeguards we put in place the same day.
What happened
We added a feature that downloaded only a few selected videos out of large multi-video torrents. Two things went wrong at once:
- Pre-allocation. Our torrent client was set to reserve disk space for a download in advance. For a pack where we wanted 10 GB of files, it reserved space for the whole 36 GB pack anyway. After a restart it did it again.
- An orphaned download. We deleted a database row for one pack while the torrent client was still downloading it. With no record left, none of our limits applied to it, and it grew to 108 GB.
The disk hit 100% in the evening. PostgreSQL could no longer write, crashed, and kept restarting into recovery. The sites that depended on it returned errors, background jobs failed, the nightly backups failed, and some login tokens for third-party platforms were lost because their renewals couldn’t be saved. It was 18 hours before we found it.
The safeguards, in layers
No single rule protects you; each of these works even if the others fail.
1. Alert early, not at the edge
We used to alert at 25 GB free. That is too late on a server that can write tens of gigabytes an hour. Now we warn below 100 GB and raise a critical, emailed alert below 60 GB. Pick thresholds based on how fast your server can fill, not on a percentage.
2. Stop starting new work well before the danger zone
The downloader refuses to start new downloads below 80 GB free. Work already in progress continues, but nothing new piles on.
3. An automatic brake
Below 40 GB free, a check that runs every minute pauses every download at once and records a critical alert. It only resumes once there is comfortable room again. This is the layer that saves the database when everything else has gone wrong.
4. Keep a reserve you can delete instantly
A 20 GB placeholder file sits on the disk doing nothing. If the disk ever fills anyway, deleting it gives the database room to recover in seconds while you find the real cause. Write the one-line instruction next to it so anyone can use it.
5. Hard caps on anything big
- Disable pre-allocation for downloads you only partly want, or put such work on a separate disk.
- Cap every job by size before it starts, not after.
- Clean up temporary files on a schedule (we keep downloaded files only until they are uploaded, and short clip files for 72 hours).
6. Never orphan a process
If a database row controls an external process (a download, an encode, an upload), never delete the row first. Mark it finished or dead, stop the process, then delete its files. And check regularly for processes that no record controls; we now flag any untracked download over 2 GB every hour.
Separate what can fill from what must not
The strongest protection is putting downloads, renders and backups on a different volume from the database. If that volume fills, the site stays up. If you can’t, the layers above are the minimum.
After an outage: the recovery checklist
- Free space first (the reserve file, then the cause).
- Wait for the database to finish recovery and confirm it accepts connections.
- Restart the application servers and background workers so they reconnect cleanly.
- Clear failed scheduled jobs so they run again on their next schedule.
- Re-run the backups that failed. Don’t wait for tonight.
- Check anything that stores credentials or tokens: renewals that failed during the outage may have invalidated logins.
- Write down what happened and which safeguard would have stopped it.
Back up off the same disk
Our nightly database dumps were on the same disk that filled, and they failed that night. Backups are only as good as the place they live. Keep at least one copy somewhere else.