Data Engineering

Data Retention Tier Calculator

Prices a hot, warm, cold and archive retention policy tier by tier.

Loading the tool…

Processing happens locally in your browser. What you paste or load is processed by this page and is not uploaded to a server. Nothing is stored unless you use a control that says it stores something, and you can clear anything this site has kept from the privacy page.

How to use this tool

  1. Set rows per day, average row size and the compression ratio.
  2. Give each tier the retention you were planning: hot, warm, cold and archive. Leave a tier out to drop it.
  3. Replace the default unit prices with the ones on your bill — the defaults are typical cloud object-storage rates, not yours.
  4. Select Cost the policy, and check the per-month column tier by tier before deciding where to cut.

What retention policy calculator does

Retention policies are usually written in months and paid for in gigabyte-months, and the two rarely get compared. The result is a policy that reads sensibly — thirty days hot, a year warm, seven years archived — attached to a bill nobody has broken down by tier.

This puts the volume and the price side by side for each tier. The finding is often counterintuitive: because hot storage can cost thirty times what archive storage costs, taking a fortnight off the hot tier commonly saves more than deleting several years of archive, and the compression ratio matters more than either. Default unit prices are typical cloud object-storage rates and are there to be replaced with yours.

Frequently asked questions

They are typical published rates for cloud block, standard object, infrequent-access and archive storage, rounded. They are there so the tool produces something on first run. Replace them with the rates on your own bill before quoting any figure to anyone.

Because the price gap between tiers is enormous — hot storage can cost two hundred times what archive costs per gigabyte. A week of hot data is often worth more in savings than several years of archive, which is the opposite of where retention discussions usually start.

No, and for archive tiers that omission matters. Archive storage is cheap to hold and expensive and slow to read, so a tier you actually query is not as cheap as this makes it look. Storage cost alone is only half the picture for cold and archive.

They add up. Each retention figure is the time data spends in that tier before moving on, so a row is in exactly one tier at a time and the total is the full lifetime of the data.