Data Quality

Data Quality Scorecard

One score across six weighted dimensions, with what each one measured and what pulled it down.

Loading the tool…

Processing happens locally in your browser. What you paste or load is processed by this page and is not uploaded to a server. Nothing is stored unless you use a control that says it stores something, and you can clear anything this site has kept from the privacy page.

How to use this tool

  1. Paste the dataset.
  2. Select Score.
  3. Read the dimension that scored lowest, not the total.
  4. Follow it into the tool that covers that dimension in detail — the related links below are ordered for exactly that.

What quality scorecard does

A single quality number is useful for one thing only: deciding which dataset to look at next when you have forty of them. It is useless on its own, because nobody can act on 74. So every dimension here is reported with what it measured and what dragged it down, and the weights are printed alongside the score rather than buried, which is what makes two files comparable.

The six dimensions are completeness, uniqueness, validity, consistency, structure and privacy exposure. Privacy is a deduction rather than a measurement, and it is floored: a dataset full of customer records is not a bad dataset, but the handling requirement it carries deserves to be visible in the headline number rather than discovered later. Read the dimensions, not the total — the total exists to tell you which dimension to read.

Frequently asked questions

Completeness at 30%, validity at 20%, uniqueness and consistency at 15% each, and structure and privacy exposure at 10% each. The weights are fixed and printed with the result, which is what makes the score comparable between two files rather than a number that means whatever the tool felt like that day.

Not because the data is wrong. A dataset full of customer records is a perfectly good dataset — it simply carries a handling requirement, and this keeps that requirement visible in the headline number instead of leaving it to be discovered later. The dimension is floored so it cannot dominate the total.

The CSV profiler answers "what is in this file" with a column-by-column table of types, ranges and distinct counts. This answers "is it fit to use" with a score and the reasons behind it. Most people want the profiler when exploring and this one when deciding.

Put the dimensions in the report and the total in the summary line. On its own the number cannot be acted on — nobody knows what to do about 74 — but it is a reliable way of ranking forty datasets so you know which one to open first.