Data classification that people actually use: a practical UK guide
A step-by-step guide to designing, labelling, automating and measuring a data classification scheme that survives contact with real users.
Classification schemes fail when they ask people to make fine-grained judgements dozens of times a day. If the scheme is not obvious at the moment of creating a document, it will be ignored or applied at random. This guide sets out how we design classification schemes that people use without thinking about them, and how to prove that they are working.
What data classification is for
Classification is not an end in itself. It exists so that a control can be applied automatically: encryption, restricted sharing, retention, watermarking, data loss prevention, or blocking uploads to unapproved destinations. If a label triggers nothing, it is metadata theatre.
In UK organisations three drivers usually sit behind the work:
- UK GDPR and the Data Protection Act 2018, where you must know where personal and special category data lives before you can defend how you handle it.
- Contractual and supply chain obligations, including ISO 27001 Annex A controls on information classification, labelling and handling.
- Incident containment. When something goes wrong, the first question is always "what data was in scope?" — and classification is how you answer it in hours rather than weeks.
Step one: keep the scheme small
Three or four levels are enough for most organisations: public, internal, confidential and, where genuinely needed, restricted. Each level must have a plain-language definition and a short list of concrete examples from your own business. If two levels are hard to tell apart, merge them.
A workable four-level scheme looks like this:
- Public. Already published, or approved for publication. Marketing material, published accounts, job adverts.
- Internal. Default for day-to-day work. Team plans, internal process notes, most email.
- Confidential. Harm if disclosed outside a defined group. Customer records, commercial terms, employee personal data, security findings.
- Restricted. Serious harm to the organisation or to individuals. Special category personal data, M&A material, credentials, source code for core products.
Write the definitions in the language of harm, not in the language of policy. "Would this embarrass a named customer, breach a contract or attract a regulator's attention?" is a better test than an abstract sensitivity rating.
Step two: decide where the label is applied
The moment of labelling determines adoption. There are three practical patterns, and most organisations need all three:
- Automatic by location. Everything in a given SharePoint site, S3 bucket or database schema inherits a label. This is the cheapest coverage you will ever buy.
- Automatic by content. Pattern matching for card numbers, NI numbers, health data or known document templates. Accept that it will be imperfect and tune it over a few months.
- Manual with a default. Users see a pre-selected label and can change it with a reason. Never present a blank field.
Where the tooling supports it — Microsoft Purview sensitivity labels are the common case in UK enterprises — configure the default and the override path before you announce the scheme. A rollout where the first experience is a compulsory blank dropdown will not recover.
Step three: start where the risk is
Do not attempt to classify everything. Identify the handful of data sets that would cause real harm if exposed — customer records, source code, commercial terms, personal data of employees — and get those right first. Broad coverage can follow once the pattern is established.
A useful sequence for the first ninety days:
- Weeks one to three: agree the scheme, the definitions and the owner of each. Map the top ten data sets by harm.
- Weeks four to eight: apply location-based labelling to those data sets and connect one control per label. Encryption and external sharing restrictions are usually first.
- Weeks nine to twelve: turn on content-based detection in report-only mode, review the false positives with the data owners, then enforce.
This is the same asset-visibility discipline described in foundations first — you cannot classify what nobody has inventoried.
Step four: connect labels to handling rules
Publish a one-page handling table that says, for each level, what is allowed for email, external sharing, removable media, printing, storage location, retention and disposal. One page. If the handling guidance runs to twenty pages nobody will read it, and the helpdesk will invent its own answers.
Make the rules enforceable by technology wherever you can. A rule that depends entirely on user goodwill is a training obligation, not a control.
Step five: measure adoption, not policy
Track the proportion of new documents labelled, the volume of overrides, the top override reasons and the incidents prevented or contained faster. If adoption is poor, the scheme is too complex, not the users. Simplify and try again.
Report a small number of these measures to your governance forum every quarter alongside the rest of your governance, risk and compliance reporting, so classification is treated as an operating capability rather than a project that ended.
Common failure modes
- Too many levels, with sub-categories and caveats. Users pick at random and the data becomes noise.
- Labels with no consequence. Nothing happens when a document is marked restricted, so nobody bothers.
- Big-bang rollout across the whole estate on one date, with no report-only phase and no tuning window.
- No owner. Classification needs a named accountable owner in the business, not only in security.
- Retention forgotten. Classification without a disposal rule quietly increases the amount of sensitive data you hold.
Where to go next
If you are starting from a blank page, a short discovery exercise usually pays for itself: where sensitive data actually lives, which controls already exist, and what can be automated with the licences you already own. That is the core of our organisational data security work, and it often runs alongside a health check and gap analysis so the classification effort is prioritised against everything else competing for the same budget.
Read more of our writing on data and governance, or talk to us about your classification programme.