Data classification and retention gives a 10 to 70 person company one clear picture of the sensitive data it holds: client records, personal data, and contracts. It maps where each lives, who can see it, and how long to keep it before deleting. You get a simple classification scheme, a retention schedule, and a light ongoing process. Counsel sets the periods for regulated data. Live in two to three weeks.
The problem
Most companies cannot say what sensitive data they are sitting on. Client files from a project that closed three years ago still live in a shared drive. Old contracts sit in an inbox nobody archives. Spreadsheets full of personal data (names, emails, and other personal identifiable information, or PII) get copied into a folder for one report and never deleted. No one decided to keep any of it. It just piled up, because deleting felt risky and no one owned the call.
Put a number on it. The Veritas Value of Data study, conducted by Vanson Bourne in 2019, found that on average over half (52 percent) of all data within organizations remains unclassified or untagged. For a 40-person firm, that means more than half of everything you store is data no one has labeled, described, or set a deletion date for. You cannot protect, or safely get rid of, what you have never looked at.
The cost stays invisible until it is not. Every extra year of forgotten client data is a bigger blast radius the day there is a breach, and more to explain when a client's own security review asks what you hold and why. Regulations like GDPR expect you to keep personal data only as long as you have a reason to. Sitting on years of it you forgot you had is both a liability and a compliance gap, and it surfaces at the worst possible time: during an incident, an audit, or a deal.
How the automation works
Find where the sensitive data actually lives.
The system scans the places a small company keeps files: drives, your CRM, email, contract folders, and shared docs. It looks for the patterns that signal sensitive content, such as personal data, client records, and signed agreements, and maps where each type sits and who can reach it.
Sort it into a simple scheme and propose a retention window.
Each store gets tagged against a short classification scheme you agree on, for example public, internal, confidential, and regulated. For each type it proposes how long to keep it before deleting, based on your rules and, for regulated data, the periods your counsel sets.
Turn it into a schedule and a light ongoing check.
You get a written retention schedule, a map of where sensitive data lives, and a recurring check that flags data past its keep-by date and new stores that need classifying. Deletion stays a reviewed decision, never a silent automatic one.
The pieces are proven: scanning drives and inboxes, pattern-matching for personal data and contracts, tagging against a scheme, and flagging data by age. The real work is the wiring. A one-time classification that is never maintained rots fast, so the ongoing check matters as much as the first pass. And auto-deletion is genuinely dangerous: delete something you were under a legal hold to keep, and you have turned a cleanup into a bigger problem. The hard part is agreeing the scheme, mapping where sensitive data actually lives, and setting retention that satisfies both privacy rules and legal-hold needs. That is what gets set up, tested, and handed over during implementation, with counsel in the loop on regulated data.
What this looks like in practice
Files had accumulated across drives and inboxes for six years, and no one could say what was in most of it.
- No classification scheme. More than half of the firm's stored files were untagged, matching the pattern most companies see.
- Client records from projects closed years ago still sat in an open shared drive, reachable by anyone in the company.
- No retention rule anywhere. Personal data was kept indefinitely by default, with no reason on file and no deletion date.
- A simple four-tier scheme (public, internal, confidential, regulated) applied across drives, the CRM, and contract folders, with sensitive stores mapped and access tightened.
- Old client records moved to restricted access, with a written retention schedule setting a keep-by date for each type.
- A recurring check now flags data past its date and any new store that needs classifying, so the map stays current instead of going stale.
Typical impact
Typical ranges for this pattern, not client claims. Your numbers get modeled in the audit.
Systems it connects
Plus most tools with an API. The audit maps your exact stack.
Who this fits
- You hold client data, personal data, or contracts, so what you keep and how long is a real exposure
- 10 or more employees, past the point where one person remembers what is stored where
- Years of files have piled up across drives, inboxes, and your CRM with no scheme and no deletion rule
- You want a schedule and a light ongoing process you can actually follow, not a one-time cleanup that goes stale in a month