Most sensitive data discovery software was designed to answer one question: what is sitting in our systems that shouldn't be? It crawls databases, file shares, and storage buckets, classifies what it finds, and files a report. That model works well for data at rest. It works badly for the thing a sales or marketing org actually ships all day — the deck an AI assistant exported to HTML twenty minutes ago and someone pasted into a sharing tool before the coffee finished brewing.
The same gap exists at the other end of the size curve, which is the part buyers' guides usually skip. A four-person startup has no warehouse to crawl, no endpoint agents, and no classification program — and it still ships pricing pages, investor updates, roadmap previews, and lead lists that an AI assistant generated an hour ago. Different scale, identical exposure. If you are a founder, the artifacts that keep you up at night are almost never in a database: they are the discount you extended to one customer, the names on the cap table, the two features you have promised three prospects, and the contact list your assistant helpfully enriched.
This is an evaluation guide rather than another taxonomy. It gives you a buyer's checklist for sensitive data discovery software — seven questions to ask a vendor, a worked example of the same page run through two detection methods, and an honest account of what no scanner can do for you. If you are a founder with no security team, read the checklist as a self-assessment instead: the questions are the same, and you are the reviewer.
Why the market is built for data at rest
Discovery products cluster into three at-rest categories, and the shape of each one tells you what it was designed to protect.
- Database and warehouse scanners. Connect to production stores, sample columns, classify them ("this column is email addresses"). Output is an inventory feeding a records-of-processing or retention obligation. Latency: hours to days, on a schedule. Can block: no.
- Endpoint DLP. Agents on laptops and mail gateways inspect files and messages as they move. Latency: real time on covered channels. Can block: yes, where the channel is covered. Weak at anything that happens inside a browser tab — a rep pasting HTML into a web app is, to most agent policies, a person typing in a text box.
- Cloud storage classifiers. Scan buckets, drives, and collaboration suites, often with permissions context. Latency: minutes to hours after the object lands. Can block: rarely — usually quarantine or alert after the exposure already happened.
Read that as a Venn diagram with a hole in it. A file generated in a chat window, never written to a managed drive, never attached to an email, and pasted directly into a browser crosses no boundary any of the three watches. It is a system-of-record problem solved for content that never enters a system of record.
Small companies land in the same hole from the opposite direction. All three at-rest categories assume an inventory worth crawling, an agent fleet worth managing, and a person whose job is to read the findings. A startup has none of those and still generates the same artifacts — so the only checkpoint that exists for a founder is the one at the moment of sharing.
The fourth category — pre-publish content scanning — inspects one specific artifact at the moment it is about to be shared, before a URL exists. Secret scanning in a code pipeline is the mature version of the idea. Applying it to generated documents (proposals, dashboards, lead lists, training material) is newer, because until recently those documents weren't being produced by machines at volume. It is strong at prevention and precision, and weak at coverage: it only sees what passes through it.
The buyer's checklist: seven questions for any discovery vendor
These questions apply whether you are evaluating an enterprise discovery suite, a lightweight PII scanner, or the scanning built into a sharing tool. They are ordered roughly by how often the answer decides the outcome. Founders without a formal review process should still walk them in order — answering them yourself in twenty minutes is cheaper than answering them later in a customer's security questionnaire.
1. What is the detection method — pattern or contextual?
Pattern matching (regular expressions and checksums) is deterministic: the same input produces the same result every time, it runs in milliseconds, it costs nothing per scan, and it is auditable — you can show a reviewer the rule that fired. It is excellent at structured identifiers and blind to meaning.
Contextual AI detection reads the document and reasons about it. It catches the sentence "we're holding the Q3 roadmap item back until the funding round closes" that no regex will ever match. It also costs tokens per scan, adds latency, and is non-deterministic — two runs can differ.
The correct answer is usually both, in that order: deterministic first because it is cheap and certain, contextual second because it is expensive and smart. A vendor offering only one should be able to say plainly which half of the problem it solves. If you are pre-revenue and watching every line item, start with the deterministic layer alone — it covers the credential-and-identifier case, which is the one that ends deals.
2. Does it read unstructured HTML, or only structured fields?
Ask for a demo against an actual generated page, not a CSV. Generated HTML hides sensitive data in places a field-oriented classifier never looks: data- attributes, inline JSON payloads powering a chart, HTML comments left by the model, hidden table columns styled out of view, a connection string in a commented-out script block. If the tool parses only visible text, it will pass a page that leaks in the markup. Founders shipping AI-generated pages hit this hardest, because nobody on a small team reads the markup before sending — see how credential leaks get into generated HTML and how to scan HTML for API keys before sharing.
3. Does it scan at the point of publish, or on a nightly crawl?
This is the question that separates prevention from forensics. A crawl tells you on Wednesday what went out on Tuesday. A publish-time scan can stop the link from existing. Both have a place — inventory work genuinely needs the crawl — but do not accept a scheduled scan as a control for outbound content. Ask specifically: can this tool interrupt the share, or only describe it afterward? How a secret scanning workflow should work walks the publish-time sequence in detail.
4. How does it handle false positives?
A scanner that flags every 16-digit number as a card number gets ignored within a week. Ask how findings are presented (blocking versus advisory), whether categories can be tuned per team, and whether a reviewer can dismiss a finding with a recorded reason. Suppression without an audit record is how a policy quietly erodes into a habit of clicking through. On a small team the erosion is faster, because the person clicking through is also the person who wrote the policy.
5. Where does the content go during the scan?
This is the question governance reviewers ask and vendors answer vaguely. If detection is pattern-based and local, the content never leaves the boundary you already approved. If it is AI-based, the content is sent to a model provider — and then it matters enormously whose account, whose data-processing terms, and whose retention policy applies. "We send it to our model provider" and "we send it to yours, under your existing agreement" are two different procurement conversations, and only one of them is short.
Founders selling upmarket should note that this exact question arrives in your own security questionnaires within a year or two. The subprocessor list you can answer with today is the one you will be defending then.
6. How are findings surfaced to a non-technical user?
The person who needs to act on a finding is usually a rep or a marketer, not a security engineer. In a startup it is the founder, between two customer calls. A JSON blob of regex match offsets is a finding a security team can use and a sales team will not. Look for: plain-language category names, the matched location highlighted in context so the author can fix the source, and a clear next action (redact, rescan, proceed). If remediation requires opening a ticket, the tool will be routed around.
7. What evidence does it leave for an audit?
Compliance value lives in the record, not the scan. Ask what is retained per scan: timestamp, who ran it, which categories fired, what the reviewer decided, and whether the record survives the artifact's deletion. Then ask how long that record is kept and who can export it. A scanner with no durable trail gives you a control you cannot evidence — see how an audit trail for shared content should be structured for what a defensible record contains.
The founder's version of this problem
Enterprise buyers worry about a regulated identifier escaping a governed system. Founders worry about something narrower and, in the early years, more expensive: the handful of facts that only exist in generated documents and nowhere else.
- Customer-specific pricing. The 40% first-year discount you gave one logo, restated verbatim in a proposal template your assistant reused for the next three prospects.
- Investor and cap-table detail. Names, allocations, and the SAFE terms sitting in an update deck that gets exported to HTML and shared with "just a couple of advisors."
- The unannounced roadmap. Dates, feature names, and the internal caveat about which of them are real, pasted into a customer-facing page because it was the same source document.
- PII in enriched lead lists. A CSV run through an enrichment step comes back with personal emails, direct dials, and home addresses your privacy policy never mentioned.
- Credentials in the markup. The analytics or database string the generator left in a comment, which is the only item on this list that can become someone else's incident.
None of these live in a warehouse. Every one of them lives in a document that was written by a model, reviewed at a glance, and turned into a URL. That is the whole argument for scanning at the moment of publish rather than on a schedule: it is the only point in the lifecycle where all five categories pass through one place.
Worked example: one page, two scanners
Take a single artifact — a quarterly pipeline dashboard generated from a CRM export, roughly 900 lines of HTML with an embedded chart, an account table, and a short written summary. It could equally be a seed-stage investor update with the same anatomy. Here is what each detection layer returns.
Deterministic pattern scan (runs on every link, every plan, zero tokens). Findings:
api_key— 1 match, in a commented-out<script>block containing the analytics connection string the generator left behind.email— 34 matches, contact addresses in the account table.phone— 11 matches, direct dials in the same table.person— 41 matches, contact names.financial— 2 matches, both false positives: order reference numbers that pass a card-number shape check.
That result is useful and incomplete. It found the leaked key — the single highest-severity item on the page — in milliseconds, and it produced a list a reviewer can walk. It also flagged two harmless strings, which is the cost of a method that reads shape rather than meaning.
Contextual AI scan with a customer-supplied model key (Teams and Enterprise). Running on top of the pattern pass, against your own Anthropic, OpenAI, or Google account, it adds:
- A line in the written summary reading "hold the renewal quote until their layoffs are announced" — internal commentary that no pattern can match and that nobody wants forwarded to the customer.
- A discount figure in the notes column that is 15 points below list, attached to a named account — the kind of number that resets every negotiation once it travels.
- A note that two accounts in the table are flagged in a column labeled
churn_risk, which is an internal judgment rendered as data. - Context on the two
financialhits: both appear in an "Order ref" column, likely benign.
Two layers, two jobs. The pattern scan found the credential and the identifiers; the contextual pass found the sentences and downgraded the noise. Neither layer alone would have cleared the page. The AI layer also cost tokens on your account — which is the honest tradeoff, and the reason it is a per-team decision rather than a default. A founder on Pro gets the first layer on every link at no token cost; the second layer is a choice you make when volume and stakes justify it.
"PII discovery tools" and "PII software": the same checklist, narrower scope
Buyers arrive at this category under several names, and the naming difference is mostly about scope rather than architecture.
PII discovery tools
Products marketed as PII discovery tools usually emphasize regulated personal identifiers specifically — names, addresses, national IDs, dates of birth, contact details — because the buying trigger is a privacy regulation rather than a security incident. They tend to be strong on classification taxonomy and jurisdiction mapping, and weaker on interception: the output is a register, and the register is the deliverable. Apply questions 3 and 7 hardest here. If the tool exists to satisfy a regulator, ask what it leaves behind when the regulator asks, and whether it can act at the moment of exposure or only enumerate it afterward. What counts as PII is broader than most teams assume, which is why taxonomy questions matter more in this corner of the market.
PII software (the lightweight end)
"PII software" as a search term tends to surface smaller point tools: redaction utilities, form scrubbers, scanning libraries a developer can embed. These are often excellent at one narrow job and cheap enough to skip procurement — which is exactly the risk, and the risk is highest at a startup, where there is no procurement to skip. Question 5 governs here. A free scanning service that posts your document to an API you have not reviewed has converted a privacy control into a privacy incident. If a tool is small enough to adopt without a review, it is small enough that someone will adopt it without a review; that dynamic is covered in sanctioning tools before the team finds their own.
How HTMLvault answers the checklist
HTMLvault is not a discovery platform. It will not inventory your warehouse, watch your endpoints, or classify your buckets — if you need those, buy those. It occupies the fourth position: every piece of HTML that becomes a link is scanned at the moment of publishing. That is the checkpoint that exists for everyone, from a solo founder on the Free plan to a security org with three at-rest tools already deployed.
Detection method — both layers. Layer one is deterministic pattern scanning across nine categories: SSN, financial, api_key, passport, address, person, dob, email, and phone. It is regex-based, runs in milliseconds, and consumes zero tokens, which is why it runs on every link on every plan rather than on a sample. Layer two is optional contextual AI review: Teams and Enterprise customers connect their own Anthropic, OpenAI, or Google API key for a semantic pass on top of the pattern scan. HTMLvault funds no tokens and takes no margin on them, so the "where does the content go" answer is your model account, under your data-processing agreement, with your provider's retention terms.
Position — publish-time, not scheduled. Findings surface before the URL is generated, not on a crawl afterward.
Surfacing — in context. Findings are grouped by category with the matched location highlighted, so the author fixes the source instead of guessing.
How to run a scan
- In the app: paste or upload HTML when creating a link. Findings appear before the link exists.
- Via API or MCP: call
scan_htmlon its own to check content without creating anything. This is the one to wire into an assistant's instructions — the model scans its own output before it ever callscreate_link. The MCP integration and REST API both expose it. - Review and decide: redact and rescan, or proceed knowingly. Pair the decision with the controls that limit blast radius — auto-expiry, password protection, and a retention window down to auto-delete.
Limits and caveats
Being honest about the boundaries is what makes a tool defensible in a review. It is also what makes a founder trust the tool enough to keep using it after the first flagged link.
- A scanner reads the file, not the recipients. It has no view of what happens after the link goes out — who it gets forwarded to, who screenshots it, who leaves it open on a shared screen. Detection tells you what is in the artifact. Containment is a separate job, done by expiry, passwords, and retention windows.
- A last-mile scanner only sees what passes through it. If half your team still emails attachments, none of that half is scanned. Coverage comes from making the sanctioned path the easy path — a smaller lift at a ten-person company than a thousand-person one, which is an argument for setting the default early. Replacing the attachment habit is most of the work, and a documented publishing workflow is how the default sticks.
- Pattern detection has both error modes. It misses sensitive prose it cannot pattern-match and flags strings that merely look like identifiers, as the worked example shows. The BYOK layer narrows the first gap; nothing eliminates human review entirely.
- Contextual AI is a paid tier and a real cost. It requires Teams or Enterprise plus your own model key, and results can vary between runs. Budget it against volume before enabling it on every link.
- Scanning is not access control. A clean scan means no detected identifiers, not "safe to share with anyone." Public links versus controlled access is the decision that follows a clean scan, not one it makes for you.
- Scanning is not a view log. Knowing a page was clean at publish tells you nothing about where it travelled afterward; that answer comes from per-link analytics, which is a separate control with a separate purpose.
- It complements the at-rest categories, it does not replace them. Present it in a security review as coverage for a currently unmonitored channel, not as a substitute for your DLP or classification program. The enterprise sharing checklist and the criteria in choosing a compliant HTML sharing solution are useful companion documents for that conversation.
For the RevOps lead whose dashboards go out as links, the marketer whose landing page drafts contain more than the copy, the IT reviewer who signs off on both, and the founder who is currently all three: the gap in most sensitive data discovery software isn't in the products you already bought. It's in the step that happens after they finish scanning — the moment an artifact stops being a file and becomes a URL. Evaluate for that step and you stop guessing what left the building, because the discount, the roadmap date, and the investor name get their last read before the link exists rather than after someone forwards it.
