Every pipeline that publishes without a person in the loop has to answer that question before IT will approve it. This guide covers what a scan result should contain, how to wire scan-then-publish through HTMLvault's REST API or MCP server, and a worked Clay example with a hard block on unredacted records. It ends with the limits you should disclose before IT finds them.
What a PII API call should return
A pipeline can't do anything useful with a scan that only says "clean" or "not clean," because code can't redact, route, or explain a yes/no answer. Whatever PII software sits in the middle, a useful response gives you four things for each finding:
- Match type. The category that fired. HTMLvault's scanner detects
ssn,financial,api_key,passport,address,person,dob,email, andphone. Your policy keys off these names. - Location. Where the match sits: an element, a section, or a character range. With a location, a person or a redaction step can fix the source without hunting for it.
- Confidence. With pattern matching, confidence is yes or no: the text matched the pattern or it didn't. A regex hit labeled "0.97" is a made-up precision. Graded confidence only means something once a model-based layer is involved (see the limits section).
- A decision. Pass, hold, or block. Work it out yourself from the findings, even if a tool offers its own verdict. That way the policy lives in your pipeline, is versioned with it, and can be shown to an auditor.
A good response also describes each match without repeating the raw value. Automation platforms save every step's input and output in their run history. If the scan result contains the full card number, you've just copied the problem into a second system.
The scan-then-publish workflow, step by step
- Authenticate with a dedicated API key. Give the pipeline its own key so you can revoke it without breaking anyone else's integration.
- Create the link as a draft. Send the generated HTML through the REST API, or call
create_draft_linkfrom an MCP client such as Claude. Creating a draft first means the check runs before anything is published. To scan without creating anything, callscan_htmlinstead. - Read the findings. Parse the category, location, and count for each match. Log the masked result next to the decision.
- Apply policy. Pass: publish the link (
update_linkover MCP) and hand the URL to the next step. Block: remove the draft withdelete_linkand alert the owner. Hold: leave the draft in place and send it to a person. - Subscribe to downstream events. Once a brief is live, webhooks push link events to your CRM, Slack, or Clay table, so nothing has to keep checking for updates.
An annotated request and response
The paths and field names below are illustrative, not the published schema. The REST API reference is the source of truth. What to take from the sketch is the shape: findings come back with the draft, and your code turns them into a decision.
# 1. Create the brief as a draft (illustrative shape)
POST /links # Authorization: Bearer $HTMLVAULT_API_KEY
{
"title": "Account brief: target account 0412",
"html": "<article>...generated brief...</article>",
"draft": true,
"expiry": "14d" # match the deal window
}
# 2. Response: findings with the draft, not a verdict
{
"link": { "id": "lnk_...", "state": "draft" },
"scan": {
"engine": "regex", # pattern matching, zero tokens
"findings": [
{ "category": "person", "location": "section.contacts", "count": 2 },
{ "category": "financial", "location": "section.notes, chars 412-430",
"count": 1, "preview": "•••• 1234" } # masked, never the raw value
]
}
}
# 3. Your gate decides, then logs decision + policy version
BLOCK any finding in ssn, financial, passport, dob, api_key
HOLD person, email, or phone over 4 each; address over 2
PASS everything else -> publish, send the URL to the rep
Worked example: a Clay account brief with a hard block
This is the pipeline Liz built, and most RevOps teams end up with something close to it. Each row in a Clay table is a target account.
- Enrich. Clay pulls company data, recent news, and the two or three contacts the rep should meet.
- Generate. An AI column writes the brief as self-contained HTML. Give it only the columns the brief needs. Data minimization in practice explains why the CRM notes field shouldn't be one of them by default.
- Create the draft. An HTTP API column posts the HTML to HTMLvault with the pipeline's key and writes the response back to the row.
- Gate. A formula column applies the policy: block, hold, or pass.
- Act. Passing rows publish with an expiry that fits the deal and go to the account owner. Blocked rows delete the draft and write the category and location to a status column. Held rows wait for a reviewer.
Zapier works the same way: a webhook step creates the draft, then Paths branch on the decision.
Two rules make the block hard rather than advisory. First, a blocked row never retries automatically. Regenerating reads the same source record, and on the second try a model that writes the number out in words may slip past the pattern. Second, the fix happens upstream: redact the source field or drop the column from the prompt. The next brief for that account will read the same record.
financial match in a CRM notes field, where a rep had typed a customer's full card number in 2019 under the heading "for safekeeping." Her test record is still in the table, blocked for the same reason, and frankly outclassed.Limits: regex precision, recall, and the BYOK AI pass
The regex scanner is fast, gives the same result every time, and costs zero tokens because no model is involved. That's why it should run on every link. The same design also sets its limits:
- Precision (how many flags are real). Patterns catch things they shouldn't. A prospect's public headquarters trips
address, and an order number shaped like a phone number tripsphone. That's why the example policy allows contact categories up to a limit instead of blocking on the first hit. - Recall (how much real PII gets caught). Patterns miss anything written in a different format: a card number split across two table cells, a birthday written as "the third of March," an SSN spelled out in words.
The BYOK AI pass improves recall. On Teams and Enterprise, you connect your own Anthropic, OpenAI, or Google API key, and a model-based scan runs on top of the regex layer. It reads meaning and context, not just the shape of the text. The tokens are yours: your provider bills them, the cost grows with how much HTML you scan, and HTMLvault never pays for them.
Turn the AI pass on where content is free-form and a leak would be costly, such as briefs built from CRM notes. Skip it where fields are fixed and the regex layer already covers them, such as templated dashboards. Model output isn't fully predictable, so send AI-only findings to hold rather than block, and reserve automatic blocks for regex hits you can explain. The PII scanner buyer's guide goes deeper on where each approach fails.
Plan boundaries, and what to hand IT
- Free: 50 links a month, 30-day expiry, 90-day data retention. That's enough to build and test the pipeline, but a busy week of briefs will use it up.
- Pro: unlimited links, expiry from one hour to never, retention from auto-delete up to two years, and up to five webhook endpoints. Regex scanning only, with no BYOK AI pass.
- Teams: flat seat bands, custom roles, audit logs, and the BYOK AI pass. SSO/SAML is a paid add-on.
- Enterprise: the BYOK AI pass, with SSO/SAML included.
Plan for rate limits on any plan. Build in backoff and retries for throttled calls instead of sending a thousand-row Clay table at once. Batch where you can: the MCP server's create_links tool creates several links in one call.
Then put together the approval packet. IT's question is always some version of "show me the run where it stops." Your packet answers it with evidence:
- the policy file and its version number
- one blocked run, with its category and location
- one held run, and who cleared it
- the limits section, with your AI-pass decision written down
With that packet, approval rests on runs IT can inspect, not on your word.
