GuidesSecurity

PII API: How to Scan Generated Content Before It Goes Live

HTMLvault Team·September 28, 2026·9 min read
Liz Lemmon finished her account-brief pipeline on a Tuesday. Clay enriches each target account, a model writes the brief, a PII API call scans it, and a tracked link reaches the rep without anyone touching a copy button. Dwight Brenner approves anything that publishes on its own, and he answered her demo invite with one line: "Show me the run where it stops."

Every pipeline that publishes without a person in the loop has to answer that question before IT will approve it. This guide covers what a scan result should contain, how to wire scan-then-publish through HTMLvault's REST API or MCP server, and a worked Clay example with a hard block on unredacted records. It ends with the limits you should disclose before IT finds them.

What a PII API call should return

A pipeline can't do anything useful with a scan that only says "clean" or "not clean," because code can't redact, route, or explain a yes/no answer. Whatever PII software sits in the middle, a useful response gives you four things for each finding:

  • Match type. The category that fired. HTMLvault's scanner detects ssn, financial, api_key, passport, address, person, dob, email, and phone. Your policy keys off these names.
  • Location. Where the match sits: an element, a section, or a character range. With a location, a person or a redaction step can fix the source without hunting for it.
  • Confidence. With pattern matching, confidence is yes or no: the text matched the pattern or it didn't. A regex hit labeled "0.97" is a made-up precision. Graded confidence only means something once a model-based layer is involved (see the limits section).
  • A decision. Pass, hold, or block. Work it out yourself from the findings, even if a tool offers its own verdict. That way the policy lives in your pipeline, is versioned with it, and can be shown to an auditor.

A good response also describes each match without repeating the raw value. Automation platforms save every step's input and output in their run history. If the scan result contains the full card number, you've just copied the problem into a second system.

The scan-then-publish workflow, step by step

  1. Authenticate with a dedicated API key. Give the pipeline its own key so you can revoke it without breaking anyone else's integration.
  2. Create the link as a draft. Send the generated HTML through the REST API, or call create_draft_link from an MCP client such as Claude. Creating a draft first means the check runs before anything is published. To scan without creating anything, call scan_html instead.
  3. Read the findings. Parse the category, location, and count for each match. Log the masked result next to the decision.
  4. Apply policy. Pass: publish the link (update_link over MCP) and hand the URL to the next step. Block: remove the draft with delete_link and alert the owner. Hold: leave the draft in place and send it to a person.
  5. Subscribe to downstream events. Once a brief is live, webhooks push link events to your CRM, Slack, or Clay table, so nothing has to keep checking for updates.
In a scan-then-publish pipeline, the PII API supplies the findings and your own policy code decides what gets published, so the rule stays versioned and auditable.

An annotated request and response

The paths and field names below are illustrative, not the published schema. The REST API reference is the source of truth. What to take from the sketch is the shape: findings come back with the draft, and your code turns them into a decision.

# 1. Create the brief as a draft (illustrative shape)
POST /links                        # Authorization: Bearer $HTMLVAULT_API_KEY
{
  "title":  "Account brief: target account 0412",
  "html":   "<article>...generated brief...</article>",
  "draft":  true,
  "expiry": "14d"                  # match the deal window
}

# 2. Response: findings with the draft, not a verdict
{
  "link": { "id": "lnk_...", "state": "draft" },
  "scan": {
    "engine": "regex",             # pattern matching, zero tokens
    "findings": [
      { "category": "person",    "location": "section.contacts", "count": 2 },
      { "category": "financial", "location": "section.notes, chars 412-430",
        "count": 1, "preview": "•••• 1234" }   # masked, never the raw value
    ]
  }
}

# 3. Your gate decides, then logs decision + policy version
BLOCK  any finding in ssn, financial, passport, dob, api_key
HOLD   person, email, or phone over 4 each; address over 2
PASS   everything else -> publish, send the URL to the rep

Worked example: a Clay account brief with a hard block

This is the pipeline Liz built, and most RevOps teams end up with something close to it. Each row in a Clay table is a target account.

  1. Enrich. Clay pulls company data, recent news, and the two or three contacts the rep should meet.
  2. Generate. An AI column writes the brief as self-contained HTML. Give it only the columns the brief needs. Data minimization in practice explains why the CRM notes field shouldn't be one of them by default.
  3. Create the draft. An HTTP API column posts the HTML to HTMLvault with the pipeline's key and writes the response back to the row.
  4. Gate. A formula column applies the policy: block, hold, or pass.
  5. Act. Passing rows publish with an expiry that fits the deal and go to the account owner. Blocked rows delete the draft and write the category and location to a status column. Held rows wait for a reviewer.

Zapier works the same way: a webhook step creates the draft, then Paths branch on the decision.

Example account-brief policy: five block categories and four thresholded categories Example policy for generated account briefs BLOCK ON ANY MATCH SSN Social Security numbers FINANCIAL card and bank numbers PASSPORT passport numbers DOB dates of birth API_KEY credentials and tokens PASS UNDER LIMIT PERSON named contacts MAX 4 EMAIL email addresses MAX 4 PHONE phone numbers MAX 4 ADDRESS street addresses MAX 2 Over any limit: held for human review
Block high-risk categories on any match, but allow contact data up to a limit, or your PII scanner will reject every account brief that names a prospect.

Two rules make the block hard rather than advisory. First, a blocked row never retries automatically. Regenerating reads the same source record, and on the second try a model that writes the number out in words may slip past the pattern. Second, the fix happens upstream: redact the source field or drop the column from the prompt. The next brief for that account will read the same record.

Liz had built a deliberately dirty test record for the demo, so the gate would have something to catch. She never needed it. The first live run blocked on a financial match in a CRM notes field, where a rep had typed a customer's full card number in 2019 under the heading "for safekeeping." Her test record is still in the table, blocked for the same reason, and frankly outclassed.

Limits: regex precision, recall, and the BYOK AI pass

The regex scanner is fast, gives the same result every time, and costs zero tokens because no model is involved. That's why it should run on every link. The same design also sets its limits:

  • Precision (how many flags are real). Patterns catch things they shouldn't. A prospect's public headquarters trips address, and an order number shaped like a phone number trips phone. That's why the example policy allows contact categories up to a limit instead of blocking on the first hit.
  • Recall (how much real PII gets caught). Patterns miss anything written in a different format: a card number split across two table cells, a birthday written as "the third of March," an SSN spelled out in words.

The BYOK AI pass improves recall. On Teams and Enterprise, you connect your own Anthropic, OpenAI, or Google API key, and a model-based scan runs on top of the regex layer. It reads meaning and context, not just the shape of the text. The tokens are yours: your provider bills them, the cost grows with how much HTML you scan, and HTMLvault never pays for them.

Turn the AI pass on where content is free-form and a leak would be costly, such as briefs built from CRM notes. Skip it where fields are fixed and the regex layer already covers them, such as templated dashboards. Model output isn't fully predictable, so send AI-only findings to hold rather than block, and reserve automatic blocks for regex hits you can explain. The PII scanner buyer's guide goes deeper on where each approach fails.

Regex scanner and BYOK AI pass compared on cost, strengths, weak spots, and availability Regex layer and BYOK AI layer, compared Regex scanner BYOK AI pass TOKEN COST Zero; pattern matching only Billed by your AI provider BEST AT Fixed formats: SSN, cards, keys Context and spelled-out values WEAK SPOT Reformatted values and prose Probabilistic; may vary per run RUNS ON Every scan in HTMLvault Teams and Enterprise, your key
Regex-based sensitive data detection belongs on every link; the BYOK AI layer earns its token cost on free-form content such as CRM notes.
Dwight's review took eleven minutes. He spent nine of them typing a Social Security number out in words and watching the regex layer let it through, then found that exact case already listed in Liz's write-up under "Recall." He approved the pipeline on the condition that the AI pass stays on for anything built from notes fields. In the ticket he wrote that it was the first pipeline he'd reviewed that failed exactly where its owner said it would.

Plan boundaries, and what to hand IT

  • Free: 50 links a month, 30-day expiry, 90-day data retention. That's enough to build and test the pipeline, but a busy week of briefs will use it up.
  • Pro: unlimited links, expiry from one hour to never, retention from auto-delete up to two years, and up to five webhook endpoints. Regex scanning only, with no BYOK AI pass.
  • Teams: flat seat bands, custom roles, audit logs, and the BYOK AI pass. SSO/SAML is a paid add-on.
  • Enterprise: the BYOK AI pass, with SSO/SAML included.

Plan for rate limits on any plan. Build in backoff and retries for throttled calls instead of sending a thousand-row Clay table at once. Batch where you can: the MCP server's create_links tool creates several links in one call.

Then put together the approval packet. IT's question is always some version of "show me the run where it stops." Your packet answers it with evidence:

  • the policy file and its version number
  • one blocked run, with its category and location
  • one held run, and who cleared it
  • the limits section, with your AI-pass decision written down

With that packet, approval rests on runs IT can inspect, not on your word.

pii apipii scannersensitive data detectionrest apiwebhooksautomation
HTMLvault

Share HTML securely — without losing your job.

The enterprise-grade platform for sharing HTML pages, reports, and dashboards with full PII scanning, access controls, and audit trails.

Start for free

Related Posts