Solutions — AI Providers · GPAI

The training sample that never shipped

How a model provider kept an unredacted user conversation out of a public release — and turned training-data provenance from a slide into a scan ledger.

Try the free EU AI Act check →

FRIDAY, 18:30 — MODEL DROP

The fine-tune is done. For reproducibility, someone commits the eval set next to the weights. Row 1,204 is an unredacted conversation with a real user. It's Friday evening; the release tag is 10 minutes away.

FRIDAY, 18:30:09 — CI

The merge is blocked: unredacted user content in a training artifact, provenance rule cited, masked evidence only. The release ships Monday — with the eval set scrubbed and the ledger showing exactly when it was caught.

The model still shipped. Row 1,204 didn't.

The problem we kept living

Before the gate, there was the provenance question

If you train and ship models, you already know this question:

  • "What's in your training data?" had a slide-deck answer, not a ledger answer.
  • Eval sets, fixtures and debug dumps travelled next to weights — committed by whoever was closest to the Friday deadline.
  • The GPAI obligations turned documentation from a courtesy into a duty: transparency about training data, and evidence you enforce your own policy.
  • Review caught what reviewers happened to open. Nobody opens row 1,204 of an eval set at 18:30 on a Friday.

So provenance stopped being a slide and became a merge check.

What it does

The guided tour — from commit to evidence

01

A rule your policy team can read

Compliance rules are plain declarations — what to detect, where to look, what happens on a hit. Add your own signatures for user-content markers and internal dataset tags without asking us.

RULE · DC-GPAI-003BLOCKS MERGE
Detectunredacted user content
Wheredatasets · eval sets · dumps
CitesGPAI transparency obligations
02

The merge that fails politely

A hit fails the check with the file, row, and rule — and masked evidence only. The reviewer sees enough to scrub it; the content itself never leaves your CI runner.

PR #2117 · CHECK RESULTFAILED
eval/holdout_v3.jsonl:1204user text · masked
RuleDC-GPAI-003 · critical
Evidencemasked · runner-local
03

Drift, scored between releases

Hourly process rules score dataset-documentation coverage and review habits against your Jira and GitHub activity — so the gap between policy and practice shows up as a trend, not at the next release.

COMPLIANCE · BY DOMAINTHIS SPRINT
Training-data provenance96%
Dataset documentation90%
Review coverage88% ↘
04

Sign-off that stays locked

What needs human judgment — the model-card review, the release risk checklist — gets a named reviewer and a locked gate. The result exports signed and timestamped: evidence, not screenshots.

RELEASE GATE · MODEL v3.1LOCKED
CI scan — 6 reposPASSING
Model-card review · S. Okaforawaiting sign-off
Audit exportsigned · sha256:c7d0…
What changed

The same release, one cycle later

Before

  • —"What's in your training data?" — a slide-deck answer
  • —Eval sets travelled next to weights, unscanned
  • —Provenance policy was a document nobody enforced
  • —Friday releases were a race against whoever committed last

After

  • ✓A scan ledger: every artifact merge since v2.9, with rule hits and outcomes
  • ✓Training artifacts pass the same gate as source code
  • ✓It's a blocking check with a named rule
  • ✓The tag waits for the gate — Monday beats a takedown
Is this for you

An honest fit check

This fits if

  • ✓You fine-tune or train models and ship them (weights, APIs, or both)
  • ✓Datasets, eval sets and fixtures move through the same repos as your code
  • ✓GPAI transparency obligations apply to you and "trust us" isn't documentation
  • ✓You'd rather block a merge than scrub a published artifact

And honestly, if

  • ·You need semantic judgment — "is this dataset ethically sourced?" is a human's call. PulseCheck routes it to a locked, role-restricted attestation instead of pretending to detect it.
  • ·Your training data never touches a git repo or CI — the gate has nothing to hook into.
  • ·You want developer-level scorecards — deliberately not built, and it won't be.
NewExpert Review

Put the judgment calls to an independent lawyer

Some questions stay a human's call. With Expert Review you can also put them to an independent, licensed lawyer, and their verdict appears next to each rule.

How Expert Review works →

GPAI provider compliance FAQ

Can PulseCheck block a pull request?

Yes. The CI Action runs your organization's ruleset against every diff and fails the check when a defined signature (an unredacted training-data sample, a hardcoded credential) is present — configure it as a required status check and the PR cannot merge.

Does PulseCheck read our source code?

The scan runs inside your own CI runner. Findings are masked before they ever leave it — PulseCheck's servers see a match/no-match result and a masked snippet, never your raw source or training data.

Does it detect systemic-risk classification or model-safety issues?

No — honestly. That's a judgment call, not a pattern match. PulseCheck routes it to a locked sign-off task a named reviewer must explicitly attest, with the attestation recorded in an audit-ready export. With Expert Review you can also put that question to an independent, licensed lawyer.

How does PulseCheck help with EU AI Act GPAI obligations?

PulseCheck ships an EU AI Act rule template pack covering model-documentation and transparency obligations, evaluated hourly against your Jira/GitHub activity, plus a compliance gate with a locked sign-off task and a signed, timestamped export for your audit trail.

Which compliance frameworks ship as templates?

GDPR, EU AI Act, DORA, and a general QA/incident-management pack ship as ready-to-install rule templates — install one and PulseCheck starts scoring your existing data against it immediately.

Bring us your next model drop

A 20-minute walkthrough is enough to see your own repository scanned. First rule live the same day — before the next Friday tag.