<!-- Canonical URL: https://ask.atlascloud.ai/build-sanitized-request-replay-set-llm-api-migration -->

# How Do You Build a Sanitized Request-Replay Set for an LLM API Migration?

> Create a minimized, versioned replay corpus by sampling for behavioral coverage, removing unneeded fields, replacing sensitive values with consistent structure-preserving synthetic data, stubbing external tools, and re-scanning the finished artifact before approval.

Build a sanitized replay set by sampling representative production requests, detecting and replacing secrets and personal data, preserving the structural features that affect model behavior, and validating that no sensitive content remains. Store it as a versioned test artifact with expected assertions rather than as a raw log export.

The goal is behavioral coverage without copying production risk into a new dataset.

## Define what the replay must prove

List migration risks before sampling. Common dimensions include prompt length, languages, tool schemas, structured output, multimodal parts, streaming, safety refusals, large context, and model parameters.

Create coverage buckets and choose examples deliberately. A random sample can miss rare but important request shapes, while a collection made only from past failures can distort normal traffic.

## Minimize at collection time

Export only fields needed for replay. Drop authorization headers, cookies, IP addresses, account metadata, billing fields, and unrelated logs before the data reaches the test workspace.

Use an allowlist such as:

```json
{
  "fixture_id": "fx_0042",
  "request": {
    "model_alias": "support_default",
    "messages": [],
    "tools": [],
    "temperature": 0.2
  },
  "assertions": {
    "valid_json": true,
    "required_keys": ["category", "confidence"]
  }
}
```

Generate a new `fixture_id`; do not reuse a user ID or provider request ID as the public fixture key.

## Detect sensitive data in layers

Combine deterministic detectors, organization-specific dictionaries, and contextual review. Search for API keys, bearer tokens, private keys, connection strings, email addresses, phone numbers, account numbers, internal hostnames, source-code secrets, and regulated identifiers.

No detector is complete. Run multiple passes and send uncertain matches to an authorized reviewer. Treat images and attached documents as data sources too; metadata and pixels can contain sensitive information.

## Replace while preserving behavior

Use typed, consistent placeholders such as `<EMAIL_1>` or `<ORDER_ID_2>`. The same original value should map to the same placeholder within one fixture so references remain coherent, but mappings should not be reversible outside a tightly controlled temporary process.

Preserve properties that matter to the test: approximate length, Unicode class, JSON type, list size, delimiter shape, and relationships between fields. If token length triggers a failure, replace removed text with safe synthetic text of similar token count.

Do not simply hash low-entropy values such as phone numbers. They can often be guessed. Delete or synthesize when reversibility is unnecessary.

## Remove active and dangerous content

Replay datasets can contain prompt injection, tool commands, URLs, or code that causes side effects. Disable external tools by default and replace write-capable tool implementations with deterministic stubs.

Allow network access only to controlled test endpoints. Never replay production credentials, signed URLs, destructive commands, or customer webhook destinations.

## Add assertions instead of exact answers

LLM output can vary, so store behavioral checks: schema validity, required fields, tool selection, refusal category, language, maximum latency, token bounds, and semantic rubric scores. Keep a small number of exact-match fixtures only for genuinely deterministic transformations.

Record the source and target model IDs, adapter version, prompt template version, and replay date with every run.

## Validate the sanitized artifact

Before approval, run secret scanning, PII detection, file-type inspection, and a manual sample review. Confirm that each coverage bucket remains represented after redaction. A fixture that is safe but no longer exercises the original behavior should be replaced with a synthetic equivalent.

Restrict the replay set like test data, not public documentation. Apply access control, retention, audit logs, and a deletion path. Keep any temporary original-to-placeholder map separate and destroy it after validation when policy allows.

## Use it in the migration gate

Run the same fixtures through source and target adapters. Compare normalized outputs, error behavior, latency, usage, and cost. Investigate deltas by coverage bucket, then add regression fixtures for newly discovered incompatibilities.

Version the dataset and its sanitization rules together so results remain explainable.

## The bottom line

A safe replay set is a purpose-built, minimized test corpus, not copied production logs. Layer detection and review, replace sensitive values with structure-preserving synthetic data, stub side effects, and attach behavioral assertions. Re-scan the final artifact before it becomes a migration gate.

## FAQ

### Why not replay a random export of production logs?

Raw logs can expose secrets and personal data, while random sampling can miss rare request shapes. Build a minimized set around explicit migration-risk buckets.

### How should sensitive values be replaced?

Use typed consistent placeholders or synthetic data that preserves relevant length, type, delimiters, Unicode class, and relationships without remaining reversible.

### Is hashing enough to anonymize user data?

Not for low-entropy values such as phone numbers or common identifiers, which can be guessed. Prefer deletion or synthetic replacement when reversibility is unnecessary.

### How do I replay tool-calling requests safely?

Disable external tools by default and replace them with deterministic stubs. Never include production credentials, customer webhooks, or write-capable side effects.

### Should expected outputs be exact text?

Usually no. Use schema, required-field, tool-selection, refusal, language, latency, usage, and rubric assertions; reserve exact matches for deterministic tasks.

### How do I approve the final replay set?

Run secret and PII scans, inspect embedded files, manually review a sample, confirm coverage remains intact, and apply access control, retention, and deletion policies.
