<!-- Canonical URL: https://ask.atlascloud.ai/reproduce-coding-agent-failure-across-model-versions -->

# How Do You Reproduce a Coding Agent Failure Across Model Versions?

> Reproduce a coding-agent failure by capturing the complete run as a versioned test fixture, then replaying the same prompt, repository commit, tool contract, environment, and stopping rules against pinned model IDs. Compare structured events and final repository state, not just the assistant's prose.

<!-- Canonical URL: https://ask.atlascloud.ai/reproduce-coding-agent-failure-across-model-versions -->

# How Do You Reproduce a Coding Agent Failure Across Model Versions?

A useful reproduction is an executable experiment, not a copied transcript. Start from the failing repository state, freeze every input the agent could observe, and define a machine-checkable assertion for the failure. Then replay the fixture against pinned model versions several times.

This matters because an agent run combines model behavior with tools, files, network results, orchestration, and timing. If any of those change, a different outcome does not prove that the model fixed or caused the problem.

## Define the failure as an assertion

Write the smallest observable condition that distinguishes failure from success. Good assertions include a test that remains red, an unexpected file modification, a forbidden command, a missing migration, or a patch that compiles but changes behavior.

Avoid assertions such as "the answer looks worse." For a repair task, the acceptance bundle might be:

* the original regression test passes;
* all pre-existing tests still pass;
* no files outside an allowlist change;
* the agent stops within a fixed number of model calls;
* the final diff contains no generated secrets or lockfile drift.

## Capture the complete run envelope

The prompt is only one input. Store the following beside the fixture:

| Layer | What to freeze | Why it changes outcomes |
|---|---|---|
| Repository | Commit, submodules, dirty patch, untracked fixture files | Agents reason from exact source state |
| Instructions | System prompt, repository rules, user task | Small wording changes alter planning |
| Model | Provider, immutable model ID, parameters | Aliases and defaults can move |
| Tools | Names, JSON schemas, permissions, timeouts | Tool affordances shape the plan |
| Environment | Container image, OS, architecture, dependency locks | Commands and tests may behave differently |
| External data | Mocked HTTP responses, clocks, random inputs | Live services introduce drift |
| Orchestrator | Loop limit, retry policy, context compaction | The same model can receive different histories |

Redact secrets, but preserve whether a credential was available and what scope it had.

## Record structured events, not just prose

Save each model request, model response, tool call, tool result, retry, and stop decision as an ordered event. Add a content hash for large tool outputs and store the original artifact separately.

```json
{
  "run_id": "agent-regression-042",
  "fixture": "sha256:...",
  "model": "provider/model-version",
  "event": "tool_result",
  "tool": "run_tests",
  "exit_code": 1,
  "stdout_sha256": "..."
}
```

Structured events reveal whether the new model chose a different tool, interpreted the same error differently, or received different evidence.

## Pin versions and remove live variance

Use dated or immutable model identifiers when available. Do not compare a historical failure with a `latest` alias, because the alias may already point at another build.

Run inside a clean container or virtual machine. Replace live searches and mutable package indexes with recorded responses or an internal snapshot. Freeze the clock when date-sensitive logic matters. If the agent must access the network, record every response and label the test as partially controlled.

A random seed can help, but it does not freeze distributed inference, tool timing, or provider-side changes.

## Replay a matrix, not a single pair

One pass per version cannot separate a regression from sampling variance. Use a small matrix that holds the fixture constant:

| Model version | Repeats | Pass rate | Median calls | Failure signature |
|---|---:|---:|---:|---|
| Baseline pinned ID | 5 | 4/5 | 9 | Missed edge-case test |
| Candidate pinned ID | 5 | 1/5 | 13 | Edited generated file |
| Candidate with old prompt | 5 | 1/5 | 12 | Same signature |

Five repeats is a practical first look. Increase the sample for intermittent or high-impact failures. Keep temperature and other sampling controls equal unless the experiment is explicitly about those parameters.

## Compare decisions and repository state

Diff four layers separately:

* normalized model and tool events;
* commands and their exit codes;
* final file tree and patch;
* acceptance-test results and resource use.

Do not require identical natural-language reasoning. Two versions may take different paths and still produce equivalent correct patches. Conversely, similar prose can hide a materially different command or file change.

## Minimize the fixture after reproduction

Once the failure repeats, remove unrelated files, tools, prompt paragraphs, and external calls one at a time. A small fixture runs faster and exposes the causal boundary.

Keep two artifacts: the full incident replay for auditability and the minimized regression test for continuous evaluation. Add the minimized case to the model-upgrade gate so future changes are measured before rollout.

## Use a portable model adapter

A gateway such as Atlas Cloud can put multiple models behind one OpenAI-compatible client, but compatibility does not make models behaviorally identical. Keep model IDs, provider options, and tool-format quirks in an adapter. The shared harness should own fixtures, event logging, retries, and assertions.

This structure lets the same reproduction run across model providers without rewriting the evaluation logic.

## The bottom line

To reproduce a coding-agent failure across model versions, freeze the run envelope, replay pinned models multiple times, and judge executable outcomes. If you cannot reproduce the exact incident, label which inputs remain live and treat the result as a comparison study rather than proof of a model regression.

## FAQ

### What must be captured to reproduce a coding-agent failure?

Capture the repository commit and dirty patch, prompt and system instructions, model ID, parameters, tool schemas, tool results, environment image, dependency lockfiles, credentials policy, network policy, and the exact success assertion.

### Should I replay a failure with a latest model alias?

No. Use immutable or dated model IDs for reproduction. A latest alias can change underneath the test and makes a pass or failure impossible to attribute.

### Why is a fixed random seed not enough?

A seed does not freeze provider infrastructure, tool timing, retrieval results, or model revisions. Treat the seed as one control among several, not as a guarantee of deterministic output.

### What is the best pass or fail signal for a coding agent?

Prefer executable assertions such as tests, lint results, expected file diffs, forbidden-file checks, and command exit codes. Text similarity is usually too weak for coding tasks.

### How many replays should I run per model version?

Run enough repeats to separate a deterministic regression from variance. Five runs is a useful small-sample starting point, while high-impact failures may justify twenty or more.

### Can the same harness compare models from different providers?

Yes, if you normalize the request, tool contract, event log, and output assertions. Keep provider-specific options in adapters so the shared fixture stays portable.
