Skip to content
AI Primer
release

ContextQA ships Ship to test deployed user flows

Ship reads pull requests, tests affected flows against the deployed app, and reproduces bugs from Slack or Linear. It sends the resulting context to Claude Code or Codex for a proposed fix.

5 min read
ContextQA ships Ship to test deployed user flows
ContextQA ships Ship to test deployed user flows

TL;DR

  • Ship turns a pull request into targeted user-flow tests against the deployed application, rather than stopping at code-diff analysis, according to rohanpaul_ai's description.
  • It reproduces bugs from Slack and Linear, captures evidence, investigates likely causes, and hands the result to an engineer or coding agent, as testingcatalog's launch post describes.
  • The handoff is designed to give Claude Code or Codex the context needed to fix the bug, a workflow omarsar0's report identifies as a missing feedback loop for agent-written code.
  • Ship spans PR impact analysis and production anomaly detection, while pauliusztin_'s eval framework puts regression checks between internal benchmarks and live production scoring.

The Ship product page shows a concrete activity stream: 148 tests across checkout, auth, and billing, two defects, two regression tests, then a root cause at auth/refresh.ts:87. A second panel turns 1,284 sessions into one reproduced anomaly and a permanent regression test. An independent launch summary describes the same PR and bug-handoff loop, while the ContextQA homepage positions Ship beside the company's broader test automation platform.

Deployed PRs

Ship reads a pull request, maps its blast radius, identifies the affected user flows, and runs targeted regressions against the deployed app, according to rohanpaul_ai's description.

  • Impact analysis: identify which flows a change can touch.
  • Targeted execution: run those flows in the deployed product before merge.
  • Defect evidence: record the failing behavior and attach actionable findings to the change.
  • Coverage growth: create or update regression tests when a new failure becomes known.

The official example shows three affected flows, 12 targeted regressions, and an expired-session defect blocking the payment step, all attached to a change in checkout handling. That puts browser behavior next to the code change that may have caused it, rather than leaving a person to discover the broken flow after release.

Bug reports

A bug report becomes a reproduction job. Ship can pick up reports from Slack, Linear, Jira, or a service desk, then replay the behavior in the running application, as testingcatalog's launch post reports.

  • Reproduce: replay the reported steps in the deployed app.
  • Capture: collect a session replay, trace, screenshots, and network evidence.
  • Investigate: use code, tickets, tests, documentation, and product context.
  • Hand off: return the likely root cause, reproduction steps, and suggested fix to an engineer, Claude Code, or Codex.

The product page's example follows a checkout spinner through session restoration, payment retry, and a stalled request before pointing to a null token in auth/refresh.ts. The documented role for Ship is investigation and context assembly. The coding agent remains the component that proposes or applies the code change.

Production signals

Ship also watches real user sessions for anomalies, reproduces suspicious behavior, and turns confirmed failures into permanent regression coverage, according to the Ship product page.

That production path gives the system a second trigger besides a pull request:

  • Session analysis: inspect real user trajectories.
  • Anomaly detection: flag behavior that looks unusual.
  • Automatic reproduction: replay the suspicious session.
  • Regression coverage: add the confirmed behavior to future runs.

omarsar0 framed the underlying bottleneck plainly: verifying agent-written code can take more time than writing it in the launch discussion. Ship's product loop is built around moving that verification from a manual browser check into a repeatable test and evidence pipeline.

Evaluation

The evaluation framing around Ship is broader than a pass or fail test. pauliusztin_'s eval framework separates internal benchmarks, pre-ship regression tests, and post-ship production evaluation. Vtrivedy10's model starts with production trajectories, turns them into evals with a task, environment, expected outcome, and verifier, then uses the resulting simulation to compare models, tools, prompts, harnesses, and infrastructure Vtrivedy10's eval thread.

Ship's page includes a sample release report rather than an independently documented benchmark. Its reported deltas are:

  • Quality score: 84 → 92, +8 points.
  • Test coverage: previous baseline → current release, +12%.
  • Cost per release: previous baseline → current release, -18%.
  • Escaped defects: previous baseline → current release, -36%.

The same card lists 31 code changes analyzed, 47 impacted tests identified, three regressions caught, one production anomaly reproduced, and two regression tests added. The page does not publish a cohort, sample size, or evaluation methodology for those figures.

Integrations and access

Ship is designed to sit inside an existing engineering stack. The launch materials describe code and CI integrations, issue trackers, Slack, service desks, and direct handoffs to Claude Code and Codex, with testingcatalog's follow-up emphasizing the engineer or coding-agent destination for each finding.

The current product page lists:

  • GitHub repositories and deployed environments for change analysis.
  • Linear, Jira, Slack, and service-desk inputs for bug investigation.
  • Claude Code and Codex as downstream repair agents.
  • A 15-day trial with 3,000 credits.
  • Starter at $49 per month for 2,000 credits, and Pro at $249 per month for 12,000 credits.
  • Additional credits at $0.025 each, with unlimited users and no seat pricing.

Limitations

The public launch material leaves two mechanics open. omarsar0 asked whether the coding agent receives a failing test and likely cause or the full reproduction trace omarsar0's question. rohanpaul_ai asked how Ship maps a shared-component change to the user flows it affects rohanpaul_ai's question.

The product page shows reproduction steps, traces, and suggested fixes in its examples, but does not specify the exact handoff payload or the flow-selection method. It also lists real-device mobile testing as coming soon for Pro, with dedicated devices reserved for enterprise customers.

The available demo clips are similarly narrow: badlogicgames's clip and nptacek's clip show code-editor and browser transitions, but do not expose a test result, trace, or Claude/Codex handoff.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR2 posts
Deployed PRs2 posts
Bug reports1 post
Evaluation2 posts
Integrations and access1 post
Share on X