references/VERIFICATION-AND-CUTOVER.md
references/VERIFICATION-AND-CUTOVER.mdBrowse 4 files
878 tokens
4,546 bytes
Token encoding: o200k_base
Snapshot 9e9fdc4
← Back to SKILL.md
Verification and cutover
Prove parity through the application's supported boundary. Do not require high-risk exercises for behavior the source does not use.
Build a contract ledger
For each traced source behavior, maintain one row:
| Source lifecycle | Observable contract | Owner | Target difference | Evidence |
|---|---|---|---|---|
| What initiates and completes it | Inputs, outputs, events, errors, ordering, state or effects callers observe | Application, Core, Harness, or Gap | What will not remain identical | Baseline and focused target test |
Evidence states are preserved, accepted change, application-owned, tested gap, or unverified. Do not call unverified equivalent.
Match testing to the slice
Ordinary core agent
- Characterize the public request and result.
- Use
TestModelorFunctionModelto prove prompt, dependency-backed tool calls, validated output, retry/error mapping, messages, and usage without network access. - Add a provider recording or live probe only when transport, native tools, structured-output mode, or provider deltas are part of the contract.
Stateful or branching agent
- Serialize the first result's normalized messages through the application's real store and load them in a fresh process for the next run.
- Test branch point, branch independence,
conversation_id/run_idsemantics, incomplete tool-call history, retention, tenant authorization, and store failures that callers can observe. - If the source resumes in-flight work, kill and restart at each promised boundary. Select a durable runtime deliberately and assert that an external effect is not duplicated.
Approvals and protected effects
- Assert the deferred call ID and validated arguments.
- Exercise deny, approve, and argument override where supported.
- For inline resolution, assert the handler runs and the agent completes in one call. When the decision arrives later, assert the first run ends with
DeferredToolRequests, persist that complete request or an equivalent pending-action record with category, validated arguments, and metadata, then assert a new run over the persisted messages withDeferredToolResultscompletes with the final output. - Assert no protected side effect before resolution, zero after denial, and exactly one after approval.
- Exercise the application's authenticated identity, policy lookup, and audit trail. A local yes/no callback is not authorization proof.
Streaming, hooks, and subagents
- Record a source golden trace containing only event types and stable fields the caller consumes.
- Test text reconstruction, tool start/result order, hook order and mutations, terminal detection, trailing events, cancellation, retry, and parallel calls.
- For subagents, test task-only input, separate history, result handback, dependencies, budget/cancellation/error propagation, event forwarding, and recursion limits.
Coding and shell agents
- Use a disposable workspace. Exercise normal reads/edits/search/commands plus
.., absolute paths, escaping symlinks, denied/protected files, output limits, timeouts, background cleanup, environment stripping, and working-directory behavior. - Demonstrate that an allowlisted interpreter can spawn another command. If untrusted code is in scope, run the containment test inside the selected container, VM, or cloud sandbox.
- If rollback is promised, verify filesystem or VCS state independently of conversation history.
Dependency and cutover checks
- Resolve the project from a clean environment against the chosen supported Pydantic AI and Harness versions. Prefer
pydantic-ai-slimwith only the needed extras when the dependency surface is bounded. - Run the original characterization tests and focused target tests. Test sync, async, callback, and streaming forms only where callers use them.
- Search imports, factories, configuration, transcript readers, event/result adapters, and deployment scripts for retained Claude Agent SDK dependencies.
- Remove
claude-agent-sdk, Claude Code subprocess setup, and Claude transcript assumptions only when no retained path needs them. - Report which checks used deterministic fakes, recordings, live providers, restart tests, or sandbox tests, and state any remaining limitation.
Completion criterion
Cut over only when each ledger row has a non-unverified evidence state, the original supported boundary passes, requested gaps are resolved or explicitly accepted, dependencies resolve cleanly, and rollback remains possible.
Referenced from SKILL.md
SKILL.mdView in source ↗
Source excerpt starting at line 22.SKILL.mdView in source ↗22Read [Research and concept mapping](references/RESEARCH-AND-MAPPING.md) for the detected source features. Read [Verification and cutover](references/VERIFICATION-AND-CUTOVER.md) before implementation.
Source excerpt starting at line 57.57Apply the completion criterion in [Verification and cutover](references/VERIFICATION-AND-CUTOVER.md). Label evidence from fakes, recordings, and live providers accurately.