Evidence the run may read
- Failure logs from multiple runs
- Test source and fixtures
- Runtime and timing metadata
Cluster repeated failures, isolate timing or state signals, and propose the smallest experiment that can confirm the cause.
Search intent: AI agent to diagnose flaky tests
A test crosses the configured intermittent-failure threshold.
engineering owner
The test owner approves quarantine, fixture changes, or retry-policy changes.
open flaky-test age (days). Track days from first threshold breach to verified disposition.
Each stage produces an artifact another stage can inspect. The final node is a person, not an autonomous write to an external system.
Validate and normalize the flaky test diagnosis inputs.
Produce the failure cluster summary.
Produce the likely nondeterminism sources.
Challenge the flaky test diagnosis result and prepare an approval packet.
The test owner approves quarantine, fixture changes, or retry-policy changes.
Source material is read-only. Drafts land in a case-specific output path. Tools may read or propose; the human gate owns the external write.
workspace/engineering/flaky-test-diagnosis/inputs/**Read only the evidence attached to this workflow instance.
workspace/engineering/flaky-test-diagnosis/outputs/**Write drafts and evidence artifacts without modifying source records.
engineering:flaky-test-diagnosis:read-or-proposeInvoke only tools explicitly granted for this run; external writes remain gated.
The blueprint starts private, caps its DAG, disables replanning, and exposes no public endpoint. Add only the tools and data adapters this workflow has approved.
name: engineering-flaky-test-diagnosis
version: 0.1.0
entrypoint: agent:BlueprintAgent
expose:
public: false
composition:
planning: deterministic_dag
max_nodes: 6
max_parallel: 1
max_replans: 0Every material conclusion cites an input artifact or a scoped tool result from this run.
The diagnosis uses evidence from multiple runs and does not treat a retry pass as proof of repair.
The run stops at a proposal and records the human decision before any external side effect.
Containment: Return a partial result with unresolved items; do not broaden scope or perform an external write.
Operator: Attach the missing evidence, narrow the brief, or explicitly approve a new scoped run.
Containment: Stop the affected branch and preserve completed artifacts in the case output workspace.
Operator: Grant only the missing resource or continue with that branch marked out of scope.
Current platform receipts sign caller identity or classification, skill, bounded input evidence, verified grant IDs when present, outcome or result preview, and timing. Optional file, tool, artifact, handoff, evaluation, and review fields require separate instrumentation and are not populated by default. Price, fees, payouts, and later human approvals remain separate platform records.
The example uses only fields populated by the current platform sealing paths. It is illustrative, not a record of a real customer run.
{
"receipt_id": "rcpt_01J...",
"schema_version": 1,
"agent_name": "engineering-flaky-test-diagnosis",
"caller": "user:workflow-owner",
"task_id": "case_flaky_test_diagnosis",
"skill_name": "flaky_test_diagnosis",
"input_hash": "4d7c...9a2f",
"grant_ids": [
"grt_case_inputs",
"grt_tool_propose"
],
"status": "ok",
"result_preview": "Output prepared: Failure cluster summary. Human decision remains separate.",
"elapsed_ms": 4218
}Correlate an alert with deploy and service evidence, rank hypotheses, and hand an incident commander a bounded action packet.
Review a proposed change against its tests, ownership boundaries, and operational risk before a maintainer merges it.
Compare an API change to the prior contract, identify affected consumers, and prepare a compatibility decision for an owner.
Deploy the workflow as a bounded internal agent, verify its outputs and the receipt fields actually emitted, then expand only the scopes your acceptance test proves it needs.