Open-source skills · Applied in SEOcluster.ai
AI agent validation & review
Evidence before acceptance. Review before release.
I built an AI development and QA workflow for SEOcluster.ai and published Agent Integrity Skills: reusable instructions for checking agent work, passing evidence between tasks and deciding when a result is ready to act on.
At a glance
- Define the claim and checks
- Review the work and evidence
- Approve the next permitted action
The work
What I built
Three skills, three separate responsibilities
The toolkit separates keeping work running, checking its output and authorising the next action. autonomous-worker-ops supports bounded jobs with status reporting and recovery. dual-agent-review separates the producer from the reviewer. fail-closed-promotion requires evidence before work can move towards release. These are portable instructions and worker templates that need adapting to each project.
Evidence travels with the task
The producer hands over a specific artefact, the proposed next action, forbidden actions, tests or probes, and the likely impact of a mistake. The reviewer examines that package and tries to reproduce a check or find a counterexample. Review ends with acceptance, required changes or a human decision, with a clear statement of what may happen next. Agreement between agents is not enough on its own.
Validation informed by scientific testing
For research and evaluation tasks, the protocol asks whether the test rules were fixed in advance, whether a holdout or separate confirmation window is missing, and whether repeated testing or cherry-picking could explain the result. These checks help distinguish an exploratory finding from evidence strong enough to support a decision. For software tasks, the corresponding evidence may be a reproducible test, a behavioural contract or an observed user journey.
Applied to SEOcluster.ai
In my SEOcluster workflow, one agent implements changes on staging and another tests real user journeys. Findings return to implementation for correction and further review. The workflow has run more than 240 review cycles, with production promotion subject to human approval. This is the product-specific application; the published skills also support reviews of plans, designs and research claims.
What counts as independent review
The protocol distinguishes a reviewer from a different model product or vendor from a same-model review, which is labelled internal QA. It also separates a successful staging check from permission to deploy. A task can pass a check while still lacking the evidence or human approval needed for the next action.
A validation process, not a guarantee
The skills define an evidence-seeking process. They do not provide mathematical proof of correctness or establish that every result has been scientifically validated. The strength of a conclusion still depends on the actual tests, data, review independence and retained evidence. That distinction is part of the design.
Follow the work
Related expertise
Have something similar in mind? Tell me what you need to build.
Discuss your project ↗Explore all work