At xcactus, we use agents to analyse requirements, prepare plans, implement changes and review the results. People define the business outcome and approve decisions within their authority. Controllers assign work and check whether it can move to the next stage.
We call this an AI software factory. The term describes how we organise delivery, not a promise that software can be released without human approval.
This article explains the process, the roles and the system components. A worked example connects them. The final section covers shared limitations.
In this article
A request starts with a problem or an expected outcome. Before implementation, we need to know who is affected, what must change, what must stay unchanged and how we will check the result.
People use Slack to provide that context, answer questions and follow progress. They do not need to operate agent terminals or contact individual workers. Plans, code, tests and business records stay in their appropriate systems; Slack is the conversation around the work.
Gather requirements
↓
Classify and scope
↓
Analyse
↓
Plan and approve
↓
Implement and verify
↓
Approve and release
↓
Check outcome and closeThe steps need different amounts of work depending on the request. They do not require a separate meeting or document for every transition.
| Stage | Who acts | Result |
|---|---|---|
| StageGather requirements and classify | Who actsRequester and coordinating assistant | ResultA recorded request with an owner, scope, constraints, acceptance criteria and delivery path. |
| StageAnalyse | Who actsAssigned worker or planner | ResultFindings about the existing system, risks, dependencies and alternatives, including a smaller change or no change. |
| StagePlan and authorise | Who actsPlanner, reviewer where required, and authorised decision-maker | ResultAn agreed approach. Planned changes require an independently reviewed plan and stakeholder approval of its scope. |
| StageImplement | Who actsAssigned worker | ResultChanges within the agreed scope, tests and a report of what was done. |
| StageVerify | Who actsWorker, controller and reviewer where required | ResultEvidence against the acceptance criteria, findings to resolve and any remaining uncertainty. |
| StageRelease and close | Who actsAuthorised decision-maker and assigned executor | ResultSeparate release permission, the applied change, a check in the target environment and a completion record. |
Business participants decide questions that change behaviour, cost, risk, data handling or scope. Routine technical choices stay with the delivery team. A changed proposal may require renewed approval. Implementation complete, approved for release and verified after release are different states.
When we use it: new or materially changed behaviour, architecture or system capability.
How it runs: intake, analysis, design, implementation and verification. Separate planning, plan-review, implementation and code-review roles provide independent scrutiny. The planner examines the existing system and proposes a concrete approach. Implementation follows the reviewed and approved scope.
Required checks: the plan and code are reviewed, findings are resolved and results are checked against the acceptance criteria. Material scope changes, unresolved blockers and exhausted correction cycles return to a human. Integration into the main branch and deployment retain separate approval gates; routine implementation steps do not each need a new business decision.
When we use it: something is broken or behaving incorrectly.
How it runs: triage, diagnosis, the smallest safe fix, verification of recovery and a record of the outcome. One persistent worker owns diagnosis, implementation, regression tests, self-audit and in-scope corrections under workflow supervision. Corrections stay with that worker so context is retained.
Required checks: the controller checks the reported result against recorded changes and verification evidence. This is not a second design or code review. High-impact incidents require a post-mortem; smaller incidents use a proportionate record. A durable remedy requiring a material architectural, data-model, authorisation or cross-repository change becomes a linked CHG, not an unapproved expansion of the incident.
When we use it: a bounded request with a clear expected outcome, such as a defined adjustment or controlled maintenance task. Unlike an incident, it starts with a request rather than a failure.
How it runs: confirm scope and authority, execute, verify and close. One persistent worker handles the request under workflow supervision.
Required checks: the controller checks the returned evidence. The worker's self-check is not independent review. If the request requires a material design change, it is re-scoped before proceeding.
For every path, registering a ticket is not permission to implement, and completing implementation is not permission to deploy. Urgency does not grant production or destructive-action authority.
An agent uses a model and tools to perform an assigned role. A role describes a responsibility, not a product or model name.
| Role | Responsibility |
|---|---|
| Business owner or authorised decision-maker — a person | Defines the outcome, resolves business trade-offs and approves actions within their authority. |
| Coordinating assistant | Clarifies and registers requests, routes work, reports status and brings questions back to people. |
| Workflow controller | Owns the current process state, assigns bounded work, receives results and checks the conditions for proceeding. |
| Analyst / planner | Inspects the problem and existing system, challenges assumptions and prepares an approach with risks and verification steps. |
| Implementer | Makes the agreed change, runs checks and reports the result and remaining uncertainty. |
| Independent reviewer | Examines a plan or implementation against requirements, risks and evidence. Planned changes separate plan review from code review. |
These roles do not require a separate agent for every task. A coordinating assistant can also control a lightweight workflow. One worker can analyse, implement and self-check a bounded request. Planned changes deliberately separate authors from independent reviewers. Human expertise can supplement the work where the risk or subject requires it.
Tool names are defined in the component descriptions below.
Each component has a specific job. The descriptions explain why we need that job, what the component does, how it connects and where its responsibility ends. They are not claims that these products are the only way to build the system.
Slack is the channel people use to discuss requests and decisions.
A business participant should not need to find an execution machine or repeat a decision to several agents. A shared conversation gives them a place to provide requirements, ask for status and resolve questions. The same need exists in business workflows, not only software delivery.
Slack presents questions, options and progress updates. A useful decision request identifies the question, explains the context and consequences, gives a recommendation and states which action is blocked.
Illustrative question — not a client transcript:
Q-EXAMPLE — Accounts without verified addresses
Some existing accounts lack the verified address required by the planned integration.
A — Exclude them and produce an exception report. Keeps the current scope; operations must handle the exceptions.
B — Add address verification. Expands implementation and introduces a new dependency.
C — Pause the affected processing. Avoids acting on incomplete data until the business rule is settled.
Recommendation: A, if partial processing is acceptable.
Blocked: the implementation decision for incomplete accounts.
The coordinating assistant links the conversation to the work record. An answer must be bound to the correct question, workflow, conversation and authorised participant before it can affect execution.
Slack is not the execution engine or the sole system of record. Membership in a channel does not grant release or purchasing authority. An ambiguous “yes” is not blanket permission, and silence is not approval. Another channel could replace Slack if it preserves identity, context, decision routing and traceability.
Hermes is the agent platform behind our coordinating assistant.
Someone must connect a person's request to the right process and return questions and results in terms that person can act on. Requesters should not have to manage individual workers.
The assistant clarifies requirements, helps classify and register work, uses permitted tools, routes assignments and reports status. It brings unresolved business decisions back to an authorised person. It can also control an appropriate workflow directly.
Hermes connects the human conversation with work records and execution workflows. Where XC Bus is used, assignments, questions, answers and delivery state pass through that coordination layer. Workers report through their controller rather than opening competing approval conversations.
The assistant cannot invent approval or extend its authority because an answer seems obvious. A decision it makes within existing authority is attributed to the agent, not presented as human approval. It must distinguish a worker's claim from evidence of completion.
Pi is the agent runtime we use as the controller in the ordinary multi-role CHG path.
Planning, implementation and review need an identified owner of workflow state. Without one, roles can work from different plans or disagree about whether the next stage is authorised.
With our orchestration layer, the controller assigns work, receives results, checks handoffs and coordinates corrections. Review loops have limits; repeated unsuccessful corrections require operator intervention.
The controller coordinates the planner, plan reviewer, implementer and code reviewer in their execution sessions. It uses versioned artifacts for handoffs and returns questions and status through the coordination path.
A controller does not replace a reviewer or a human approval. Not every workflow uses Pi or needs all four roles. The authorised process determines the controller; there must be one clearly identified controller for a workflow, not competing controllers.
Claude Code and Codex are agent environments used to perform assigned work. Claude names a model family; the worker's responsibility comes from its assigned role.
Agents can inspect code and requirements, propose changes, use engineering tools and examine results. Keeping their assignments bounded makes their work easier to review and correct.
A planner prepares an approach, an implementer changes the system and runs tests, and a reviewer examines the plan or code. Claude Code and Codex can serve these different roles; a product name does not determine whether the agent is an author or reviewer.
Workers receive a task contract, project context and the relevant plan or code revision from the controller. They return findings, actual check results, changed artifacts and blockers. In-scope corrections return to the same live role so it retains context.
Workers cannot silently replace an approved plan, widen scope or authorise release. They must report tests they could not run and evidence they could not obtain. An independently assigned reviewer can still share the author's mistaken assumption.
Herdr provides the workspaces and agent sessions in which execution roles operate.
Concurrent roles need identifiable sessions and separate places to work. Corrections need to return to the right role rather than a newly launched agent assumed to have the same context.
Herdr organises workspaces, panes and sessions. Our orchestration layer adds controlled launches, links roles to their sessions, tracks workflow state and checks handoffs. Roles work in isolated Git worktrees.
The controller assigns work to roles in those sessions. Workers produce versioned artifacts that can be handed to another role for review.
A workspace manager is not a model or a reviewer. Saving a role identifier does not restore a lost session. Recovery must establish which execution is current, what completed and which evidence remains valid. Separate workspaces alone do not provide security isolation.
AI rules define common working practices. A task contract defines one worker's assignment.
Copied prompts drift between projects. Broad instructions such as “fix the feature” also leave the worker to guess scope and completion criteria. Shared rules and bounded assignments reduce those ambiguities.
Versioned rules cover engineering principles, process management, worker context, runtime safety, communication, project memory and delivery metrics. They tell agents to challenge requirements, define success, inspect existing code and conventions, remove unnecessary work, make the smallest sufficient change, test failure cases and report uncertainty.
The task contract names the objective, inputs, scope, exclusions, completion conditions, stop conditions and required report. This simplified example is not a literal API schema:
role: implementer
objective:
Prevent duplicate order creation when a request is retried.
inputs:
- approved implementation plan
- existing order API and tests
scope:
- order creation path
- related regression tests
out_of_scope:
- payment redesign
- unrelated refactoring
- deployment
done_when:
- retry behaviour matches the approved requirement
- relevant tests have been executed
- the change is published for review
- verification gaps are reported
stop_when:
- the fix requires a broader data-model change
- the approved plan cannot be followed safely
- required evidence cannot be obtained
report:
- branch and exact revision
- changed behaviour
- checks and actual results
- remaining risksTool-specific entry points reference shared instructions through AGENTS.md rather than maintaining separate policies. Repositories receive managed shared rules while AGENTS.project.md retains local context and conventions, protected from shared-rule updates.
Project memory records architecture, constraints and progress. Workers update it only when assigned that responsibility, to avoid conflicting concurrent accounts. Rule updates follow:
Inspect target → Check drift → Review diff → Approve → ApplyInstructions do not enforce themselves. Import behaviour differs between tools, so a reference does not prove that a worker received a rule. Required contracts must be present in the actual execution context. Consistency checks can detect missing clauses or mismatched declarations; they do not prove compliance.
A task contract cannot grant authority absent from the governing ticket. That authority comes from trusted launch context and recorded decisions, not arbitrary instructions in repository content. Workers report blockers to the controller instead of restarting intake or seeking a different approval channel. Memory supports continuity but does not replace code, test results or decisions as evidence.
Git and GitHub hold versioned work that roles can inspect and reference precisely. Our handoff code checks the required approval evidence against that work.
A reviewer and an implementer must work from the same plan revision. Approval of an earlier version cannot silently cover a later change.
Git records revisions, and GitHub makes remote branches and reviewable artifacts available for handoffs. Our handoff checks require approval of the exact plan revision and evidence that the reviewer was launched against it.
Simplified check before implementation:
before starting implementation:
require recorded plan
require approval for that exact plan revision
require recorded reviewer launch for that revision
require remote branch still matches the revision
if any requirement is missing:
refuse the handoff
report the failed conditionThe planner provides a versioned plan, the reviewer examines that revision and the implementer receives the approved version. Reports identify the branch, exact revision, changes, checks and remaining risks. If the plan changes after approval, the earlier approval does not automatically cover it. The same principle applies to code reviews and business decisions.
These are checks in our orchestration code, not guarantees supplied by Git alone. They can block a handoff with missing or stale evidence, but cannot establish whether the review was insightful or the plan solves the right problem. A reviewable change is not permission to merge or deploy.
XC Bus carries assignments, questions, answers and status between central coordination and execution stations. NATS JetStream provides persistent message delivery and redelivery.
A worker may ask a question while a decision-maker is unavailable. A station may disconnect before the answer arrives. The system needs to retain the answer and deliver it to the correct workflow when the station reconnects.
The XC Bus hub records assignments, questions and delivery state. Station runtimes handle local delivery and persistence. NATS carries the messages. Receiving, storing and acknowledging a message is deterministic code; it does not wait for a model to decide whether the message should be saved.
Question state distinguishes delivery from application:
Open → Presented → Answered → AppliedIf a station is offline, a recorded answer remains answered, not applied. It becomes applied when the workflow handles it and returns an acknowledgement. The hub prevents a later submission from silently replacing an accepted answer; changing a decision requires an explicit follow-up.
The ordinary Pi-controlled path is shown below. Other authorised workflows may use a different controller.
Slack ↔ Hermes ↔ XC Bus hub ↔ NATS JetStream ↔ Station runtime
↕
local inbox / outbox
↕
Workflow Pi
↕
WorkersHermes communicates with the hub over HTTP. The controller owns workflow progression; the bus delivers coordination messages.
Broker acceptance, application storage, an applied decision and completed work are different outcomes. Even an applied answer does not prove implementation is complete.
Delivery can happen more than once. A lost acknowledgement does not prove that an action failed. Consequential operations need idempotency (safe handling of repeated requests), reconciliation or duplicate detection where the business effect occurs. Retrying “send payment” is not equivalent to retrying “read status”.
XC Bus does not make business decisions or provide worker isolation. It should carry typed assignments and decisions, not become an unrestricted remote shell.
A Git worktree separates working copies. A sandbox restricts the environment in which a worker executes.
Concurrent workers should not overwrite one another's files. Separately, code execution may need restrictions on filesystem access, network destinations and credentials. These are different problems.
Worktrees separate code changes. Containers, virtual machines, microVMs or managed sandboxes can host workers with configured limits on filesystem and network access, credentials and their lifetime, resources, artifact export, teardown and retention. A worktree can exist inside a sandbox.
Workers can run outside a developer's host while assignments, questions, answers and status are coordinated centrally through XC Bus. Each environment needs a compatible station or execution adapter, a supported launch path, reliable artifact exchange and appropriate permissions. Central coordination does not require central possession of every worker's secrets.
A worktree does not restrict network access, protect host credentials or contain arbitrary code. A sandbox with broad host mounts, privileged access, shared credentials or unrestricted production connectivity can undermine its own isolation.
Sandbox placement is an execution option, not a claim of automatic compatibility with every provider. Command guards can block selected dangerous operations, but do not prove containment against malicious agents or untrusted code. Instructions, command guards, operating-system isolation and external access controls are not interchangeable.
Consider a request to prevent duplicate orders when a request is retried. This is an illustrative walkthrough, not a client delivery report.
The same pattern can support a business workflow. In an illustrative supplier-offer process, AI extracts terms and flags discrepancies; deterministic code calculates totals and checks purchasing thresholds. A business owner resolves an exception in Slack, a scoped integration creates the authorised record, and a readback verifies it. This is an application example, not a claim about a deployed client system.
The agent does not need broad ERP privileges to explain a purchasing recommendation. If the supplier, price or order contents change while a decision is pending, the system must reassess whether the earlier approval still applies.