Contact Us
Thank you! Your message has been received!
Ok
Oops! Something went wrong while submitting the form.

Beyond coding agents: how we structure an AI software factory

At xcactus, we use agents to analyse requirements, prepare plans, implement changes and review the results. People define the business outcome and approve decisions within their authority. Controllers assign work and check whether it can move to the next stage.

We call this an AI software factory. The term describes how we organise delivery, not a promise that software can be released without human approval.

This article explains the process, the roles and the system components. A worked example connects them. The final section covers shared limitations.

1. How work moves through the system

A request starts with a problem or an expected outcome. Before implementation, we need to know who is affected, what must change, what must stay unchanged and how we will check the result.

People use Slack to provide that context, answer questions and follow progress. They do not need to operate agent terminals or contact individual workers. Plans, code, tests and business records stay in their appropriate systems; Slack is the conversation around the work.

Gather requirements
↓
Classify and scope
↓
Analyse
↓
Plan and approve
↓
Implement and verify
↓
Approve and release
↓
Check outcome and close

The steps need different amounts of work depending on the request. They do not require a separate meeting or document for every transition.

StageWho actsResult
StageGather requirements and classifyWho actsRequester and coordinating assistantResultA recorded request with an owner, scope, constraints, acceptance criteria and delivery path.
StageAnalyseWho actsAssigned worker or plannerResultFindings about the existing system, risks, dependencies and alternatives, including a smaller change or no change.
StagePlan and authoriseWho actsPlanner, reviewer where required, and authorised decision-makerResultAn agreed approach. Planned changes require an independently reviewed plan and stakeholder approval of its scope.
StageImplementWho actsAssigned workerResultChanges within the agreed scope, tests and a report of what was done.
StageVerifyWho actsWorker, controller and reviewer where requiredResultEvidence against the acceptance criteria, findings to resolve and any remaining uncertainty.
StageRelease and closeWho actsAuthorised decision-maker and assigned executorResultSeparate release permission, the applied change, a check in the target environment and a completion record.

Business participants decide questions that change behaviour, cost, risk, data handling or scope. Routine technical choices stay with the delivery team. A changed proposal may require renewed approval. Implementation complete, approved for release and verified after release are different states.

CHG — a planned change

When we use it: new or materially changed behaviour, architecture or system capability.

How it runs: intake, analysis, design, implementation and verification. Separate planning, plan-review, implementation and code-review roles provide independent scrutiny. The planner examines the existing system and proposes a concrete approach. Implementation follows the reviewed and approved scope.

Required checks: the plan and code are reviewed, findings are resolved and results are checked against the acceptance criteria. Material scope changes, unresolved blockers and exhausted correction cycles return to a human. Integration into the main branch and deployment retain separate approval gates; routine implementation steps do not each need a new business decision.

INC — an incident

When we use it: something is broken or behaving incorrectly.

How it runs: triage, diagnosis, the smallest safe fix, verification of recovery and a record of the outcome. One persistent worker owns diagnosis, implementation, regression tests, self-audit and in-scope corrections under workflow supervision. Corrections stay with that worker so context is retained.

Required checks: the controller checks the reported result against recorded changes and verification evidence. This is not a second design or code review. High-impact incidents require a post-mortem; smaller incidents use a proportionate record. A durable remedy requiring a material architectural, data-model, authorisation or cross-repository change becomes a linked CHG, not an unapproved expansion of the incident.

SR — a service request

When we use it: a bounded request with a clear expected outcome, such as a defined adjustment or controlled maintenance task. Unlike an incident, it starts with a request rather than a failure.

How it runs: confirm scope and authority, execute, verify and close. One persistent worker handles the request under workflow supervision.

Required checks: the controller checks the returned evidence. The worker's self-check is not independent review. If the request requires a material design change, it is re-scoped before proceeding.

For every path, registering a ticket is not permission to implement, and completing implementation is not permission to deploy. Urgency does not grant production or destructive-action authority.

2. Who does what

An agent uses a model and tools to perform an assigned role. A role describes a responsibility, not a product or model name.

RoleResponsibility
Business owner or authorised decision-maker — a personDefines the outcome, resolves business trade-offs and approves actions within their authority.
Coordinating assistantClarifies and registers requests, routes work, reports status and brings questions back to people.
Workflow controllerOwns the current process state, assigns bounded work, receives results and checks the conditions for proceeding.
Analyst / plannerInspects the problem and existing system, challenges assumptions and prepares an approach with risks and verification steps.
ImplementerMakes the agreed change, runs checks and reports the result and remaining uncertainty.
Independent reviewerExamines a plan or implementation against requirements, risks and evidence. Planned changes separate plan review from code review.

These roles do not require a separate agent for every task. A coordinating assistant can also control a lightweight workflow. One worker can analyse, implement and self-check a bounded request. Planned changes deliberately separate authors from independent reviewers. Human expertise can supplement the work where the risk or subject requires it.

Terms used below

  • SDLC: software development lifecycle, from understanding a need through release and maintenance.
  • Workflow: one identifiable piece of work moving through defined stages, with an owner, decisions and evidence.
  • Worker: an agent assigned part of that work, such as planning, implementation or review.
  • Handoff: a transfer between roles or stages, including the context and evidence needed to continue.
  • Execution station: a machine or runtime that hosts workers or connects to their execution environment.
  • Approval gate: a point where the required permission must be established before an action.
  • Exact revision: the specific version of a plan or code change being reviewed or approved, rather than whichever version happens to be current later.

Tool names are defined in the component descriptions below.

3. System components

Each component has a specific job. The descriptions explain why we need that job, what the component does, how it connects and where its responsibility ends. They are not claims that these products are the only way to build the system.

Human communication — Slack

Slack is the channel people use to discuss requests and decisions.

Why we use it

A business participant should not need to find an execution machine or repeat a decision to several agents. A shared conversation gives them a place to provide requirements, ask for status and resolve questions. The same need exists in business workflows, not only software delivery.

What it does

Slack presents questions, options and progress updates. A useful decision request identifies the question, explains the context and consequences, gives a recommendation and states which action is blocked.

Illustrative question — not a client transcript:

Q-EXAMPLE — Accounts without verified addresses

Some existing accounts lack the verified address required by the planned integration.

A — Exclude them and produce an exception report. Keeps the current scope; operations must handle the exceptions.

B — Add address verification. Expands implementation and introduces a new dependency.

C — Pause the affected processing. Avoids acting on incomplete data until the business rule is settled.

Recommendation: A, if partial processing is acceptable.

Blocked: the implementation decision for incomplete accounts.

How it connects

The coordinating assistant links the conversation to the work record. An answer must be bound to the correct question, workflow, conversation and authorised participant before it can affect execution.

Limits

Slack is not the execution engine or the sole system of record. Membership in a channel does not grant release or purchasing authority. An ambiguous “yes” is not blanket permission, and silence is not approval. Another channel could replace Slack if it preserves identity, context, decision routing and traceability.

Intake and coordination — Hermes

Hermes is the agent platform behind our coordinating assistant.

Why we use it

Someone must connect a person's request to the right process and return questions and results in terms that person can act on. Requesters should not have to manage individual workers.

What it does

The assistant clarifies requirements, helps classify and register work, uses permitted tools, routes assignments and reports status. It brings unresolved business decisions back to an authorised person. It can also control an appropriate workflow directly.

How it connects

Hermes connects the human conversation with work records and execution workflows. Where XC Bus is used, assignments, questions, answers and delivery state pass through that coordination layer. Workers report through their controller rather than opening competing approval conversations.

Limits

The assistant cannot invent approval or extend its authority because an answer seems obvious. A decision it makes within existing authority is attributed to the agent, not presented as human approval. It must distinguish a worker's claim from evidence of completion.

Workflow control — Pi

Pi is the agent runtime we use as the controller in the ordinary multi-role CHG path.

Why we use it

Planning, implementation and review need an identified owner of workflow state. Without one, roles can work from different plans or disagree about whether the next stage is authorised.

What it does

With our orchestration layer, the controller assigns work, receives results, checks handoffs and coordinates corrections. Review loops have limits; repeated unsuccessful corrections require operator intervention.

How it connects

The controller coordinates the planner, plan reviewer, implementer and code reviewer in their execution sessions. It uses versioned artifacts for handoffs and returns questions and status through the coordination path.

Limits

A controller does not replace a reviewer or a human approval. Not every workflow uses Pi or needs all four roles. The authorised process determines the controller; there must be one clearly identified controller for a workflow, not competing controllers.

Planning, implementation and review — coding agents

Claude Code and Codex are agent environments used to perform assigned work. Claude names a model family; the worker's responsibility comes from its assigned role.

Why we use them

Agents can inspect code and requirements, propose changes, use engineering tools and examine results. Keeping their assignments bounded makes their work easier to review and correct.

What they do

A planner prepares an approach, an implementer changes the system and runs tests, and a reviewer examines the plan or code. Claude Code and Codex can serve these different roles; a product name does not determine whether the agent is an author or reviewer.

How they connect

Workers receive a task contract, project context and the relevant plan or code revision from the controller. They return findings, actual check results, changed artifacts and blockers. In-scope corrections return to the same live role so it retains context.

Limits

Workers cannot silently replace an approved plan, widen scope or authorise release. They must report tests they could not run and evidence they could not obtain. An independently assigned reviewer can still share the author's mistaken assumption.

Agent workspaces and sessions — Herdr

Herdr provides the workspaces and agent sessions in which execution roles operate.

Why we use it

Concurrent roles need identifiable sessions and separate places to work. Corrections need to return to the right role rather than a newly launched agent assumed to have the same context.

What it does

Herdr organises workspaces, panes and sessions. Our orchestration layer adds controlled launches, links roles to their sessions, tracks workflow state and checks handoffs. Roles work in isolated Git worktrees.

How it connects

The controller assigns work to roles in those sessions. Workers produce versioned artifacts that can be handed to another role for review.

Limits

A workspace manager is not a model or a reviewer. Saving a role identifier does not restore a lost session. Recovery must establish which execution is current, what completed and which evidence remains valid. Separate workspaces alone do not provide security isolation.

Shared instructions — AI rules and task contracts

AI rules define common working practices. A task contract defines one worker's assignment.

Why we use them

Copied prompts drift between projects. Broad instructions such as “fix the feature” also leave the worker to guess scope and completion criteria. Shared rules and bounded assignments reduce those ambiguities.

What they do

Versioned rules cover engineering principles, process management, worker context, runtime safety, communication, project memory and delivery metrics. They tell agents to challenge requirements, define success, inspect existing code and conventions, remove unnecessary work, make the smallest sufficient change, test failure cases and report uncertainty.

The task contract names the objective, inputs, scope, exclusions, completion conditions, stop conditions and required report. This simplified example is not a literal API schema:

role: implementer

objective:
  Prevent duplicate order creation when a request is retried.

inputs:
  - approved implementation plan
  - existing order API and tests

scope:
  - order creation path
  - related regression tests

out_of_scope:
  - payment redesign
  - unrelated refactoring
  - deployment

done_when:
  - retry behaviour matches the approved requirement
  - relevant tests have been executed
  - the change is published for review
  - verification gaps are reported

stop_when:
  - the fix requires a broader data-model change
  - the approved plan cannot be followed safely
  - required evidence cannot be obtained

report:
  - branch and exact revision
  - changed behaviour
  - checks and actual results
  - remaining risks

How they connect

Tool-specific entry points reference shared instructions through AGENTS.md rather than maintaining separate policies. Repositories receive managed shared rules while AGENTS.project.md retains local context and conventions, protected from shared-rule updates.

Project memory records architecture, constraints and progress. Workers update it only when assigned that responsibility, to avoid conflicting concurrent accounts. Rule updates follow:

Inspect target → Check drift → Review diff → Approve → Apply

Limits

Instructions do not enforce themselves. Import behaviour differs between tools, so a reference does not prove that a worker received a rule. Required contracts must be present in the actual execution context. Consistency checks can detect missing clauses or mismatched declarations; they do not prove compliance.

A task contract cannot grant authority absent from the governing ticket. That authority comes from trusted launch context and recorded decisions, not arbitrary instructions in repository content. Workers report blockers to the controller instead of restarting intake or seeking a different approval channel. Memory supports continuity but does not replace code, test results or decisions as evidence.

Versioned work and approval checks — Git/GitHub

Git and GitHub hold versioned work that roles can inspect and reference precisely. Our handoff code checks the required approval evidence against that work.

Why we use them

A reviewer and an implementer must work from the same plan revision. Approval of an earlier version cannot silently cover a later change.

What they do

Git records revisions, and GitHub makes remote branches and reviewable artifacts available for handoffs. Our handoff checks require approval of the exact plan revision and evidence that the reviewer was launched against it.

Simplified check before implementation:

before starting implementation:

  require recorded plan
  require approval for that exact plan revision
  require recorded reviewer launch for that revision
  require remote branch still matches the revision

  if any requirement is missing:
      refuse the handoff
      report the failed condition

How they connect

The planner provides a versioned plan, the reviewer examines that revision and the implementer receives the approved version. Reports identify the branch, exact revision, changes, checks and remaining risks. If the plan changes after approval, the earlier approval does not automatically cover it. The same principle applies to code reviews and business decisions.

Limits

These are checks in our orchestration code, not guarantees supplied by Git alone. They can block a handoff with missing or stale evidence, but cannot establish whether the review was insightful or the plan solves the right problem. A reviewable change is not permission to merge or deploy.

Communication between execution stations — XC Bus and NATS

XC Bus carries assignments, questions, answers and status between central coordination and execution stations. NATS JetStream provides persistent message delivery and redelivery.

Why we use them

A worker may ask a question while a decision-maker is unavailable. A station may disconnect before the answer arrives. The system needs to retain the answer and deliver it to the correct workflow when the station reconnects.

What they do

The XC Bus hub records assignments, questions and delivery state. Station runtimes handle local delivery and persistence. NATS carries the messages. Receiving, storing and acknowledging a message is deterministic code; it does not wait for a model to decide whether the message should be saved.

Question state distinguishes delivery from application:

Open → Presented → Answered → Applied

If a station is offline, a recorded answer remains answered, not applied. It becomes applied when the workflow handles it and returns an acknowledgement. The hub prevents a later submission from silently replacing an accepted answer; changing a decision requires an explicit follow-up.

How they connect

The ordinary Pi-controlled path is shown below. Other authorised workflows may use a different controller.

Slack ↔ Hermes ↔ XC Bus hub ↔ NATS JetStream ↔ Station runtime
                                                     ↕
                                            local inbox / outbox
                                                     ↕
                                              Workflow Pi
                                                     ↕
                                                  Workers

Hermes communicates with the hub over HTTP. The controller owns workflow progression; the bus delivers coordination messages.

Limits

Broker acceptance, application storage, an applied decision and completed work are different outcomes. Even an applied answer does not prove implementation is complete.

Delivery can happen more than once. A lost acknowledgement does not prove that an action failed. Consequential operations need idempotency (safe handling of repeated requests), reconciliation or duplicate detection where the business effect occurs. Retrying “send payment” is not equivalent to retrying “read status”.

XC Bus does not make business decisions or provide worker isolation. It should carry typed assignments and decisions, not become an unrestricted remote shell.

Execution isolation — worktrees and sandboxes

A Git worktree separates working copies. A sandbox restricts the environment in which a worker executes.

Why we use them

Concurrent workers should not overwrite one another's files. Separately, code execution may need restrictions on filesystem access, network destinations and credentials. These are different problems.

What they do

Worktrees separate code changes. Containers, virtual machines, microVMs or managed sandboxes can host workers with configured limits on filesystem and network access, credentials and their lifetime, resources, artifact export, teardown and retention. A worktree can exist inside a sandbox.

How they connect

Workers can run outside a developer's host while assignments, questions, answers and status are coordinated centrally through XC Bus. Each environment needs a compatible station or execution adapter, a supported launch path, reliable artifact exchange and appropriate permissions. Central coordination does not require central possession of every worker's secrets.

Limits

A worktree does not restrict network access, protect host credentials or contain arbitrary code. A sandbox with broad host mounts, privileged access, shared credentials or unrestricted production connectivity can undermine its own isolation.

Sandbox placement is an execution option, not a claim of automatic compatibility with every provider. Command guards can block selected dangerous operations, but do not prove containment against malicious agents or untrusted code. Instructions, command guards, operating-system isolation and external access controls are not interchangeable.

4. How the components work together

Consider a request to prevent duplicate orders when a request is retried. This is an illustrative walkthrough, not a client delivery report.

  1. Define the outcome. The requester describes the problem in Slack. The coordinating assistant clarifies what counts as a duplicate, what behaviour must remain unchanged and how the result will be checked. The work is classified and recorded. An existing failure may follow the incident path; a planned capability change follows CHG.
  2. Analyse and plan. For the planned-change path, the controller assigns a planner to inspect the existing order API and tests. Payment-system redesign remains outside scope. A separate reviewer examines the proposed plan, and the authorised stakeholder approves its scope.
  3. Check the handoff. The handoff code checks the exact plan revision and its approval evidence before implementation starts. The worker receives the task contract shown above, in an identified execution session and separate worktree.
  4. Execute or raise a blocker. The worker changes the order path and runs the relevant tests. If a broader data-model change is needed, it stops and reports the finding. The controller returns the question through the coordination path. XC Bus retains a recorded answer while a station is offline; delivery alone does not close the task.
  5. Review the result. The worker reports its branch, exact revision, check results and gaps. The code reviewer examines the change. Corrections return to the same responsible role, within the review limits.
  6. Release separately. The result goes back to the authorised person with evidence and remaining risks. Integration and deployment require their own approvals. The completion record distinguishes preparation from release and verification in the target environment.

The same pattern can support a business workflow. In an illustrative supplier-offer process, AI extracts terms and flags discrepancies; deterministic code calculates totals and checks purchasing thresholds. A business owner resolves an exception in Slack, a scoped integration creates the authorised record, and a readback verifies it. This is an application example, not a claim about a deployed client system.

The agent does not need broad ERP privileges to explain a purchasing recommendation. If the supplier, price or order contents change while a decision is pending, the system must reassess whether the earlier approval still applies.

5. Shared limitations

Review and tests can miss the requirement. A separate reviewer can share the author's blind spot. Passing tests can confirm the implementation without establishing that it solves the business problem. Representative tests, security analysis and human judgment remain necessary.

Models should not decide mechanical checks. Models help interpret requirements and explore alternatives. Permission checks, routing, retries and mechanical validation should run in code rather than depend on a model improvising the right answer.

Records do not certify compliance. Decisions and execution evidence support accountability, but do not automatically satisfy contractual, security or regulatory obligations. Retention, access control, data residency and separation of duties depend on the organisation and use case.

To apply this approach elsewhere, start with one workflow. Define what the worker may change, who may approve it, what evidence is needed to proceed and how the result will be checked. Add roles, messaging infrastructure and sandboxes only where they solve a necessary problem. You do not need our full stack to use those principles.

By Maciej Sowa, Founder & CTO, xcactus.

Talk to xcactus about designing an AI-assisted delivery process or a business workflow with clear responsibilities, approval gates and verifiable results.