01 · DesignDescribe an idea. The agent lays out the alternatives, their tradeoffs and a recommendation, then turns the plan you approve into tasks.
The AI software factory
Signals in. Software out.
ReasonOS is the software factory for teams that build with agents. Ideas, issues and failing tests go in. Reviewed, tested, shipped changes come out, and every cycle leaves the next one better equipped.
Runs the agents you already useClaude CodeCodexGeminiReasonOS agent
- 01SignalQA agentcheckout-smoke failed at step 6: the card form retried a 503 eleven times
- 02DesignAgent + SarahThree alternatives with tradeoffs. Recommended: a retry budget per request
- 03PlanTeam decisionPlan approved, split into four tasks: two for agents, two for people
- 04BuildClaude Code + Priyafeat/retry-budget running on its own server, one workspace for the agent and the person
- 05ReviewAgent + SarahChange #418 approved at 9f2c1e4. The approval expires if new code lands
- 06TestQA agentcheckout-smoke passed, 14 of 14 steps, in Chrome, Firefox and WebKit
- 07ShipBuild systemRebuilt the 3 targets the change touched. The other 209 came from cache
- 08LearnEvery agentThe retry lesson is saved with the branch, for the next agent that touches checkout
What one cycle learns, the next one starts with: lessons, tests and every build carry forward.
Why now
Agents now write more code than people can read. The line around them was built for one person and one change at a time.
Clone, install, wait for CI, chase a reviewer, hand off to QA, file a ticket for the deploy. Every step assumed a person was typing. A software factory rebuilds the line around a branch that is already running, so people and agents move at the same speed, under the same rules.
What makes it a factory
Not another tool on the line. The line itself.
Most AI tools speed up one station and leave the hand-offs around it untouched. ReasonOS runs the whole line, so a change moves from signal to production without being re-explained, re-cloned or re-tested from scratch.
- 01
A running line for every branch.
Every branch boots as its own server: cloned, built and running its tests before anyone opens it. People and agents join the same line with a link, and nobody sets up a laptop.
- Its own server per branch
- Join with a link
- Shared files, terminal and preview
- 02
Any agent. Any model.
Claude Code, Codex, Gemini or the ReasonOS agent, chosen per role, on the subscriptions your team already pays for. Each runs inside the branch with the same tools, permissions and checks as a person.
- Claude Code, with your Claude plan
- Codex, with your ChatGPT plan
- Gemini, with your Google account
- ReasonOS agent, with project API keys
- 03
One memory across the lifecycle.
The plan, every conversation, what agents learned and a map of the code live with the branch, committed beside it. Nothing is re-explained between stages and no agent starts from zero.
- The plan and its tasks
- Every agent conversation
- Lessons that carry forward
- A map of every symbol and test
Inside the factory
Every station is the real product. Not a dashboard on top of other tools.
02 · PlanAn issue starts when its branch opens and closes when its change lands.
Work
Cycle 14
ActiveSep 15 – Sep 2803 · BuildPeople and agents work in the same running branch: one workspace, one terminal, one conversation, one shared context.
04 · DelegateHand a coding agent a task in its running branch. It reads the code, makes the change, runs the checks and opens a review. Add specialist roles and independent reviewers when the work needs a team.
05 · ReviewAn approval names the commit it judged, and expires when new code is pushed.
06 · TestA test written in sentences, run by an agent in a real browser on every push — every click drawn where it landed, and a design reviewer’s findings on every screen.
07 · ShipA push builds and tests only what it affects; everything else comes from the cache.
Building blocks
Adopt one block, or the whole factory.
Every part works on its own and gets better beside the others. Start with the build system, QA or review, and add the rest when the line is ready for it.
05Agents
04Intake
03Quality gates
02Delivery
Agentic orchestration
A team of agents, working as one.
Give the work distinct roles, configure the agents and models each role uses, and decide where a person has final authority. Independent work moves in parallel; every agent works from the same task, code and evidence.
One projectOne shared branch and evidence trail
- 01DesignFrames alternatives and a planAvailable
- 02Impact analysisFinds the systems and tests affectedAvailable
- 03Risk assessmentChecks the change independentlyAvailable
- 04CodeMakes the change on its branchAvailable
- 05Test generationAdds checks and edge casesAvailable
- 06Code reviewChecks the change before approvalAvailable
- 07AI test & iterationRuns checks and manages bounded repairsAvailable
One or more agents can work in each role. Their work and evidence stay together, so people and later agents can check it.
Self-improving
A factory that gets better every cycle.
Tools that forget make every change start cold. In ReasonOS what a cycle learns stays in the branch: the next signal is triaged faster, the next agent starts warmer, and the next build is mostly done.
- 01
A failing test files its own proposal.
When a QA run fails because of the application, the triage agent files a proposal in the project that owns the code, with the evidence attached: the steps, the screenshots and the requests behind them.
- 02
Agents remember.
Lessons and preferences are saved with the branch and carried into the next session, so the second agent to touch a service starts where the first one finished.
- 03
Nothing is built twice.
Every build lands in the shared cache. A teammate, an agent or CI that needs the same output gets it in seconds instead of rebuilding it.
- 04
The map keeps up.
Atlas maps every symbol, caller and test in the code, and any agent can query it. Agents spend their turns on the change instead of searching for it.
Outcomes
Measure what ships, not tokens or lines of code.
- 01
Tests passing on every push
Every QA run, its steps and its evidence, per branch and per environment.
Available - 02
Agent spend against budget
Model tokens metered per person and project, against a monthly budget the organization sets.
Available - 03
Build time saved by the cache
What the shared cache served instead of rebuilding, for every person and every run.
Coming soon - 04
Cycle time, signal to production
How long a change takes from the signal that raised it to its release.
Coming soon - 05
Cost per merged change
Agent spend, build minutes and test runs, added up for each change that lands.
Coming soon
Enterprise
Autonomy needs control. Every control is built in, not sold separately.
Single sign-on
SAML or OIDC through your identity provider, or Google, Microsoft and GitHub sign-in. Two-factor is required by default, and people with your verified domain join automatically.
Roles you define
Build roles from allow and deny rules, grant them to nested teams, and protect branches with required reviews, checks and signed commits.
A complete activity log
Every sign-in, push, merge, permission change and secret change, with who did it and whether it was allowed.
Branches with no public address
Every request goes through one authenticated gateway. Secrets are encrypted and never exposed to terminals or agents as environment variables.
Changelog
Shipping every week.
QA
Every click drawn on the page, and a design reviewer on every screen
Open any step of a QA run to see exactly where the agent clicked, drawn on the page it clicked, with a crop of the element beside each action. A design reviewer looks at every screen the run visits and flags what a top-tier team would not ship — misrendered or misaligned elements, confusing flows, sloppy copy — as findings beside the step, numbered on the page. Requests open to their headers, bodies and timing, and each step keeps its console, its document and the browser’s load.
Agents
AI orchestration for role-based agent teams
Set up Design, Risk Assessment, Impact Analysis, Code, Test Generation, Code Review and AI Test & Iteration roles. Assign primary agents and independent reviewers, keep people at decision gates, preserve workflow evidence, and compare recorded outcomes.
QA
A step-by-step view of every QA run
Open any run to see each step’s outcome, its retries and which steps were reused from a previous run. Repeat runs of the same test now behave the same way.
Agents
Claude Code, Codex and Gemini in the editor
Chat with Claude Code, Codex or Gemini inside a branch, approve their actions as they work, and set a default agent for each project.
Start your factory. The first branch runs in seconds.
Open a branch and it boots with your tests already running. Your team and your agents join with a link; keep it as long as the work takes, then throw it away.