Atlas

An engineering team built from Claude Code

Every goal gets a line to main.

You talk to one PA. Behind it, specialists each own a line of work in their own Claude Code session. Every task passes the checkpoints its line requires, and finished work rides the landing train to your main branch, tested together.

The Atlas line map Five coloured lines leave the PA: plan, build, prove, guard and keep healthy. Each passes its own checkpoint and meets the others at the gate, where your project's checks run. Work that clears the gate rides the landing train to main; work that fails is sent back down its line.
PA Plan Build Prove Guard Keep healthy Main
Departures · calculatorEvery task has an owner, a line and a state
TimeLineTaskNext stopStatus

An illustration of an ordinary afternoon. In the app this is the live board, and every row links to its task, its owner and its evidence.

15role definitions, each with its own deliverable and the evidence it must produce
26.1 hof work done in 10.6 hours on the clock, by specialists working in parallel (measured on our own projects)
162attack classes on the security engineer's coverage ledger, module by module
6finished branches per landing-train batch, tested together through your own checks

Why not just Claude Code?

Claude Code is a superb engineer. Atlas is the team around it.

Everything Atlas does runs on Claude Code sessions. What it adds is everything a team adds around a good engineer: someone to talk to, colleagues with other jobs, a board that remembers, rules about what "done" means, and a way to land work safely.

What changesClaude Code on its ownWith Atlas

Who you talk to

You drive every session and read every terminal.

One PA. The engineers can't message you; they escalate to PA, which asks you one question at a time with buttons, and merges bursts of messages into one.

How many engineers

One session, or as many terminals as you can juggle.

Specialists in parallel, each in its own Claude Code session and working folder, with model and effort tuned to the role. On our own projects: 26.1 hours of work in 10.6 hours.

What remembers the work

The session's context. Close the laptop or hit a crash and the thread is gone.

A board in a database. Every task moves through states, each claim holds a 45-minute lease, owners are re-pinged every 3 minutes, and engineers checkpoint so a restart resumes where they stopped.

What "done" means

Done when the model says it is.

Completion is refused until the work is on main or queued for the train, a screen has its test cases and screenshots, and the product's rules pass. Otherwise it is recorded as "done, not verified", with what is missing.

What waits for what

You remember the order and enforce it yourself.

A dependency plan. Gates hold building until the architecture is approved, the design is signed off and the failing tests exist; a screen automatically waits for the API it calls.

Review

You review it, when you have time.

Every build that finishes is reviewed at once, and comes back RELEASABLE or with fix tasks filed to its builder by name. Reviews hold the release, not the builder.

Getting work onto main

You merge branches and fix main when they collide.

The landing train merges up to six finished branches and runs your checks once on the result. A red batch is split until the branch that broke it is found and sent back. Main's history is never rewritten.

Getting better

Every session starts from nothing.

Lessons from finished work are given to future engineers. One is held back from each start as a control, and every six hours the lessons that track worse results or higher cost are retired automatically.

Money, secrets, the irreversible

Whatever the session is allowed to do.

Production go-lives, spending, deleting data, sending outside and dropping scope are never decided for you. An optional daily budget warns at 50% and 80% and pauses at 100%. Secrets pasted into chat go to the secret store before PA sees them.

Evidence you can open

A terminal transcript.

A Tests screen with the case library, runs, screenshots and step-by-step replay; a Security screen that leads with what has never been reviewed; a Performance screen with load against your targets.

The team

Every engineer owns a line, and the line ends in evidence.

Each role is its own Claude Code session with a written job: what it does, what it must hand over, and what it must never do. The engine creates the standard roles when work arrives for them; anything else is a specialist PA proposes and you approve.

PA

PA, the one you talk to

The interchange · every line starts here

The single point of contact between you and every engineer. It routes, coordinates, escalates and reports, and never builds anything itself.

  1. Works out whether you want to discuss, plan or have something built, and asks when it can't tell
  2. Confirms scope once, then lays the whole plan as tasks that depend on each other
  3. Sets goals with measurable criteria and opens the gates that apply: architecture, design, test-first
  4. Answers the engineers' questions or brings them to you, one at a time, with buttons
  5. Hands over: one message per reply, milestones with their evidence, estimates only from the forecast

Never: writes code; approves production, spending, data deletion or a change of scope; says "done" while a goal is below 100%.

AR

Architecture

Plan

System design and technical direction, written down before anything is built.

  1. Holds backend, UI and integration work until the design is decided
  2. Writes the design with real choices and numbers: topology, infrastructure, callers, cost, alternatives
  3. On an existing codebase, maps it into the team's shared memory
  4. Hands over: the design as a document with diagrams, for your approval
"Use real choices and numbers, not 'TBD'."

Never: lets building start on a system decision you haven't reviewed.

RG

Research & gap analysis

Plan

Finds the unknowns before they become rework.

  1. Researches and compares technologies, with web access
  2. Measures the gap between what is required and what is built
  3. Runs small proofs of concept and surfaces the unknown unknowns
  4. Hands over: comparison tables and diagrams; open research holds the release
UD

UI design

Plan

The actual design, as screens you can open and see, not a spec about the design.

  1. Renders every screen, including its loading, empty and error states, with real tokens and realistic content
  2. Maps every value on a screen to a real field in the API
  3. Finishes screen by screen, so building can start on the first one
  4. Hands over: mockups for your sign-off; building waits until you give it
"A finished design task is not approval; the operator's sign-off is."
BE

Backend development

Build

The server side: APIs, the data model, the business logic.

  1. Proposes the API contract before implementing it; changing an agreed shape is refused
  2. Proves each feature on new inputs, one for every variant
  3. Builds against the failing tests when the project works test-first
  4. Hands over: working code on its branch, reviewed the moment it finishes

Never: weakens a test to make it pass.

UI

UI development

Build

The real, running screen: the approved design, built faithfully, wired to live data.

  1. Builds every component the mockup shows, with no hard-coded numbers or placeholder text
  2. Checks the agreed API shapes and uses the generated types
  3. Passes a visual check that scores the screen against the design and catches reskins
  4. Hands over: a screenshot of the live screen
"A reskin of the old screen is NOT a build."

Never: fakes data. A missing endpoint becomes a fix task for backend.

SP

Specialists

Build · made for your project

Domain experts created for one project, with a set boundary, skills and a reason to exist.

  1. Proposed by PA; a new kind of role starts only with your approval
  2. Writes the contract first and never weakens a test
  3. Product plugins (Salesforce, for example) bring their own specialists, rules and gates
  4. Hands over: the requirement working on real inputs
"Done means the REQUIREMENT works on real inputs."
TS

Testing

Prove

Owns the quality of the tests themselves.

  1. Writes and shares the test cases before running anything, tagged by technique
  2. Tests the capability on new inputs, not the cached answer
  3. Writes failing tests first; red releases the builder, green releases the release
  4. Grades its own tests with mutation testing
  5. Hands over: cases and runs in the Tests screen; a claim without an artifact doesn't count
"Test the CAPABILITY, not the cache."
QA

Browser QA

Prove · the end-to-end gate

Proves the app works by driving it in a real browser against the running app.

  1. Drives the core user flows with new inputs; any console or network error is a fail
  2. Writes specs for every screen and role, unasked
  3. Backs up the design check route by route
  4. Hands over: runs with screenshots. Completion is refused without them
"'Unit tests pass' is NOT your bar."

Never: fixes the app. It verifies; the builders fix.

PF

Performance

Prove

Answers whether the system takes the load you expect, and where it chokes first.

  1. Turns the load model drafted from the code into your terms: orders per hour, people signed in
  2. Writes load scenarios: a tenth of the load locally, the full numbers on a server
  3. Reviews the code against four sets of performance rules
  4. Hands over: measurements against your targets, and the first bottleneck

Never: points anything at production.

CR

Code review

Guard

Quality, patterns and standards, on every build as it finishes.

  1. Picks up the review Atlas files automatically for each finished build
  2. Treats every rule line as blocking, and files fixes to the builder by name
  3. Checks that the designed feature is really wired, with no silent downgrades or falsely green tests
  4. Hands over: findings by file and line, and a verdict: RELEASABLE or not
CT

Critic

Guard · when you want a second opinion

An adversarial reviewer whose only job is to find what is wrong with plans and claims.

  1. Checks a plan against the project's constraints, goals, decisions and blockers
  2. Names unstated assumptions and the revisions it needs
  3. Hands over: a score out of 10 and a verdict: ship, revise or block
"You exist because Claude is congenitally agreeable."

Never: agrees out of politeness, or invents a flaw to seem useful.

SE

Security

Guard

Vulnerabilities, threat modelling and secure code, from a ledger rather than a hunch.

  1. Works through every module against 162 attack classes
  2. Closes a unit only with evidence; findings are fixed only after a re-test
  3. Asks before it runs: now, at the next build, or hold
  4. Hands over: the coverage ledger and verified findings in the Security screen
"'Looks fine' is not a result."
BD

Build & deploy

Ship

The running app, because a green test suite is not a running app.

  1. Waits until no build, review or fix is open, and the train is idle
  2. Starts the frontend and API on free ports and hands the address to browser QA
  3. Treats a designed component that isn't attached as a blocker
  4. Hands over: the app running, and the deployment document updated

Never: hand-fixes application code.

LT

The landing train

Ship · run by Atlas itself

Not an engineer but the engine: it brings finished branches to main, tested together.

  1. Merges up to six finished branches into one batch
  2. Runs your project's checks once on the result
  3. Splits a red batch until the branch that broke it is found, and sends it back
  4. Hands over: work on main, and a message saying what arrived
DC

Documentation

Keep healthy

Keeps the documents in step with the code.

  1. READMEs, API documents, the changelog and decision records
  2. Runs after testing and before the release
  3. Hands over: documents in the app's Documentation tab
DM

Dependency management

Keep healthy

Keeps the project's dependencies healthy.

  1. Versions, outdated packages and the breaking changes that come with upgrades
  2. Licence compliance and lock-file integrity
  3. Takes the dependency and vulnerability fixes the other engineers route to it
  4. Hands over: every package with its current, latest, advisory and action

One goal, end to end

What happens between "build me this" and main.

The real sequence, as the engine runs it. The state names are the ones you will see on the board.

Intake

PA works out what you're asking for and confirms the scope once. It records the goal with measurable criteria, and asks whether to build test-first or the standard way.

Plan

PA lays the work as tasks that depend on each other, API contracts first, and opens the gates that apply: architecture, design, test-first.

PENDING

Delivery

When a task's dependencies finish, it becomes due. The engine checks the gates, creates the engineer if it doesn't exist yet, respects how many can run on your machine, and rings the owner.

SCHEDULED

Claim and work

The owner claims it under a 45-minute lease, reads the team's shared memory and checkpoints as it goes. Questions travel up through PA and answers come back the same way.

RUNNING

Tests

Test-first: the testing engineer's failing tests turn the gate red, which releases the builder; green releases the release. Otherwise, testing follows the build.

RED → GREEN

Complete, if it really is

Completion checks the work is on main or queued for the train, that a screen has its cases and screenshots, and that the product's rules pass, then records a verdict and releases what was waiting.

COMPLETED · confirmed or needs validation

Review

Each finished build is reviewed at once: RELEASABLE, or fixes filed to its builder. Security (when you consent), performance and code review run alongside.

Land

The landing train brings each feature to main whole. Then build & deploy runs the app and browser QA drives the core flows in a real browser.

You're told

Progress, then a message for each feature that reached main and each release. When a goal is met, its work is merged and tagged only if its suite is green.

Where it runs

The same team on your Mac, or in your company's cloud.

Available now

On your Mac

  • One window, one tab per project. A team runs while its tab is open
  • Pausing waits until every running task has a fresh checkpoint
  • After a restart every project opens paused; nothing starts behind your back
  • How many engineers run at once is sized to your Mac's memory
Coming next

In your company's cloud

  • Installed into your own DigitalOcean account; your code and records stay there
  • Projects keep running with the laptop closed, and sleep after 30 idle minutes
  • A monthly budget cap: a warning at 80%, everything sleeps at 100%
  • A managed database per environment (UAT, testing, production), created when you say so
  • Each project runs on the Claude account of the person who set it up

Bring the goal

Atlas brings the team, and the line to main.