The product that improves itself.
Drums watches how people use your product in production: where it fails, where they get stuck, where it could do better. It writes down what it expects before a change ships, makes the change with the coding agent you already use, and measures whether it helped. Every change is proposed to you first. The record is waiting with your coffee.
Observe
Failures and the requests behind them are captured from your running app. Every deploy is recorded, so each failure is traced to the change that shipped it.
Change
Your own coding agent writes the change. Drums checks it against real behaviour, replaying the failure where there is one. Never against the agent's word.
Measure
A small share of traffic first. The outcome is measured once, when the window the change declared closes, and never re-cut afterward. Error rates from Drums' own record, completion from your PostHog.
Bring your own agent. Drums runs the loop.
Claude Code, Codex, Gemini CLI, Cursor, Amp and OpenCode write the change. Drums decides what to change, whether it helped, and whether it ships.
Don't declare it better.
Measure that it helped.
Drums starts by only watching. Nothing changes until you say so.
Every claim says how it is known. Never one green check.
A metric that moved is not proof the change caused it. Drums says how sure it can be, and a full rollout never proves cause.
Measured, not declared
Every change carries the plan it shipped under: the outcome it expects, the baseline, the window. It is measured once, when that window closes, and never re-cut to look better. Drums checks back at 7, 30, and 90 days, because a change that looked good at close can quietly regress later.
Says how it knows, not green checks
Verified, observed, inferred, approved, unresolved. Every claim carries one of the five, and only verified or observed can ship alone.
Git, not another database
Changes land as real commits. The evidence, meaning what was seen, what was changed, how it was
verified and who approved, attaches as a git note. git log is the audit trail.
Earned autonomy, not a settings page
Each failure class climbs from observe to act-alone on its own record, and drops a rung automatically after rollbacks.
Approval, not blanket autonomy
Customer data, billing, permissions, and infrastructure always wait for a named person.
Agent-neutral, no lock-in
Claude Code and Codex write the repair today, OpenCode next. Switching agents changes nothing about the record.
Reproduce
A failing request is captured from your running app and tied to the deploy that shipped just before it. Then Drums makes the failure happen again, so the link is a fact, not a guess.
Detected from telemetry
An observation arrives over an event webhook, with the evidence attached. Drums reads your PostHog to measure outcomes today; Sentry and OpenTelemetry are next.
Attributed to a deploy
The failing stack trace is matched against the deploy that shipped just before it and the files it changed.
Rebuilt at that revision
An isolated copy of the app is rebuilt at the exact commit the failure first appeared on.
Replayed, not guessed
The captured request runs again. If it doesn't fail, Drums says so instead of writing a patch.
Intent carried forward
The PR body, the linked ticket, or the agent prompt behind the change is kept as what it was supposed to do.
Repair
Your own coding agent writes the change. Drums hands it the observation, the hypothesis, and the acceptance criteria, then checks the result against real behaviour, replaying the original failure where there is one. Never against the agent's report.
Rolling out with design partners now; the docs say exactly what ships today.
Release
The change goes to a small share of traffic first and is measured against the observation that prompted it: the failure gone, the number moved. Then it is promoted or rolled back, and shipped changes stay reversible with one command for 30 days.
In design; the docs say exactly what ships today.
Every step, decision, and revert lands in the record.
Deployed by your systems.
Drums drives the deploy platform you already run. It never becomes your deploy target, and a held promotion cannot be silently overridden.
Stop and undoBuilt so trust survives the first mistake.
Every claim says how it is known. Autonomy is earned per failure class. One command stops everything, and approval is required where a mistake would cost the most.
Every claim says how it knows
Verified, observed, inferred, approved, or unresolved. Never one green check.
Boundaries in plain language
Proposed from your reverts, code owners, and migration history. Corrected in a sentence, not a config file.
Git is the record
Changes land as real commits. The evidence, meaning what was seen, what was changed, how it was verified and who approved, attaches as a git note that travels with the repository.
Approval where it counts
Customer data, billing, permissions, and infrastructure always wait for a named person.
Verified against the failure
The original failing request, replayed against the repair until the failure is gone. Not the agent's summary.
Reproduction inside verification
An isolated copy of that exact revision, the captured request replayed. The guess becomes a fact before anything is written.
Canary before promotion
A small share of traffic first. Promotion waits until the original failure is gone and nothing new appears.
One command to stop
drums stop pauses every environment; drums undo --since reverses what
shipped inside a window.
Shadow mode
Repairs generated and verified in isolation, never shipped, so you can compare them against what your team actually did.
Autonomy earned per class
Observe → shadow → propose → act alone, promoted on track record and demoted automatically after rollbacks.
One binary
A CLI today; a git remote and an MCP tool inside Claude Code and Codex are in design. No new place to go.
Built on your systems
A webhook in today; Sentry, PostHog, and OpenTelemetry adapters next. Your existing deploy platform out. Drums never becomes your deploy target.
Start free. Pay when it works for you.
The whole loop short of acting alone is free on your own machine. Paid plans add the hosted record, your team, and rollouts that measure themselves. Drums is invite-only right now: every paid seat starts with a call.
Developer
The loop on your own machine
- Get started with:
- Unlimited repositories
- Observe, reproduce, propose
- Your own coding agent writes the change
- Every claim labeled with how it is known
- The record stays local
Pro
For teams running it in production
- Everything in Developer, plus:
- Hosted record and console for the team
- Approvals by named people, in Slack
- Canary rollouts, measured windows, revisits
- Act-alone authority, earned per failure class
Enterprise
For organizations with boundaries
- Everything in Pro, plus:
- Private deployment, your infrastructure
- Identity-bound approvals and SSO
- Signed audit export
- Volume pricing across applications
The engineer is on the loop, not in it.
Drums proposes by default; you approve. Every shipped change stays one command away from undone.
The self-improving product, answered plainly
What Drums does, what it refuses to do on its own, and where it stops.
What is a self-improving product?
How does Drums fix production bugs automatically?
Does Drums deploy to production without approval?
drums authority promote, and automatic demotion the first time something goes wrong.
There is no flag that skips the ladder.