Reliability
Why a 2xx response isn't verification
An accepted API call isn't proof the change exists. How Ahena reads providers back, tells VERIFIED from APPLIED_UNVERIFIED, and retries safely.
In short
A 2xx status means the provider accepted a request, not that the change you wanted exists. After every write Ahena reads the provider back and compares what it finds with the plan. Each operation ends VERIFIED (seen), APPLIED_UNVERIFIED (accepted but not seen yet), FAILED (not made) or UNKNOWN_OUTCOME (Ahena can't tell), and a plan is APPLIED only when everything that ran was verified. Retries use idempotency keys where the provider offers them and read-match-create where it doesn't, so a retry doesn't create things twice.
What a 2xx actually tells you
An HTTP 200 or 201 from a provider's API means the provider accepted your request. It's tempting to treat that as "done", and most scripts do. But several ordinary situations break the equation:
- Eventual consistency. Many APIs acknowledge a write before every reader can see it. A read immediately afterwards may not show the change yet.
- Asynchronous work. Some writes start a process rather than finish one. Asking an email provider to verify a domain succeeds at once; whether the domain is verified depends on DNS propagation that happens later.
- Long-running operations. Some creates return an operation to poll, not the resource itself.
- Lost responses. A request can time out, or the connection can drop, after the provider has already made the change. From the caller's side, the outcome is unknown.
- Accepted isn't equal. The provider may normalise or partially apply what you sent, so "accepted" and "matches what you meant" are different statements.
None of this is exotic. It's why infrastructure tools that report success from status codes alone eventually report a change that isn't there, or create something twice on a retry.
Read-back verification
Ahena treats a provider's "OK" as a claim to check. After every write, it reads the provider back and compares what it finds with what was planned. In practice the test is simple: Ahena re-plans the change, and if the re-read no longer plans it, the desired state is there. For providers that are eventually consistent, it makes a few attempts with bounded polling before deciding.
Long-running operations are followed to the end. Registering an app with Firebase returns an operation, which Ahena polls with bounded backoff (about 20 seconds) before the re-read decides whether the app exists.
Verification also depends on reading everything. When Ahena checks whether a webhook endpoint, a DNS record or an app already exists, it reads the full listing, across every page. If a listing can't be read completely (it's over the page cap, or a cursor repeats), Ahena raises an error rather than concluding the resource is absent. It never decides "absent" from a partial list, because that's how duplicates get created.
VERIFIED, APPLIED_UNVERIFIED, FAILED, UNKNOWN_OUTCOME
Each operation in an apply ends in one state. The CLI shows a check mark only for VERIFIED, and the dashboard's plan page uses the same states (in plain words: verified, applied but unverified, failed, unknown).
| State | Meaning |
|---|---|
| VERIFIED | Ahena read the provider back and saw the desired state. |
| APPLIED_UNVERIFIED | The provider accepted the change, but the re-read doesn't show it yet (after bounded polling), or the re-read itself failed. |
| FAILED | Not made. Either the provider refused (with its reason), or the request may have been lost and the re-read confirmed it wasn't made, so retrying is safe. |
| UNKNOWN_OUTCOME | The request may have reached the provider and the provider couldn't be read to find out, or the worker stopped mid-call. |
| SKIPPED | Not approved, or not attempted (after an earlier failure, or in an interrupted run). |
APPLIED_UNVERIFIED is not a failure, and it isn't hidden either. Asking Resend to re-check a domain is a good example: the request is accepted at once, and the operation stays honestly APPLIED_UNVERIFIED while DNS propagates. Doctor and the next plan will see the real state later.
Recovering from an unknown outcome
The hardest case is a write that may or may not have happened: a network error, a timeout, or a 5xx on a request that wasn't a read. Ahena's core journals every such write. After the apply, it re-plans against the provider and resolves each one:
- The change is visible: VERIFIED. If the provider returns something only once (a webhook signing secret, for example), the note says it wasn't received, and Doctor flags the missing secret.
- The change is definitely still needed: FAILED, with "it wasn't made, so retrying is safe".
- The re-read itself fails: UNKNOWN_OUTCOME, with "check before retrying".
A plan interrupted mid-call ends as NEEDS_REVIEW. The next apply re-reads the provider and does only what's still missing. Only one apply runs per environment at a time, so a retry can't race a slow first run; a plan with no progress for 5 minutes is treated as interrupted, closed out, and audited as plan.interrupted.
Idempotency: retrying without duplicates
An operation is idempotent when doing it twice has the same effect as doing it once. Setting a value is naturally idempotent; creating a resource usually isn't. Some APIs solve this with idempotency keys: the client sends a unique key with the create, and if the same key arrives again, the provider replays the original result instead of creating a second resource.
Ahena uses whichever mechanism each provider actually has:
- Stripe supports idempotency keys. Webhook endpoints are created with a stable key, so a retry after a lost response replays the original endpoint and its signing secret. Products and prices use keys plus deterministic ids. This replay behaviour has been exercised against Stripe's test mode (same objects back, no duplicates) through Ahena's test harness, not yet through
ahena apply. - Cloudflare and Resend don't offer idempotency keys for these creates, so Ahena reads, matches, creates only if absent, then re-reads. Setting R2 CORS is a PUT, which is idempotent by nature.
- Firebase: an "already exists" (409) after a retry counts as applied, and the re-read confirms it.
- Supabase auth settings use set semantics, so applying the same site URL and redirect URLs twice changes nothing.
Retries follow the same logic. For every remote provider, rate limits (429) and gateway errors (502, 503, 504) are retried with bounded, jittered backoff that honours Retry-After. But a non-idempotent write is retried only on 429, because a 429 means the provider refused the request without processing it. Reads, PUTs, DELETEs and requests carrying an idempotency key are retried on any of those errors.
A plan's status is never rounded up
A plan's status follows from its operations:
APPLIEDonly when every operation that ran was verified;APPLIED_UNVERIFIEDwhen something couldn't be confirmed;PARTIALLY_APPLIEDorFAILEDon failures;NEEDS_REVIEWwhen any outcome is unknown.
Partial failure is reported, never hidden. If Stripe fails after Cloudflare succeeded, the plan ends PARTIALLY_APPLIED with each operation's outcome, and Ahena doesn't delete what succeeded to fake atomicity across providers. To retry, run apply again:
ahena apply -e production # plans afresh from the providers' live state, then appliesThe fresh plan is built from what the providers have now, so completed steps aren't repeated. The audit log records the same honesty: an apply is recorded as success, partial, failure or unknown, with one line per outcome. For how plans are made and approved before any of this runs, see Plan, approve, apply, verify.
Try it on your own stack
Start free with one project. Connect the providers you already use, run Doctor, and see the plan before anything changes.