← All articles

7 min read

Who owns a cell?

How conditional ownership records and epochs stop a recovered Peren node from competing with its previous owner.

Node A owns a stateful Worker. It commits a change and then disappears from the network. Node B recovers the work. A few seconds later, A resumes.

Both processes are alive. Both have the Worker bundle. Both may have local storage for the same object. Which one is allowed to answer the next request?

Peren resolves the conflict with a cell ownership record and an ownership epoch. The record answers “who owns this now?” The epoch tells a returning node whether its answer is out of date.

01 · REQUESTClient or eventHTTP · queue · alarm · schedule
route
02 · RUNTIMEWorker + bindingsV8 isolate with granted capabilities
dispatch
03 · CELLOwner + epochSerial execution · SQLite · output gate
publish
04 · OBJECT STORERecoverable stateOwnership record · snapshot · WAL
Ownership is decided between the cell and the shared record. An epoch fences a node whose authority is stale.

Start with one ownership rule

Every Durable Object style instance maps to a deterministic cell ID. Any node can derive the same ID, but only the current owner may execute that cell and acknowledge its durable output.

The rule is narrow:

At most one ownership epoch may acknowledge durable output for a cell at a time.

A stale process may still run instructions briefly; a network partition cannot erase its memory. What Peren can do is stop that process from presenting stale work as a successful durable result.

Three facts must agree before a state-changing response leaves a node:

  1. the shared record names this node as owner;
  2. the node’s ownership epoch is still current; and
  3. the response’s committed state is recoverable.

One conditional update chooses the owner

The ownership record contains the current node, if any, and an epoch. The object-store provider also returns a comparison version such as an ETag or generation.

To claim an unowned cell, a node creates the record only if it is still absent. To transfer or recover a cell, it replaces the record only if the version it read has not changed. If two nodes race from the same version, only one update can be applied.

Losing that race is ordinary contention. The node reloads the record and follows the winner. A timeout is trickier: the update may have reached storage even though the reply did not reach the node. The only safe next step is to read the authoritative record and reconcile what happened.

Ownership failure · Step 1 of 4

A owns epoch 17

NODE AOWNERholds epoch 17
AUTHORITATIVE RECORDA · epoch 17comparison version v17
NODE BOBSERVERreads epoch 17
Event

A is serving the cell. B can observe the record but has no authority to execute.

Who may acknowledge?NODE A

The record and A's retained epoch agree.

Loss of contact permits a recovery attempt. Only the successful conditional update changes authority.

A missed heartbeat or expired lease gives B a reason to try recovery. It does not give B the cell. B becomes the owner only after it satisfies the recovery policy and wins the conditional record update.

Epochs fence a node that returns

Suppose A owns epoch 17. B recovers the cell and advances the record to epoch 18. A may still hold open sockets, a resident isolate and a local database, but it carries an old generation.

That one-number difference is the fence. When A checks its authority, the shared record says B owns epoch 18. A cannot publish or acknowledge work as epoch 17.

Epochs also separate replicated data from different owners. Publication identity includes the cell and ownership generation. Data produced under epoch 17 cannot overwrite epoch 18’s sequence. Different bytes at an identity that should be immutable are surfaced as corruption.

The response waits behind ownership

Ownership can change while a request is executing, so checking only at the beginning is insufficient.

For a state-changing invocation, the cell:

  1. executes under its retained ownership lease;
  2. records the highest storage revision committed by the invocation;
  3. verifies ownership;
  4. publishes the state needed to recover through that revision;
  5. verifies ownership again; and
  6. releases the response.

Why check twice? A can verify epoch 17, begin an upload and then stall while B takes epoch 18. When A’s upload eventually finishes, upload success alone cannot let it answer. The final check sees the newer owner and keeps the output gate closed.

If authority is uncertain, Peren drains or fences the cell. It does not turn an ownership failure into an empty success or a response that another node may be unable to recover.

Recovery starts from acknowledged state

After B acquires the cell, it restores a valid snapshot and the ordered write-ahead-log records that follow it. A snapshot supplies the database baseline; WAL records carry later committed changes. Peren refuses a WAL chain without a valid baseline, missing positions or conflicting object bytes.

Recovery targets the last acknowledged durable position. A local change that never crossed the output gate is not presented as a result the client successfully observed.

Recovery chain

Restore only confirmed, ordered state

BASELINESnapshotposition 40
+
NEXTWAL 41checksum valid
+
NEXTWAL 42acknowledged
→
RESTOREDRevision 42ready to activate
Refused: WAL without a snapshot · missing position 41 · conflicting digest
Recovery starts from a valid baseline and applies the confirmed sequence in order.

Object storage is therefore part of the live stateful path. Its conditional writes choose ownership, and its snapshot and WAL records make acknowledged state recoverable. The bucket is doing more than holding backups.

Clean shutdown uses the same guard

A voluntary release conditionally writes an unowned record while preserving the current epoch. The next acquisition advances the generation.

Release consumes the authority held by that owner. It cannot clear a record using only a cell ID, and it cannot erase a newer owner’s update. A late shutdown task from A therefore cannot remove B’s epoch 18 record.

The same rule therefore covers a planned drain and an unplanned recovery. Moving a cell always passes through guarded ownership state.

Try to break the rule

The dangerous cases sit between otherwise ordinary operations, so Peren tests them at several levels. Executable models explore ownership transitions and stale owners. Domain tests race acquisition, recovery and late release. Distributed scenarios write through a real Worker on one node, restore the acknowledged state on another and then bring the original node back.

The last step runs the conditional-write cases against the object store itself. A model can find a flaw in the protocol, but it cannot tell us how a hosted endpoint behaves when two writes race or a response disappears.

Coming back online does not make a node the owner again. It must still match the shared record and current epoch before the output gate will let its response leave.

Read the durability model, placement rules and degraded-operation guide for the operational details.