Code-review finding, not reproduced at runtime. An independent static review of #159 found this. The algorithm predates #159: #159 moved it unchanged from apps/host/src/worker.ts into catalog.lock.
catalog.lock at a476c32 (
|
export function lock(id: string): boolean { |
|
const { pid } = paths(id); |
|
try { |
|
writeFileSync(pid, String(process.pid), { flag: "wx" }); |
|
return true; |
|
} catch (error) { |
|
if ((error as NodeJS.ErrnoException).code !== "EEXIST") throw error; |
|
if (isAlive(Number(readFileSync(pid, "utf8")))) return false; |
|
rmSync(pid, { force: true }); |
|
return lock(id); |
|
} |
|
} |
):
writeFileSync(pid, String(process.pid), { flag: "wx" }); // claim
...
if (isAlive(Number(readFileSync(pid, "utf8")))) return false;
rmSync(pid, { force: true }); // reclaim a dead owner's file
return lock(id);
Interleaving, with a stale worker.pid holding dead PID D and contenders A and B (for example two clients starting the same dormant channel after a crash):
- A and B both get EEXIST, both read D, and both see it dead.
- A runs
rmSync, retries, and its wx write succeeds: A holds the lock.
- B runs
rmSync, removing A's fresh lock file, then its own wx write succeeds: B holds the lock too.
Two workers then own one channel's SQLite storage and socket, which breaks the single-owner invariant in docs/architecture.md. The same window applies to remove()'s use of the lock during delete.
Acceptance: a single exclusive owner of a channel's storage, held across crash recovery and concurrent starts and deletes. Verify it with real contenders started at the same moment after a worker died and left a stale worker.pid: exactly one worker may open the channel each time.
Found while reviewing #159 (#153). Dogfooding #5.
Code-review finding, not reproduced at runtime. An independent static review of #159 found this. The algorithm predates #159: #159 moved it unchanged from
apps/host/src/worker.tsintocatalog.lock.catalog.lockat a476c32 (ace2/apps/host/src/catalog.ts
Lines 214 to 225 in a476c32
Interleaving, with a stale
worker.pidholding dead PID D and contenders A and B (for example two clients starting the same dormant channel after a crash):rmSync, retries, and itswxwrite succeeds: A holds the lock.rmSync, removing A's fresh lock file, then its ownwxwrite succeeds: B holds the lock too.Two workers then own one channel's SQLite storage and socket, which breaks the single-owner invariant in
docs/architecture.md. The same window applies toremove()'s use of the lock during delete.Acceptance: a single exclusive owner of a channel's storage, held across crash recovery and concurrent starts and deletes. Verify it with real contenders started at the same moment after a worker died and left a stale
worker.pid: exactly one worker may open the channel each time.Found while reviewing #159 (#153). Dogfooding #5.