P0 from docs/async_refactor.md. The mutating routes are sync `def`, so FastAPI
runs them in a threadpool, and they reach Runtime.broadcast through
rebuild_levels — writing asyncio.Queue directly from there. That queue is not
thread-safe: it wakes a consumer by resolving a Future, which only the loop
thread may do. A dropped wakeup means a drawing made in one browser does not
reach another until the next market tick.
broadcast now posts through call_soon_threadsafe when it is off the loop, and
publishes directly when it is on it, so the stream's own path pays nothing.
Worth being straight about the tests: the race is timing-dependent and did not
reproduce in twenty attempts — a foreign-thread put_nowait usually lands in the
ready queue before the loop sleeps, and a tick every second covers the rest.
Even asyncio's debug thread-affinity check stays quiet unless a consumer is
parked on the Future at that instant. So the tests assert the contract rather
than provoke the failure: a broadcast from a worker thread must go through
call_soon_threadsafe, one from the loop must deliver synchronously, and both
must arrive.
Also adds the loop-lag probe, which reports scheduling drift as loop_lag_ms on
/api/status. It found P1 on its first run: 19,441ms worst against 1.5ms in
steady state, which is seeding blocking the loop. "The chart feels laggy" is now
a number.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An audit for blocking work on the event loop, written up rather than acted on —
the chart changes in flight land first.
The headline is a correctness bug, not a performance one. Sync route handlers
run in FastAPI's threadpool and call rebuild_levels, which reaches
asyncio.Queue.put_nowait on every subscriber. asyncio.Queue is not thread-safe:
it wakes a consumer by resolving a Future, which has to happen on the loop
thread. A dropped wakeup means a drawing made in one browser does not reach
another until the next tick — invisible today only because the stream ticks
about once a second and covers it.
Below that: level rebuilding is CPU-bound on the loop and is the whole of the 82
second startup, and disarming an alert writes to disk from a coroutine.
Also states what not to do, since the obvious reading of "make it async" is
wrong here. Sync routes stay sync — FastAPI's threadpool is what keeps their
work off the loop, and converting them would drag the rebuild cost onto it.
ManualLineStore's threading lock stays, because both the loop and threadpool
threads reach that store.
Keeping it that way is three layers: a short async section in AGENTS.md, which
is the only file both agents load every session; comments on the lines someone
would actually edit, starting with the worker count in Procfile; and a loop-lag
probe on /api/status so a stall reports itself as a number rather than as "the
chart feels laggy". The risk register gains a row per finding pointing here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>