Commit graph

4 commits

Author SHA1 Message Date
372617b08c Split the plan from the log of what actually happened
One file was trying to be two things: a spec written to be executed
top-to-bottom, and a dated record of everything that went wrong on the way. At
1,882 lines it did neither well, and the log was 36% of it — which is why the
plan's opening went unmaintained for days while the log grew every hour.

docs/plan.md keeps the decisions and the reasoning behind them, including the
risk register. docs/implementation.md takes the dated entries: the problems, the
wrong theories, the measurements that settled them. Git already says what
changed; that file says why it was hard, which is the part worth reading before
debugging something similar. Most entries describe something that looked like
one bug and turned out to be another.

Each points at the other, and the four referring files — AGENTS.md, README.md,
NEXT_STEPS.md and async_refactor.md — now point at whichever half they meant.
Git tracked the rename, so history follows plan.md rather than starting over.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:29:47 -05:00
ff1b9982d1 Seed in bulk and coalesce level rebuilds
P1 from docs/async_refactor.md. Measured on the dev stack: the port now accepts
connections 4 seconds after a restart rather than 121, and the worst loop lag
falls from 19,545ms to 526ms, with steady state between 0.2 and 0.6ms.

Seeding replayed years of history through on_bar, rebuilding every level from
scratch per bar and broadcasting each one to nobody. It now fills the store
quietly and derives price, ATR and the level set once at the end, from the
finished history. Alerts are deliberately not evaluated over replayed bars: a
level touched two years ago is not news, and firing on history is one way a
deploy re-alerts.

The seed was not all of it. Yahoo's first poll emits a whole day of minutes in a
single burst, each one taking the full live path, which was most of the
remaining twenty seconds. request_rebuild now coalesces to at most one rebuild
per 250ms and a background pass flushes anything deferred, so a burst costs a
handful of rebuilds instead of hundreds and the last bar is still never the one
dropped.

Verified unchanged after the change: bar counts across every timeframe, all five
daily moving averages with their full point sets, prior-day levels and VWAP. 125
python tests and 31 e2e tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:42:27 -05:00
a395818581 Post cross-thread events through the loop, and measure how late it runs
P0 from docs/async_refactor.md. The mutating routes are sync `def`, so FastAPI
runs them in a threadpool, and they reach Runtime.broadcast through
rebuild_levels — writing asyncio.Queue directly from there. That queue is not
thread-safe: it wakes a consumer by resolving a Future, which only the loop
thread may do. A dropped wakeup means a drawing made in one browser does not
reach another until the next market tick.

broadcast now posts through call_soon_threadsafe when it is off the loop, and
publishes directly when it is on it, so the stream's own path pays nothing.

Worth being straight about the tests: the race is timing-dependent and did not
reproduce in twenty attempts — a foreign-thread put_nowait usually lands in the
ready queue before the loop sleeps, and a tick every second covers the rest.
Even asyncio's debug thread-affinity check stays quiet unless a consumer is
parked on the Future at that instant. So the tests assert the contract rather
than provoke the failure: a broadcast from a worker thread must go through
call_soon_threadsafe, one from the loop must deliver synchronously, and both
must arrive.

Also adds the loop-lag probe, which reports scheduling drift as loop_lag_ms on
/api/status. It found P1 on its first run: 19,441ms worst against 1.5ms in
steady state, which is seeding blocking the loop. "The chart feels laggy" is now
a number.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 15:30:45 -05:00
9fc6402f03 Plan the async work, and record where it must not regress
An audit for blocking work on the event loop, written up rather than acted on —
the chart changes in flight land first.

The headline is a correctness bug, not a performance one. Sync route handlers
run in FastAPI's threadpool and call rebuild_levels, which reaches
asyncio.Queue.put_nowait on every subscriber. asyncio.Queue is not thread-safe:
it wakes a consumer by resolving a Future, which has to happen on the loop
thread. A dropped wakeup means a drawing made in one browser does not reach
another until the next tick — invisible today only because the stream ticks
about once a second and covers it.

Below that: level rebuilding is CPU-bound on the loop and is the whole of the 82
second startup, and disarming an alert writes to disk from a coroutine.

Also states what not to do, since the obvious reading of "make it async" is
wrong here. Sync routes stay sync — FastAPI's threadpool is what keeps their
work off the loop, and converting them would drag the rebuild cost onto it.
ManualLineStore's threading lock stays, because both the loop and threadpool
threads reach that store.

Keeping it that way is three layers: a short async section in AGENTS.md, which
is the only file both agents load every session; comments on the lines someone
would actually edit, starting with the worker count in Procfile; and a loop-lag
probe on /api/status so a stall reports itself as a number rather than as "the
chart feels laggy". The risk register gains a row per finding pointing here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:58:04 -05:00