chart/docs/implementation.md

46 KiB
Raw Blame History

Implementation log — what actually happened

A dated record of problems encountered building this and how each was resolved. Kept deliberately: git says what changed, this says why it was hard, what the wrong theories were, and what the evidence turned out to be. Most entries describe something that looked like one bug and was another — which is the part worth reading before debugging something similar.

The plan and the reasoning behind design decisions live in docs/plan.md.

Newest entries are at the bottom.

Dated record of problems hit and how they were resolved. Times are UTC; the repo's commit timestamps are -0500.

2026-08-10 — rebuild, and a chart that looked frozen

10:30 · The rebuild was genuinely required. schwab-py had been added to requirements.txt, but the running image was built at 2026-08-09 22:05, before that line existed. The bind mount (.:/app) hides this: source edits appear live, so the Schwab commits looked deployed while pip freeze in the container showed no schwab-py at all. Anything imported rather than read from disk needs docker compose build. Rebuilt to schwab-py 1.5.1 and recreated the container.

10:30–10:31 · Startup takes ~82 seconds, and the port is closed the whole time. Runtime.start() replays every seeded bar through on_bar, and each daily-bar update re-runs rebuild_levels() → broadcast_level_delta(), which serialises and diffs five MA levels carrying ~730 points each. With a 730d/1h seed plus an 8d/1m seed that is quadratic work before uvicorn binds. Measured: 10:30:24 "Waiting for application startup" → 10:31:46 "Application startup complete". An open browser tab polling /api/status throughout logs a wall of ERR_CONNECTION_REFUSED; that is the restart window, not a fault. Unfixed. The fix is to bulk-load seeded bars and rebuild levels once at the end, rather than once per bar. Related: M7 persistence would cut the seed itself.

Diagnosing a hang that is actually slowness: docker stats reported ~0.1% CPU while the process was in fact grinding, so it pointed the wrong way. What worked was faulthandler.dump_traceback_later(25, exit=True), which named the exact frame (indicators.py:sma under runtime.py:62). py-spy is unusable here — it needs SYS_PTRACE, which the container does not have.

Do not write scratch files into the repo while diagnosing. A _probe.py dropped in the project root is inside the bind mount, so --reload restarted the lifespan and reset the 82-second clock — twice — which is what made slow startup look like an infinite hang. Pipe throwaway scripts over stdin (docker exec -i … python -) instead. Only .py changes trigger the reloader; writing screenshots into artifacts/ is safe.

10:35 · A blank chart canvas that was not a bug. The Playwright container has no usable locale, so Chromium reports en-US@posix; Lightweight Charts formats its time axis through Intl, which throws Invalid language tag and leaves the canvas empty. docker-compose.yml already sets LANG=en_US.UTF-8 for that service and it is not sufficient. Launch with chromium.launch({ args: ['--lang=en-US'] }) — with that, the page renders and reports zero console errors. Worth stating plainly: this failure looks exactly like a broken app, and it is not.

10:41–10:50 · The real bug — the chart sat ~10 hours behind a healthy feed. Symptom: header price live at 7785.00 while the last candle closed 7772.75, and the series appeared to end at 00:20. Everything downstream checked out — /api/bars newest 10:39 from schwab; store.put keeps bars strictly ascending; the WebSocket snapshot delivered 1000 ascending bars ending 10:42 and live bar events arrived every minute; the browser received all of it. Interrogating window.__chart gave the answer:

seriesLen 1000   seriesLast 08-10T10:48 (7786.25)   ← data complete
visible   08-07T20:41 → 08-10T00:35                  ← viewport wrong
logical   from 840 to 1005

The series was complete; the viewport was 617 bars too far left — exactly bars_held.1d. setBars() set a visible logical range from the candle array length, then syncVisibleLevels() attached the daily MA series, whose 617 daily points pre-date the 1m window; prepending them renumbered every logical index and dragged the view off the live edge. Fixed in static/chart.js by anchoring the viewport to a time range. Verified in a real browser: visible range 08:02 → 10:49, last candle 7786.25 matching the header. See §9.

Method note. Three checks in a row said "healthy" — the REST API, the WebSocket, and the frontend source all looked correct in isolation, because each of them was correct. Only querying the live page's own chart object separated "the data is missing" from "the data is off-screen". Screenshots alone were actively misleading here: the stale time axis was read as a session gap.

2026-08-10 (later) — real-time ticks, and what to do about cold restarts

The chart now moves between minute closes. CHART_FUTURES emits a bar only once its minute is over, so the chart stepped once a minute and sat still in between — read, reasonably, as a dead feed. LEVEL_ONE_FUTURES carries real trades on the same socket (delayed: False, verified on this account back in M6), and it was never subscribed. It is now, and it builds a forming bar for the current minute which the authoritative CHART_FUTURES bar then supersedes.

Three constraints shaped it, each of which would have caused a real bug:

  • Tick bars must never reach the aggregator. It accumulates with current.v += incoming.v, so re-sending the same forming minute would add its volume into every higher timeframe on every update. Runtime.on_bar returns early for not bar.closed: store the bar, set the price, broadcast, stop.
  • Ticks are throttled (SCHWAB_TICK_SECONDS, default 1.0). /ES trades many times a second and each emission costs a store write plus a broadcast to every open socket. Setting it negative drops the Level 1 subscription entirely and returns the source to closed bars only.
  • A tick for a minute already closed is dropped, or a late trade would overwrite a settled exchange bar with a partial one.

Alerts deliberately stay on closed bars. A level is judged on a settled bar, not on a price that may not last the minute — and on_bar already gated on closed, so this needed no change. Intra-bar alerting is a separate decision.

Bid-only Level 1 updates are skipped rather than carried forward: a bid is not a trade and must not extend a candle's high or low. Verified live — 15 forming bars and 2 closed bars in 100 seconds, and in a browser the candle's high and low visibly extend within the minute.

Cold restarts — the options, and a recommendation. Every restart costs ~82 seconds of refused connections, re-seeds from Yahoo, and starts with empty alert cooldowns, so a deploy can re-alert whatever price is sitting on.

  1. Make seeding non-quadratic. Seeding replays every bar through on_bar, and each daily-bar update rebuilds all five MA levels and diffs them. Bulk-load the seeded bars and rebuild levels once at the end. Contained, testable, and removes most of the 82 seconds. Do this first — it is the cheapest real win and needs no new storage.
  2. Persist bars (M7, SQLite). Restarts then seed only the gap. Removes the Yahoo dependency from the startup path and shrinks the window further. This is the durable answer, and the plan already scopes it.
  3. Persist alert cooldowns and armed state. Independent of 1 and 2, and the part that actually misbehaves rather than merely being slow: without it every deploy re-alerts. Small table, big behavioural win.
  4. Serve before seeding finishes. Start uvicorn immediately and seed in a background task, so the port never refuses. The chart would open cold and fill in, which is better than an unreachable page — but it changes what "warm" means to every consumer of /api/status, so it wants its own thought.

Recommended order: 1, then 3, then 2. 4 only if the window still bites after 1.

Stale bar events across a timeframe switch. Cannot update oldest data appeared in the console once ticks were live. Switching timeframe races: the server answers subscribe with a fresh snapshot from one coroutine while another is still draining bar events for the timeframe just left, so a 1m bar can land after the 1h snapshot. Applied to the 1h series it is older than every point in it, and Lightweight Charts throws rather than ignoring it — taking the app down instead of dropping one bar. The race predates the tick feed; Level 1 made bar events ~15x more frequent, which is what surfaced it.

Guarded at both ends. app.js honours the tf the event already carries and drops anything for a timeframe that is no longer selected. chart.js refuses a bar older than the series' last point regardless of where it came from — a bar behind the last one has nothing to contribute. Verified: 36 rapid timeframe switches under a live tick feed produce zero errors, and calling candles.update() directly with a stale bar still throws while the guarded updateBar() does not.

Same-price trades were being dropped. The candle still paused for 10–20 seconds at a time after Level 1 went in. Instrumenting the raw stream settled it: 87 messages in 90 seconds, only 33 carrying LAST_PRICE. Most of the rest are pure bid/ask movement and correctly ignored — but a seventh of them look like this:

['ASK_SIZE','ASK_TIME_MILLIS','BID_SIZE','BID_TIME_MILLIS',
 'LAST_SIZE','QUOTE_TIME_MILLIS','TOTAL_VOLUME','TRADE_TIME_MILLIS','key']

Trade time, trade size, cumulative volume — and no LAST_PRICE, because Level 1 sends only changed fields and the trade printed at the price of the one before. Requiring LAST_PRICE threw those away along with their volume. parse_level_one now treats size-plus-trade-time as a trade and returns a null price for the caller to carry forward. Measured on the live feed: median gap 3.1s → 2.0s, worst 21.5s → 8.1s, and bar volume climbs within the minute instead of standing still.

Worth recording for the next person who reads a gap as a bug: the remaining pauses are the market, not the pipe. In thin pre-open tape /ES genuinely goes seconds without a price-changing trade, and then moves several ticks at once — which is what a "gap up" after a quiet spell actually is.

The time axis reads local, the data stays UTC. Lightweight Charts is timezone-agnostic: it reads epoch seconds as UTC and labels them as UTC, which is why the axis disagreed with the wall clock. Fixed with tickMarkFormatter for the axis and localization.timeFormatter for the crosshair, both going through the browser's own zone.

Deliberately not fixed by shifting the bar timestamps, which is the other common recipe. Every time in this codebase is epoch UTC by convention, and the chart's own times feed trendline anchors, indexAt, hit testing and the values posted back for manual lines — an offset applied to the data would put all of them out by the offset, which is exactly the class of bug that once priced a trendline 147 points away.

One limit worth knowing: tick placement is still computed on UTC days, so the day-change divider sits at 00:00 UTC rather than local midnight, labelled with the local date. The labels are right; the divider is in the UTC place.

2026-08-10 (afternoon) — update rate, the left scale, and volume

Schwab conflates Level 1 to one update per second. Chasing "still slow in market hours" ended at a hard ceiling rather than a bug. In regular hours the gaps between updates are whole multiples of 1.005s — 2.01, 3.02, 4.03 — which only happens if the source emits on a one-second cadence and some seconds carry no trade. SCHWAB_TICK_SECONDS was the limiter at 1.0 and is now 0.25, where it no longer binds. One update per second is the source's ceiling. Anything faster would mean inventing prices between trades, which a chart must not do.

Two real losses were found on the way and fixed:

  • Higher timeframes only moved once a minute, because tick bars are 1m and the socket filters by subscriber timeframe. Runtime.provisional_higher now combines the aggregator's committed state with the live minute — without mutating it, since the aggregator accumulates volume and would double count.
  • Trades carrying only a trade stamp and a moved TOTAL_VOLUME — no LAST_PRICE, no LAST_SIZE — were skipped. 66 → 74 updates per 90s.

Daily context moved to the left price scale. The right had prior-day levels, session VWAP and five daily MAs competing with the live price and hand-drawn intraday levels. The trap: a price scale takes its range from the series on it, so moving levels across draws them against a different range and puts them at the wrong height. A transparent candlestick mirror on the left scale feeds it exactly the right scale's input; verified as a zero-pixel delta between the two. Hand-drawn levels stay right, which is the space being cleared. priceScaleId is fixed at series creation, so it is passed at construction and kept out of the options reapplied afterwards.

Volume is finally drawn. It travelled the entire pipeline — parsed from both Schwab services, aggregated, stored, broadcast in every bar — and nothing rendered it. Now an overlay histogram on its own hidden scale in the bottom fifth. An overlay rather than a pane, and emphatically not the price scale: volumes are five figures against four-figure prices, and sharing a scale would flatten the candles to a line.

Sidebar vertical space. Three cuts, all in §9.4's layer panel. The daily MA periods sit on one line — flex-wrap:nowrap with tighter gaps and 12px boxes, where 10px gaps and 22px indent had pushed 200 onto a line of its own. Auto trendlines joins Manual lines as a parenthetical (auto) rather than owning a row, which suits a control that is disabled until M8. The alert log becomes a <details> like Confluence zones, closed by default with its count in the summary — collapsed by default is the point, since leaving it open would save nothing, and the count means activity is still visible while closed.

Sidebar vertical space. The right column was taller than the viewport with nothing selected. Five changes, no functionality removed:

  • The five daily MA periods fit one line (flex-wrap:nowrap, tighter gaps, 12px boxes); 10px gaps and a 22px indent had pushed 200 onto a row of its own.
  • Auto trendlines becomes a parenthetical (auto) on the Manual lines row rather than owning one, which suits a control disabled until M8.
  • Alert log becomes a <details> like Confluence zones, closed by default with its count in the summary so activity still shows while shut.
  • Tools becomes a <details> too, open by default, and each tool's panel is bound to armedTool — only the armed tool shows its label, colour, width and side controls. armTool already toggles and permits one armed tool at a time, so the panels follow it exactly.
  • Order is Layers, Tools, then the rest, with Layers collapsed by default.

Measured with nothing armed: 1110px of content down to 900px, which is inside the viewport rather than past it. .sidebar-section:first-of-type carries the zeroed top margin so reordering cannot reintroduce a gap at the top.

2026-08-10 (evening) — chart comments, and Drawings

Comments are drawings, not levels. A comment is stored as a ManualLine with kind="comment", so it inherits persistence, the shared drawing-number sequence, the sidebar list, filtering and deletion without a parallel set of endpoints. The one rule that must never bend: ManualLineStore.levels() filters comments out. A comment reaching the level list would join a confluence cluster and push a phone notification about a piece of text. It is also created with armed=False, and PATCH /lines/{id} returns to_dict() rather than to_level() for one, so no caller is ever handed a level-shaped comment.

kind is derived when absent — zero slope was always a typed level, anything else a drawn trendline — so drawings saved before comments existed keep working.

Pinned or floating. Pinned comments carry anchor_t/anchor_p and move with the chart; floating ones carry x/y as fractions of the pane, hold their place through any zoom, and can be dragged. Comments render as DOM rather than canvas: they hold arbitrary text, collapse to a numbered dot, and a floating one has to ignore the time scale entirely. A pinned comment scrolled out of view parks on the edge it left, pointing back the way it went, so it is never simply lost.

"Lines & levels" becomes "Drawings", filtered by type and by text — the text match covers the label, the kind and the #number, so comment, cpi and 7 all narrow the list. Delete acts on what the filter shows, which is what makes deleting by type or by string a single button.

One CSS trap worth recording: .trendline-row span { grid-column:2 } captured the comment row's icon span and dragged it into the text column. Scoped to span:not(.drawing-icon).

A comment lost its place when the timeframe changed. Placed on a 30m bar, then switched to 15m, it slid to the far left. timeToCoordinate answers only for times that are data points on the current series, so a 30m bucket start returned null on another timeframe — and null was being read as "off the left edge". Anchors are now resolved to the bar that contains them, which is timeframe-independent: an 09:30 note sits on the 09:30 bar at 15m and on the 09:00 bar at 1h. setBars also re-renders comments, since a timeframe switch replaces the grid underneath every pinned one.

Verified across 30m → 15m → 1h → 30m: the anchor stays 08:30 throughout, resolving to the 08:30 bar on 15m and the 08:00 bar on 1h, never edge-parked, and returning to its original x on the way back. Edge-parking still works where it should — a comment scrolled 400 bars out parks right and comes back on return to live.

The trendline Side control became inert. Once snapping always lands on a bar extreme, the side is inferred from which extreme — a high is resistance, a low is support — so the dropdown could no longer affect anything. It now appears only when "Snap to highs/lows" is off, which is the one case where there is no extreme to infer from; otherwise the row reads "Side auto". Verified both ways: snap on shows the note and no dropdown, snap off shows the dropdown.

created_at (epoch seconds) is already stored on every drawing and returned by GET /api/drawings, so filtering by age needs UI only, not a migration.

Trendline placement, third pass — and a regression I shipped. Making a pending anchor always win (previous entry) fixed the twitch case and broke the opposite one: a genuine press-drag begun after an abandoned click was hijacked by that stale anchor, so the line started far from the drag. That reached production. The rule is now a single threshold — 12px of travel between press and release makes it a drag, which is wide enough to survive a twitch on a deliberate click and unambiguous for a real drag. A drag clears any half-placed anchor rather than silently adopting it.

The crosshair was lying about the anchor. Lightweight Charts defaults to CrosshairMode.Magnet, which snaps the crosshair to the bar's close. Hovering by a bar's low therefore drew the crosshair mid-bar, and a correctly-snapped anchor looked wrong — measured: aiming 4px above a bar low placed the anchor at the low (7773) and not the close (7773.25), while the crosshair sat at the close. Arming a tool now switches the crosshair to Normal, and a snap dot marks the exact point the anchor will use, coloured by the side it implies.

Four gesture paths are verified in a browser: two clicks with a twitch on the second, an abandoned click followed by a real drag, a plain press-drag, and hovering. All start where they should and land on a bar extreme.

Worth recording for diagnosis: a reported "line ended up high off the bar" turned out to render exactly on its bar — zero pixels off at 1h, 30m and 15m — because the anchor had snapped to the drawn timeframe extreme (the 09:00 1h low, 7744.25) while being checked against 1m bars, where it matches neither extreme. Always compare an anchor against the timeframe it was drawn on.

A zero price wrecked every timeframe's scale. A LEVEL_ONE_FUTURES update arrived with LAST_PRICE: 0. The parser rejected None but 0 is not None, so a minute opened at zero — o=0.0 h=7777.25 l=0.0 — and provisional_higher carried that low into 5m, 15m, 30m, 1h and the daily bar, flattening the price scale everywhere. Non-positive prices are now treated as absent, so the last real price carries forward, and the tick still counts as a trade.

The exchange's own bars were being dropped. store.put replaced a bar only when it matched the tail. That held while one closed bar arrived per minute, but ticks open the next minute before CHART_FUTURES delivers the previous one — so the authoritative bar no longer matched the tail and was discarded, leaving the tick approximation and its partial volume in place permanently. put now searches back a bounded number of buckets, and refuses to let a provisional bar overwrite a settled one.

Both were introduced by the tick feature and both are covered by tests: a zero price parses as a trade with no price, a late closed bar replaces its bucket and keeps the exchange's volume, and a tick cannot overwrite a settled bar.

Snapping now measures distance on screen, not in time. The rule was "take the bar sharing the cursor's time, then its nearer extreme", which ignored how far that extreme actually was. Pointing anywhere below a candle snapped to that candle's low however distant, and the extreme genuinely under the cursor was never considered — so zoomed out to ~360 bars at three pixels each, hitting the intended bar took several attempts. snapPoint now scans six bars either side and picks the extreme nearest in pixels. Proven by probe: with the cursor on one bar's low but nudged two pixels so coordinateToTime resolves to its neighbour, the snap takes the extreme under the cursor rather than the neighbour's.

Worth recording because it was misdiagnosed twice: a report of "the snap dot appears way above the bar" was, on the numbers, the dot landing correctly on the bar's low while the cursor sat 151 points below it. The right price scale keeps a bottom: 0.1 margin and the volume overlay is drawn in it, so the lower fifth of the pane is below every candle — an inviting place to point that contains no price action at all.

e2e tests

bin/e2e runs tests/e2e/*.test.mjs inside the playwright service against the dev stack. Node's built-in test runner, no dependencies added to this repo: Playwright is global in that container and tests/e2e is mounted at /repo/tests/e2e. Every case in there is a bug that shipped — the viewport parked ten hours back, hourly candles drawn as slivers, stale bar events throwing, comments drifting on a timeframe switch, and three separate ways a trendline anchor could disagree with its own preview. None of them could have been caught by pytest, which is the argument for the suite existing.

Tests clean up after themselves: withChart records the drawings that exist before the body runs and deletes anything new afterwards, because the dev store is shared with whoever is using the app. Select by title rather than class when asserting on chart overlays, for the same reason.

The snapping rule, stated once

x picks the bar, y picks which extreme. That is the whole rule. It is written here because changing it reactively three times is what made trendlines feel broken, not any inherent difficulty:

  1. An 8px proximity gate meant a cursor between the high and the low snapped to neither, so the anchor kept a raw mid-bar price and the side silently fell back to the dropdown.
  2. Removing the gate fixed that. Then a nearest-in-2D search was tried, to make a bar easier to hit when zoomed out — and broke sweeping along the bottom, because whichever nearby bar had the lowest low won on total distance and the dot skipped off the bar under the cursor. Reverted.
  3. What actually made it feel wrong was never the rule: the crosshair was in Lightweight Charts' default Magnet mode, snapping to the bar's close, so the feedback pointed somewhere the anchor would never go. It is Normal everywhere now, with the snap dot showing the real target.

An e2e test sweeps the cursor along the bottom of a zoomed-out 1h chart and requires every position to land on the low of the bar beneath it — 115 of 115. That test is the rule, executable.

The snap leapt to the live edge — found by diagnostic mode. Reported from the user's own browser, which no headless run had reproduced:

cursor_x 1409.0  chart_w 1280.0  -> cursor_t None -> snapped to the last bar
cursor_x 1161.0  chart_w 1280.0  -> cursor_t None -> snapped to the last bar
cursor_x 1128.0                  -> valid time, drift 0 bars

coordinateToTime answers null over the right-hand price axis, over the whitespace past the last bar, and anywhere outside the chart — and snapPoint read that as "the newest bar", so the dot jumped to the live edge from wherever the cursor was. Two faults behind it: the tool's pointer listener is on window and therefore fires over the sidebar (x=1409 on a 1280-wide chart), and a null time meant a default rather than no answer.

Now a pointer outside the plot hides the indicator entirely, and a null time resolves to the bar nearest in pixels rather than the newest one. Covered by an e2e test that hovers a bar, the axis, the sidebar, and back.

The lesson is about method rather than geometry: four hypotheses were tested and killed by measurement here — device pixel ratio, viewport size, resize desynchronisation, and the chart scrolling under the gesture — while the actual cause was visible in one line of the client's own numbers. When the browser is on another machine, instrument it early instead of reproducing locally.

Overlays are positioned against the plot, not the element

The trendline snap was 66 pixels out, and so was everything else drawn over the chart. Lightweight Charts reports coordinates from the plot area's origin. The chart element also contains the price scales, so once the left scale was enabled for the daily labels, the plot started 66px into the element — and every overlay positioned with left: against the element was displaced by exactly that much, in both directions at once:

  • the cursor's element-x was read as a plot-x, resolving a bar ~66px to the right of the pointer;
  • the indicator was then drawn at that bar's plot-x interpreted as element-x, landing ~66px left of where the bar is painted.

Not near the cursor, not near the bar, and varying with zoom — 66px is a couple of bars at 30m and a dozen at 1m, which is why it looked random rather than offset. Every diagnostic number agreed with itself throughout, because dot_y, expected_y and bar_low_y all derive from the same API and shared the same wrong origin. Self-consistent instrumentation cannot see a systematic error in its own frame of reference.

All overlays now live in one container positioned over the plot canvas, so they inherit plot coordinates untranslated: the snap dot and label, the comment layer, the trendline anchor handles, the preview line, the tooltip, the price tag and the context menu. eventPoint subtracts the same offset, so a pointer position and a chart coordinate finally mean the same thing. The container is repositioned on resize.

This had been mis-diagnosed for hours: device pixel ratio, viewport size, resize desynchronisation, the chart scrolling under the gesture, and the dead band below the candles were each measured and ruled out. The measurement that found it was comparing canvas.width to element.clientWidth — 0.894 — which is the first thing that ever disagreed with itself.


Handoff: the snap indicator is ~10px left of where it belongs

Status: fixed. Everything below is measured, not inferred.

The symptom

With the Trendline tool armed, the snap dot sits about one bar to the left of the cursor. The height is correct and the bar it chooses is correct — only the horizontal drawing position is wrong. Reported from a real browser and reproduced headlessly.

The measurement

bar spacing            6.96 px
dot centre - cursor  -10.5 px   (= 1.5 bars at that zoom)
chosen bar            correct   (label names the bar under the cursor)

And the cause, from ConfluenceChart.syncOverlayLayer():

at load:          containerLeft 56   true plot offset 66    <- 10px stale
after a re-sync:  containerLeft 66   true plot offset 66    <- correct

Why

Lightweight Charts reports coordinates from the plot area's origin. The chart element also contains the price scales, so with the left scale enabled the plot begins 66px in. All overlays therefore live in a container (.chart-overlays) positioned over the plot, so they can use chart coordinates untranslated — see create() and syncOverlayLayer() in static/chart.js.

syncOverlayLayer() runs once in a requestAnimationFrame during create(). At that moment the left price scale has not finished sizing itself to its label text, so the measured offset is 56. It settles at 66 once labels render, and nothing re-measures. The container stays 10px left of the plot for the life of the page, which drags every overlay with it: the snap dot and label, comments, trendline anchor handles, the preview line, the tooltip and the price tag.

Only x is affected. y never passes through this offset, which is why the height has always looked right.

The fix to write

Re-measure instead of measuring once. Options, cheapest first:

  1. ResizeObserver on the plot canvas — fires when the scale settles and on every later change. Probably the right answer.
  2. Call syncOverlayLayer() at the top of renderComments() and showSnapDot(). Correct but does DOM reads on every mouse move.
  3. Re-sync on subscribeVisibleLogicalRangeChange as well as on resize. Cheap, but misses a scale that widens without the range changing.

Beware: the left scale's width depends on its label text, so it changes when the price range gains a digit or a longer level label appears. Whatever you choose must survive that, not just the initial load.

How to verify

./bin/e2e trendline          # 7 cases, all currently pass — they do not catch this

The suite misses it because its assertions go through the same coordinate API that carries the error. Add a test that measures in page pixels: place the cursor exactly at a bar's centre and assert the dot's centre is within ~2px horizontally. The reproduction is:

const r  = el.getBoundingClientRect();
const bx = chart.timeScale().timeToCoordinate(bar.t);
await page.mouse.move(r.x + c.plotOffsetX() + bx, r.y + c.candles.priceToCoordinate(bar.l) - 6);
// dot centre x should equal the cursor x; today it is ~10px left

Live numbers from the client are available without a console: open the chart with ?diag=1, then docker compose logs api | grep SNAPDBG. Note that SNAPDBG will not show this bug — dot_y, expected_y and bar_low_y all derive from the same API and share the same origin, so they agree with each other while being wrong together. The error is only visible by comparing against something outside that frame of reference: canvas.getBoundingClientRect() against element.getBoundingClientRect(), or painted pixels.

Context worth having

  • window.__chart is a deliberate debug handle exposing the wrapper.
  • The dev stack is at http://localhost:8010, and http://api:8000 from inside the playwright container. It runs on a remote machine; the user's browser does not. Headless passes prove little about their screen — see "Where things run" in AGENTS.md.
  • Related history is above under "Overlays are positioned against the plot, not the element", which fixed the 66px case this 10px residue survived.
  • One e2e test, "clicking a comment collapses it", is flaky (roughly one run in three) and unrelated. Worth fixing before trusting the suite.

Resolution

ConfluenceChart now attaches a ResizeObserver to the plot canvas itself. Unlike the outer chart element, that canvas changes width when Lightweight Charts finishes sizing the left price scale or a longer price label appears. The observer repositions the shared overlay layer and redraws its anchored DOM content after each such change.

The trendline suite now compares the snap dot and cursor in page pixels, then forces a six-digit left-scale label and repeats the assertion. This catches the stale coordinate frame that the chart API's self-consistent coordinates could not. The comment-collapse flake was also fixed: a broad CSS rule had re-enabled pointer events on the full-size handles SVG, which intermittently covered a comment. Overlay surfaces now ignore pointers unless the actual control opts in.

2026-08-11 02:55 CDT — test-only coverage, mobile plan and feed research

Added regression coverage without changing production code. Pytest grew from 96 to 103 cases: Yahoo H1 seed bars must form the correct CME-session daily OHLCV, manual alerts must remain disarmed through rebuild and restart, daily anchors must follow Eastern DST in UTC, and two WebSocket clients must keep their layer preferences isolated from one another and from global alert state. The browser suite grew from 15 to 19 cases: partial or malformed persisted preferences must still boot, volume must remain paired with candles on its own scale across timeframe changes, and Backspace/Delete in a rename input must edit text rather than delete the drawing. Full results: 103 backend and 19 browser tests passing.

One proposed test exposed a current defect and was deliberately not committed as a failing test: withinPlot() compares against the outer chart element, which includes the right price axis. Hovering that axis can therefore leave the snap dot visible at the plot edge. The production fix is to compare plot-relative coordinates against the measured plot canvas width and height, then add the page-pixel price-axis case. Startup rebuild-count coverage is also deferred until the bulk-seeding optimization exists; a duration assertion against today's slow startup would encode the problem rather than protect a fix.

Local and deploy test commands now live in README. Browser E2E stays local because it creates and deletes drawings; production verification uses bin/wait-deploy, /api/health, /api/version, and an authenticated /api/status smoke check.

The mobile audit is recorded in docs/mobile_enhance.md. Its recommended first steps are one coordinate path for every gesture, touch-sized invisible hit areas, persistent first-anchor feedback, and a sticky mobile tool rail before more ambitious gesture changes.

No durable unauthenticated source of free real-time CME /ES data was found. Schwab remains the verified entitled live source and Yahoo the practical delayed seed source. Tastytrade/dxLink is the best free-with-broker-account candidate to test next; IBKR is the strongest low-cost fallback rather than a free one. CME, TradingView and Barchart free pages are delayed displays, not licensed backend APIs. For freshness UI, existing bar WebSocket messages can supply browser receipt time without another call, but exact trade time and source heartbeat are not retained yet. The best placement is the existing status strip below the chart, with a compact mobile form such as Updated 3s ago · Live.

2026-08-11 03:19 CDT — make recommendations discoverable

Added docs/NEXT_STEPS.md as the concise home for deferred production fixes, freshness UI semantics, mobile priorities and market-data alternatives. Linked it, the implementation/session log and docs/mobile_enhance.md directly from AGENTS.md, which every coding agent reads before working in this repository. This keeps current advice visible without turning the historical implementation plan into an undifferentiated backlog.

2026-08-11 03:39 CDT — one price scale, left labels and editable-line snapping

The left side was introduced to keep daily moving-average, VWAP and prior-day labels away from the intraday labels on the right. That decluttering intent was correct; implementing it as a second Lightweight Charts price scale was not. Each scale autoscaled independently once its own overlays were attached, so the same numeric price could occupy a different y-coordinate on each side. Daily and intraday structure then looked directly comparable while being geometrically unrelated.

All price-bearing series and flat levels now share the candle series' right scale. Long-term labels remain on the left as DOM overlays whose y positions are computed through candles.priceToCoordinate(), so they preserve the intended separation without creating another coordinate system. Context series suppress their built-in right-side titles. A browser invariant verifies every overlay's scale id, every flat level's host, and sub-pixel agreement between an MA price and the candle coordinate for that same price.

Editing a selected trendline had a separate defect: handle dragging used element-relative x, snapped only to the nearest bar time, and retained a raw cursor price. Initial placement used plot-relative coordinates and high/low snapping, so moving an anchor could visibly jump off the bar and change the line to an unusable slope. Placement, selection, handle dragging and context actions now all use eventPoint() against the measured plot canvas. Handle movement passes through the same snapPoint() high/low rule and displays the same snap feedback as placement. This also closes the known right-price-axis containment bug; a page-region test now proves the dot disappears over the actual axis and returns over the plot.

The 30-minute chart had only about one week of history because it was derived solely from Yahoo's eight-day 1-minute seed. The 730-day hourly seed cannot reconstruct half-hour candles. Yahoo's native 30m interval was verified live at range=60d (2,843 bars on 2026-08-11), so startup now seeds H1/730d, M30/60d, then M1/8d. The finer minute aggregation replaces the recent overlap; the native feed supplies older 30-minute bars. Regression tests pin both Yahoo's native interval request and the startup seed sequence. Final verification: 105 backend tests and 22 browser tests passing.

Daily moving averages had also been rendered with WithSteps on every base timeframe. Holding a daily value constant is correct when projecting it over intraday candles, but on the daily chart it made the SMA itself look like a staircase. MA line type is now timeframe-aware: stepped on intraday charts and simple point-to-point lines on 1d, with both modes pinned by browser coverage.

2026-08-11 — diagnostic captures moved behind auth

The capture endpoints were split: uploading required a token, retrieving did not. That was deliberate — an unguessable id acting as a capability URL, with a test asserting it — and the reasoning was sound: it lets someone debugging fetch a capture without holding the chart password.

Changed anyway, because of what a capture is. getDisplayMedia returns a picture of somebody's screen, and preferCurrentTab is a preference rather than a constraint, so a mis-click shares a different window. 72 bits of entropy stops guessing, but a capability URL still escapes through everything that records URLs: proxy and access logs, browser history, a link pasted into a chat.

Retrieval now uses the same dependency as the rest of the API, which already accepts the session cookie — so a browser that is logged in needs nothing extra, which was the requirement. An agent on the server reads the files directly from the capture directory, and one working over HTTP presents the API token. Neither path got harder, which is why the trade was worth making.

Both handlers moved from meta.py to routes.py. meta.py is the deliberately unauthenticated router — health, version, login and logout — and a screenshot endpoint did not belong in that company.

2026-08-11 — alerts carry a number and a local time

A push saying "zone at 7784" and a log row saying the same thing were impossible to line up when several fired together. Alerts are now numbered server-side, and the number appears in both the ntfy body and the Events list.

The number has to come from the server. The browser already had a counter, but it restarts on reload and differs between tabs, so it can never agree with a phone. It is persisted alongside the alert cooldown state, because numbers restarting from 1 after a deploy would collide with what is already sitting in a phone's notification history — which meant changing that file from a bare list to an object, with the loader still accepting the old shape.

Pushes also carry a timestamp now, in the configured zone rather than the server's. ALERT_TIMEZONE defaults to America/Chicago. Containers run UTC, and a push that says 02:14 to someone whose wall clock reads 21:14 costs a translation every time. The browser keeps formatting its own times locally, so only the push needed a zone.

A mistake worth recording: number and at were first declared before tripped in the Alert dataclass, which broke every positional caller with "got multiple values for argument 'number'". New fields on a dataclass with positional callers go last. Separately, the first version of the timezone test asserted times I had guessed rather than computed — the code was right and the test was wrong, which is worth checking before assuming a failure is a bug.

2026-08-11 — e2e flakiness is now a pattern, not a test

Two consecutive full e2e runs failed different tests: first "diagnostic capture uploads a PNG", then "a symbol can be dropped at a price", "the snapped extreme decides the side" and "dragging an anchor uses the same snapping". Each passes in isolation. Four different tests failing randomly is a harness problem rather than four bad tests.

Ruled out: rebuild coalescing. The mutating routes still call rebuild_levels() directly and synchronously — only on_bar defers through request_rebuild — so a drawing created over the API still has its levels rebuilt before the response returns. Flakiness was also observed before that change landed.

The likely candidates, untested: tests share one live dev stack whose market feed keeps moving underneath them, and cleanup runs against a store that other tests and any open browser are also mutating. Worth fixing before the suite is trusted, because a suite that fails randomly trains people to re-run it, and a re-run is indistinguishable from a fix.

2026-08-11 — flaky browser cases quarantined

The cases named above are skipped rather than allowed to make the full suite nondeterministic. The symbol failure was reproduced: an older persisted, off-screen symbol occupied the same edge position and intercepted the click intended for the symbol created by the test. That test is unsound while it shares a drawing store with arbitrary browser state. The two trendline cases address a moving live feed through fixed viewport fractions or a bar selected before a multi-step gesture, so they need stable fixture data or stronger gesture-local targeting. The diagnostic capture case needs deterministic display-media and upload-completion boundaries.

The live-edge viewport case joined the quarantine after the same pattern was observed directly: its helper reads candle data and the visible range in separate chart API calls while the live feed can update between them. It then requires exact timestamp equality, so a legitimate new bar makes the assertion compare two different instants. It needs one stable snapshot or a tolerance that still catches the original ten-hour regression before it is sound.

These are explicit node:test skips with reasons, not deleted coverage. Re-enable each case only after its stated external dependency is removed and repeated full suite runs remain green.

2026-08-11 — a time axis past the last bar, and the Yahoo bar it exposed

Trendlines project into the whitespace right of the last candle, but the axis stopped at the newest bar — so two lines could be seen converging with no way to tell when. Lightweight Charts only labels times that exist on its scale, so the fix is whitespace data: { time } points with no value, which extend the scale and draw nothing. timeToCoordinate now answers out there too, which the projection layer wants anyway.

Future times repeat the most recent bar interval, matching what timeAtIndex already does for the projections. That drifts across the daily halt and the weekend, because futures do not trade through them. Consistency with the projection matters more than being right in the abstract — an axis disagreeing with the line drawn above it would be worse — but session-accurate projection needs the server's session rules, which the client does not have.

The interesting part was what this exposed. Padding that should have been five bars measured as sixty-seven, because both the padding and the future slots took the interval from the gap between the last two bars. That gap was two seconds on a one-minute chart, so barInterval now takes a median over recent bars, which ignores a ragged tail.

Chasing why the last gap was two seconds found a real data bug, live in production, which is still on Yahoo. Yahoo stamps the in-progress candle with the moment of the request rather than its bucket start, and the poller emitted anything newer than the last thing it sent — so every poll appended a new "1m" bar a few seconds after the previous one:

04:38:11  aligned=False
04:38:50  aligned=False
04:39:00  aligned=True     <- the only real bar
04:39:15  aligned=False
gaps: 10, 15, 30, 15, 2, 26, 20, 12, 8, 17 seconds

Timestamps are bucketed on parse now, and the final candle is emitted unclosed so it revises the current minute instead of entering the aggregator — which would otherwise add its volume to every higher timeframe on every poll. last_emitted tracks the newest settled bar, so the forming minute is re-sent each poll and once more when it closes. Live check after the fix: eight consecutive bars, all aligned, all sixty seconds apart.

Worth generalising: a measurement that looks absurd is worth chasing rather than clamping. The absurd number here was "67 bars of padding", and the bug behind it had nothing to do with padding.