chart/AGENTS.md
Chris Amow 372617b08c Split the plan from the log of what actually happened
One file was trying to be two things: a spec written to be executed
top-to-bottom, and a dated record of everything that went wrong on the way. At
1,882 lines it did neither well, and the log was 36% of it — which is why the
plan's opening went unmaintained for days while the log grew every hour.

docs/plan.md keeps the decisions and the reasoning behind them, including the
risk register. docs/implementation.md takes the dated entries: the problems, the
wrong theories, the measurements that settled them. Git already says what
changed; that file says why it was hard, which is the part worth reading before
debugging something similar. Most entries describe something that looked like
one bug and turned out to be another.

Each points at the other, and the four referring files — AGENTS.md, README.md,
NEXT_STEPS.md and async_refactor.md — now point at whichever half they meant.
Git tracked the rename, so history follows plan.md rather than starting over.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:29:47 -05:00

135 lines
6.2 KiB
Markdown

# Working on this repo
## Read current context first
Before planning work, read [`docs/plan.md`](docs/plan.md) for the decisions
and the reasoning behind them, and
[`docs/implementation.md`](docs/implementation.md) for the dated record of what
actually went wrong and how it was resolved — that one is the faster read when
debugging, because most entries describe something that looked like one bug and
turned out to be another. Then
[`docs/NEXT_STEPS.md`](docs/NEXT_STEPS.md) for current recommendations and known
deferred fixes. Mobile interaction work also has its own detailed plan in
[`docs/mobile_enhance.md`](docs/mobile_enhance.md).
## Tests earn their place by catching a real bug
When a bug is found, ask whether a unit test could reasonably have caught it. If
yes, write that test with the fix. If no — a rendering artefact, a browser
quirk, a data-source oddity — say so and don't add one.
The bar is "would this have failed before the fix, and would it fail again if
someone reintroduced it". Tests that restate the implementation, assert
constructor defaults, or exercise paths nothing depends on are noise; they make
the suite slow to run and expensive to change, which is how a suite stops being
trusted.
What has actually paid off here: bar aggregation and bucket boundaries, the
store's replace-vs-append rules, level and alert arithmetic, parsing real
market-data payloads (fixtures are trimmed real responses, not invented), and
the invariants that would otherwise be silent — a comment must never become a
level, a tick must never overwrite a settled bar, volume must be counted once.
Name the test after the failure, not the function: `test_a_tick_cannot_overwrite
_a_settled_bar` beats `test_put`.
## Where things run
**The agent works on a remote machine over SSH. The user's browser runs on a
different machine.** Consequences, all learned the hard way:
- You cannot see the user's screen, console, or cursor. Screenshots and pasted
console output are the only window into it. Browser extensions that drive
"your" Chrome do not help — they attach to the machine the browser is on.
- Headless Chromium here renders on server hardware: different screen, window
size and device pixel ratio from the user's. "Works in my headless run" is not
evidence that it works for them. When a UI bug will not reproduce, match their
viewport and `deviceScaleFactor` explicitly before concluding anything.
- The dev stack is served to them over the network (e.g. `hera.local:8010`),
which is the same app the headless browser reaches as `http://api:8000`.
When a visual bug resists reproduction, prefer putting the numbers **on screen**
in the app over asking for another console paste — one screenshot then carries
the whole diagnosis.
## Verify UI in a real browser
Chart bugs are invisible from the outside — the API, the socket and the
frontend source can each be correct while the screen is wrong. Drive the
Playwright container against the dev stack:
```
docker exec -i chart-playwright-1 node - <<'EOF'
const { chromium } = require('/usr/lib/node_modules/playwright');
// launch with args:['--lang=en-US'] — see below
EOF
```
**Always launch Chromium with `args: ['--lang=en-US']`.** The container has no
usable locale, so Chromium reports `en-US@posix`, `Intl` throws, and the chart
renders as a blank canvas that looks exactly like a broken app.
`window.__chart` is a deliberate debug handle. Querying it separates "the data
is missing" from "the data is off-screen" — which is how a viewport bug that
three passing API checks had missed was finally found.
## Diagnostic mode
Chart geometry bugs live in the browser, which is usually on a different machine
from whoever is debugging them. Rather than asking for console pastes:
```
open the chart with ?diag=1 # remembered until ?diag=0
docker compose logs api | grep SNAPDBG
```
With it on, every snap the trendline tool computes is posted to
`/api/debug/snap` and logged server-side — the cursor's time, price and x, the
snapped time and price, how many bars were held, the first and last bar, and the
chart's width. Throttled to about one a second. It reads the client's own
numbers, which is exactly what "works in my headless run" cannot tell you.
Extend it when the next geometry puzzle appears; the endpoint takes whatever
fields `SnapReport` declares.
## Direction of travel
Two live planning documents, both written to be refactored toward rather than
implemented in one go:
- `docs/async_refactor.md` — nothing blocks the event loop. P0 and P1 are done;
`/api/status` reports `loop_lag_ms`, and a rise there is the signal.
- `docs/multi_user.md` — separate people with their own drawings and alerts,
authenticated by OIDC. Read it before adding state to `Runtime`: new state is
either genuinely shared (market data) or belongs to a user, and knowing which
now is much cheaper than untangling it later.
Do not build local user accounts. The destination is OIDC, so password storage
would be written and then deleted.
## Running tests
```
docker exec chart-api-1 sh -c "cd /app && python -m pytest -q"
```
pytest + pytest-asyncio, declared in `requirements-dev.txt`. Tests live in
`tests/`, import from `app.*`, and use `tmp_path` for anything that persists.
Async paths are driven with `asyncio.run(...)` directly rather than async test
markers.
## Things that will cost you an hour
- **Never write scratch `.py` files into the repo root.** It is bind-mounted, so
`--reload` restarts the app, and startup takes ~82 seconds. Pipe throwaway
scripts over stdin instead: `docker exec -i chart-api-1 python - <<'EOF'`.
Screenshots into `artifacts/` are safe; only `.py` triggers the reloader.
- **Dev and production keep separate drawing stores.** Dev writes
`data/manual_lines.json`; production has its own Coolify volume. A fix that
"didn't land" is often the other store.
- **Rebuild the image after touching `requirements.txt`.** The bind mount makes
source edits look live while an added dependency is simply absent.
- **A deploy resets alert cooldowns**, so production may re-alert on whatever
price is sitting on. There is no durable state yet.
- Times are epoch seconds, UTC, everywhere. Only the display is localised —
never shift the stored values.