Ziyan Tracker is a small family tool: it pulls a student’s assignments from the school’s portal, pushes deadlines to a shared “Homework” calendar, shows everything on a password-protected dashboard, and emails a digest at 7am. It worked, with one blind spot.
A lot of what matters at school never reaches the portal. It arrives in the class parents’ WhatsApp group: a photo of a flyer, a PDF permission form, “don’t forget picture day is moved to Thursday.”
This post covers the first four phases of fixing that: what we built, what broke when it met the real world, and how we fixed it. All of it happened in one morning, in nine commits.
The rule everything else follows from
Anything in the tracker’s items table that has a date is pushed to the family calendar and included in the morning email, automatically. That is exactly what you want for the school portal, and exactly what you don’t want for something an AI guessed from a parent’s chat message.
So the design rule is: nothing from an untrusted source may ever write to items. Messages become candidates. A parent confirms a candidate, and only then does it become an item. Every phase below is built around that one constraint.
Phase 1: Foundation
The starting point was honest but fragile: a Flask app deployed by copying files from a laptop to a server. No git repository, no tests, no record of where any item came from beyond a source string.
I put the deployed code under git exactly as it was running, added pytest with a throwaway database per test, switched SQLite to WAL mode (several processes now write concurrently), and added four tables:
sources: who we listen to, and how much we trust themingested_items: one row per thing received, with full provenance and the original file stored on disk under its SHA-256candidates: what was extracted, awaiting a parent’s reviewsource_runs: per-source sync counters, for a future “Source Health” panel
Confirmed items gained an origin_ingested_id column, so any entry on the dashboard can answer “where did this come from?”
Phase 2: Make the existing upload the first customer
Before adding a new source, I moved an old one onto the new pipeline. The dashboard’s manual upload (text, PDF or screenshot, read by Claude, then a checkbox review page) used to hold the file in memory for one request and keep review state only in an HTML form. Now it writes an ingested_item, saves the file, creates candidates, and renders the same review page from the database. Same UX, but with provenance, deduplication (upload the same flyer twice and you get the same row), and Word/Excel support.
The point was to prove the pipeline with something I already understood. WhatsApp would later arrive as just another source feeding the same tables.
I also wrote a deploy script that refuses to run with uncommitted changes, backs up the database and the current code, ships only tracked files, and prints the rollback command.
Phase 3: A connector that cannot speak
There is no official API for reading a personal WhatsApp group, so this uses an unofficial library (whatsapp-web.js, which drives WhatsApp Web in headless Chromium). That is a terms-of-service grey area with a non-zero risk to the account, and I went in with eyes open: read-only, one group, low traffic.
“Read-only” is enforced by construction rather than by good intentions:
- The rest of the code sees a four-method interface:
listChats,fetchHistory,onMessage,downloadMedia. There is no send, react, delete or mark-as-read to call. - Exactly one file may import the WhatsApp library. A test fails if any other file does.
- Another test scans that file for every write-capable call the library offers, and a companion test proves the scanner actually catches them.
- Message content is only reachable for groups on an allowlist of group IDs. Names can be changed by any admin; IDs can’t. Direct chats are never even listed.
It runs as its own service, under its own unprivileged user, with root-owned code it cannot modify, bound to localhost, behind a token, inside a hardened systemd unit.
Phase 4: Ninety days of history
The tracker gained a localhost-only POST /internal/ingest endpoint: its own token, refused outright if the request arrived through the public reverse proxy, idempotent per (group, message ID). The connector gained a backfill that walks the allowlisted group over a 7, 30, 60 or 90-day window, downloads images and documents (not voice notes, video or stickers), and reports each run. The connector will only post to a loopback address, so message content cannot leave the machine by misconfiguration.
Because the two halves are written in different languages (TypeScript and Python), one test runs the real backfill against the real Flask app on a temporary database. It checks the contract at the boundary, which is where the bugs turned out to be.
What broke, and what fixed it
The server had no usable Node. The only Node on the box lived inside another user’s home directory, which a locked-down service account cannot and should not read. The installer now downloads a pinned Node release and verifies its checksum, used by nothing else on the machine.
Puppeteer’s Chrome download “succeeded” with no Chrome. Under Node 26, its zip extraction dies silently after four files. The download itself is fine, so the installer falls back to plain unzip, then checks for missing shared libraries before starting anything.
Chromium’s sandbox versus Ubuntu 24.04. Unprivileged user namespaces are blocked, and the setuid fallback conflicts with NoNewPrivileges. Chromium runs with --no-sandbox, and the systemd unit becomes the sandbox: read-only filesystem, private /tmp, no home directories, a memory cap, a dedicated user.
First contact with a real account: {"error": "r"}. That single letter is a minified exception from inside WhatsApp Web. The library’s getChats() asks WhatsApp’s servers for the metadata of every group in the account and wraps it all in Promise.all, so one group the account had left rejected the entire list. It was also far more traffic than a quiet observer should generate. I replaced it with a read of what WhatsApp Web already holds in memory: no server queries, and history loading now stops as soon as it has covered the requested window instead of paging back thousands of messages.
Fifteen messages read, zero stored. The next run reported 13 failures in 0.17 seconds. Every message had reached the tracker without an ID. WhatsApp’s message-key objects don’t survive being copied out of the browser page intact; the serialized ID doesn’t make the trip. I now pick every field inside the page, with fallbacks, and fail loudly naming the object’s shape if none work. The same weakness would have silently marked every attachment “unavailable”, so that path was fixed too.
The failed run said ok: true. Arguably the worst bug of the day, because it is the kind that hides the others. A run with failures now reports itself as failed and lists the reasons, and the tracker says why it rejected a message. I only had to guess once.
Two human-interface mistakes. An example group ID in the instructions looked real enough to be pasted and allowlisted. And an install was run while newer work was still uncommitted, so the installer (correctly) shipped the older commit and the new command printed help text. Fixes: the allow command now refuses any ID the linked account doesn’t actually have, there is a disallow, and the habit now is never to ask for an install until the tree is committed.
A permission I chose not to grant. Deploys need root, and the development account has none, so a human runs three commands per round. The tempting shortcut, passwordless sudo on the deploy scripts, would have been full root in disguise, because those scripts live in a directory the same account can edit. I left it manual.
Where it stands
- 45 messages and 5 attachments from one allowlisted group, reaching back 11 weeks, stored with provenance
- 0 candidates, 0 calendar entries, 0 digest lines from WhatsApp: exactly as designed
- 38 Python tests and 35 connector tests, several of which run in a real browser page
- Roughly 400 lines of TypeScript in the connector
Three lessons I would repeat to anyone doing similar work. Test at the boundary between systems, because both real bugs lived there. Make failures red and self-describing before you need them to be. And when a library does something convenient on your behalf, read what it actually does; here the convenient call was both the bug and the riskiest traffic pattern.
Next: phase 5
So far the tracker has collected messages without understanding any of them. Phase 5 decides which ones matter for school, turns those into candidates for a parent to review, and forgets the rest.